Sign inSign up

zilutian/web-search

By zilutian

•Updated over 7 years ago

CloudSuite web search benchmark

Image
0

159

zilutian/web-search repository overview

(Updating the server image in process. Please check later)

⁠Instructions for using Web Search benchmark

Source: cloudsuite.ch

⁠Instructions for generating index

Replace $WIKI_DUMPS with your folder name. Please install aria2 and pbzip2. Alternatively, you can use wget and bzip2 instead, though they are slower. Please make sure you have enough disk space in the $WIKI_DUMPS folder. File enwiki-latest-pages-articles-multistream-index.txt.bz2 is 16GB at the time of writing (March 2019), but the size is subject to change.

mkdir $WIKI_DUMPS && cd $_
aria2c https://dumps.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles-multistream-index.txt.bz2
aria2c https://dumps.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles-multistream.xml.bz2 
pbzip2 -dckvm300 enwiki-latest-pages-articles-multistream-index.txt.bz2 > enwiki-latest-pages-articles-multistream-index.txt

Then run /cloudsuite/benchmarks/web-search/server/generate_index.sh script, whose content is listed here. Variable $NO_PAGES specifies the desired page numbers to generate, which is an indication of the index file size. At the time of writing (March 2019), 3500000 pages result in 11GB data, 7000000 pages result in 30GB.
With pbzip2, you might run into deadlock. So if you encounter "segmentation fault, core dump" and strace indicates that a thread is waiting too long for a lock, please fall back to bzip2.

#!/bin/bash 
NO_PAGES=3500000
INDEX_FILE=$WIKI_DUMPS/enwiki-latest-pages-articles-multistream-index.txt
DUMP_FILE=$WIKI_DUMPS/enwiki-latest-pages-articles-multistream.xml.bz2

DUMPNAME=dump_$NO_PAGES

if [[ $NO_PAGES -ge `wc -l < $INDEX_FILE` ]]; then
        echo "wikipedia dump does not have $NO_PAGES pages"
        exit 1
fi

BYTE_OFFSET=`sed "${NO_PAGES}q;d" $INDEX_FILE | sed "s/:.*//"`

head -c $BYTE_OFFSET $DUMP_FILE > $DUMPNAME.xml.bz2
pbzip2 -dc $DUMPNAME.xml.bz2 > wiki_dump.xml
rm $DUMPNAME.xml.bz2
`echo "</mediawiki>" >> wiki_dump.xml`

You should see a file wiki_dump.xml now in the $WIKI_DUMPS folder. Please check its size and verify it is your intended size.

ls -hl 
⁠Instructions for creating docker network

Create a docker network to connect the search server and client

docker network create search_network 
⁠Instructions for starting the server

First, pull the server image from Dockerhub. No need to specify the architecture, tag "latest" has been created using the Docker manifest, which will pull the correct image based on your architecture automatically.

docker pull zilutian/web-search-server 

Note that $WIKI_DUMPS is the absolute path to the dump folder $WIKI_DUMPS which contains file wiki_dump.xml. The first parameter passed to the image, '11g' is the amount of memory allocated to java process. Please pass the size of your wiki_dump. '1' indicates a single Solr node, which is always the case for index node. 'generate' will ask Solr to generate an index based on the XML passed.

docker run -it --name server -v $WIKI_DUMPS:/home/solr --net search_network -p 8983:8983 zilutian/web-search-server 11g 1 generate 

You should see the following content

Waiting up to 180 seconds to see Solr running on port 8983 [-]

If Solr server hangs after a few minutes, check the log file and see if there is a problem regarding the config of the JVM (e.g. stack size is too small) or other issues. To do that,

docker cp *CONTAINER_ID*:/usr/src/solr-7.7.1/server/logs/*log*  *LOCALFILENAME*

If all is well, you should see

Started Solr server on port 8983 (pid=104). Happy searching!
...
=================================
Index Node IP Address: 172.29.0.2
=================================
...

Alternatively, you can carve out the index generated by Solr based on wiki_dump.xml from the image and save it as a separate Docker image, e.g. "index". Then pass it as volumes-from

docker run -it --name server --volumes-from index --net search_network -p 8983:8983 zilutian/web-search-server 11g 1

From the server container, you can verify the index has been properly created and search request handler is working as expected by issuing a query, as the example below shows for query "hello"

curl http://localhost:8983/solr/cloudsuite_web_search/query?q=hello

From the response, you can see the number of documents found (numFound), and it should not be 0.

⁠Instructions for starting the client

Similarly, let's pull the client image first.

docker pull zilutian/web-search-client

Start a client container using the following command. Substitute INDEX_NODE_IP with the IP address of index node. The four numbers after the server address refer to: the scale, which indicates the number of concurrent clients or the number of threads (50); the ramp-up time in seconds (25), which refers to the time required to warm up the server; the steady-state time in seconds (20), which indicates the time the benchmark is in the steady state; and the rump-down time in seconds (20), which refers to the time to wait before ending the benchmark.

Variable JAVA_HOME can be different depending on your architecture. For x86_64 (amd64), run as

docker run -it -e JAVA_HOME=/usr/lib/jvm/java-8-openjdk-amd64 --name client --net search_network zilutian/web-search-client INDEX_NODE_IP 50 25 20 20 

For aarch64 (arm64), run as

docker run -it -e JAVA_HOME=/usr/lib/jvm/java-8-openjdk-arm64 --name client --net search_network zilutian/web-search-client INDEX_NODE_IP 50 25 20 20 

You should see the build result

Starting Faban Server
...
init:
    [mkdir] Created dir: /usr/src/faban/search/build/classes

compile:
    [javac] /usr/src/faban/search/build.xml:35: warning: 'includeantruntime' was not set, defaulting to build.sysclasspath=last; set to false for repeatable builds
    [javac] Compiling 1 source file to /usr/src/faban/search/build/classes
    [javac] 
    ....

bench.jar:
    [mkdir] Created dir: /usr/src/faban/search/build/lib
      [jar] Building jar: /usr/src/faban/search/build/lib/search.jar

deploy.jar:
BUILD SUCCESSFUL
...
INFO: Ramp up completed
...
INFO: Steady state completed
...
INFO: Ramp down completed
...
INFO: Gathering SearchDriverStats ...
...

In the meanwhile, you can see the search requests from the server terminal

2019-03-15 10:19:12.604 INFO  (qtp1277009227-153) [c:cloudsuite_web_search s:shard1 r:core_node2 x:cloudsuite_web_search_shard1_replica_n1] o.a.s.c.S.Request [cloudsuite_web_search_shard1_replica_n1]  webapp=/solr path=/query params={q=2+expect+spearman&fl=url&lang=en&rows=10} hits=** status=** QTime=**
...

Pay attention to the warning messages when you tune the parameters for rampup time, steady time, and ramp down time. Please also pay attention to the number of hits. There might be some configuration issue if the number of hits is 0.

WARNING: SearchDriverAgent[1]: Rampup too short. Could not run calibration. Please increase rampup by at least 8 seconds

Upon finishing, you should see the output of benchmark result from the client:

<benchResults>
    <benchSummary name="Sample Search Workload" version="0.3">
        <runId>1</runId>
        <startTime>Fri Mar 15 10:15:43 GMT 2019</startTime>
        <endTime>Fri Mar 15 10:19:13 GMT 2019</endTime>
        <metric unit="ops/sec">24.083</metric>
        <passed>true</passed>
    </benchSummary>
    <driverSummary name="SearchDriver">
        <metric unit="ops/sec">24.083</metric>
        <startTime>Fri Mar 15 10:15:43 GMT 2019</startTime>
        <endTime>Fri Mar 15 10:19:13 GMT 2019</endTime>
        <totalOps unit="operations">1445</totalOps>
        <users>50</users>
        <rtXtps>47.1523</rtXtps>
        <passed>true</passed>
        <mix allowedDeviation="0.0000">
            <operation name="GET">
                <successes>1445</successes>
                <failures>0</failures>
                <mix>1.0000</mix>
                <requiredMix>1.0000</requiredMix>
                <passed>true</passed>
            </operation>
        </mix>
        <responseTimes unit="seconds">
            <operation name="GET" r90th="0.500">
                <avg>0.002</avg>
                <max>0.176</max>
                <sd>0.010</sd>
                <p90th>0.003</p90th>
                <passed>true</passed>
                <p99th>0.003</p99th>
            </operation>
        </responseTimes>
        <delayTimes>
            <operation name="GET" type="cycleTime">
                <targetedAvg>2.029</targetedAvg>
                <actualAvg>2.028</actualAvg>
                <min>0.001</min>
                <max>10.000</max>
                <passed>true</passed>
            </operation>
        </delayTimes>
    </driverSummary>
</benchResults>

You can use Ctrl-C to terminate the server process.

⁠Debug tips

The schema used for the benchmark to generate inverted index can be found at /usr/src/solr-7.7.1/server/solr/configsets/_default/conf/managed-schema from the server container
or in /cloudsuite/benchmarks/web-search/server/files/conf/managed-schema in source code for CloudSuite https://github.com/ZiluTian/cloudsuite/blob/zilu-django-workload/benchmarks/web-search/server/files/conf/managed-schema⁠.

Tag summary

Content type

Image

Digest

Size

165.4 MB

Last updated

over 7 years ago

docker pull zilutian/web-search:client-arm64