Sign inSign up

zilutian/data-serving

By zilutian

•Updated over 7 years ago

CloudSuite data serving benchmark

Image
0

427

zilutian/data-serving repository overview

Docker images of CloudSuite (cloudsuite.ch) data-serving workload for aarch64 architecture. Please visit CloudSuite's website for more background information.

Relative source code: github parsa-epfl/cloudsuite/benchmarks/data-serving, branch: ZiluTian-data-serving-arm

⁠Versions

Ubuntu 18.04

Cassandra 3.11.2

YSCB 0.14.0

⁠Create a docker network to connect the server and client
docker network create serving_network 
⁠Generate a server container (Cassandra)
docker run --rm -p 9042:9042 --name cassandra-server --net serving_network zilutian/data-serving:server 

Please remember to publish the port 9042 (-p 9042:9042) for client to connect. Upon success, you should see following message

 ...
INFO  [main] Server.java:156 - Starting listening for CQL clients on /0.0.0.0:9042 (unencrypted)...
INFO  [main]  CassandraDaemon.java:529 - Not starting RPC server as requested. Use JMX (StorageService->startRPCServer()) or nodetool (enablethrift) to start it
INFO  [OptionalTasks:1] 515 CassandraRoleManager.java:356 - Created default superuser role 'cassandra'
⁠Generate a client container (YSCB) and connect it with the server container
docker run --rm --name cassandra-client --net serving_network zilutian/data-serving:client "cassandra-server" 

Upon success, you should see the latency report similar to the following

[OVERALL], RunTime(ms), 4709
[OVERALL], Throughput(ops/sec), 212.35931195582927
[TOTAL_GCS_PS_Scavenge], Count, 1
[TOTAL_GC_TIME_PS_Scavenge], Time(ms), 41
[TOTAL_GC_TIME_%_PS_Scavenge], Time(%), 0.8706731790189
[TOTAL_GCS_PS_MarkSweep], Count, 0
[TOTAL_GC_TIME_PS_MarkSweep], Time(ms), 0
[TOTAL_GC_TIME_%_PS_MarkSweep], Time(%), 0.0
[TOTAL_GCs], Count, 1
[TOTAL_GC_TIME], Time(ms), 41
[TOTAL_GC_TIME_%], Time(%), 0.8706731790189
[CLEANUP], Operations, 1
[CLEANUP], AverageLatency(us), 2288640.0
[CLEANUP], MinLatency(us), 2287616
[CLEANUP], MaxLatency(us), 2289663
[CLEANUP], 95thPercentileLatency(us), 2289663
[CLEANUP], 99thPercentileLatency(us), 2289663
[INSERT], Operations, 1000
[INSERT], AverageLatency(us), 1240.247
[INSERT], MinLatency(us), 658
[INSERT], MaxLatency(us), 50847
[INSERT], 95thPercentileLatency(us), 2295
[INSERT], 99thPercentileLatency(us), 3783
[INSERT], Return=OK, 1000
⁠Alternatively, you can also run benchmark after bringing up Cassandra server
docker run --rm -e RECORDCOUNT=10000 -e THREADS=10 -e TARGET=100 --name cassandra-client --net serving_network zilutian/data-serving:client "cassandra-server"

Update the arguments' values as you see fit. For a full description regarding the parameters RECORDCOUNT, THREADS, and TARGET, please refer to YCSB's github wiki page https://github.com/brianfrankcooper/YCSB/wiki/Running-a-Workload⁠. Please be aware of the coordinated omission problem related to YCSB.

Upon success, you should see latency report similar to the following

[OVERALL], RunTime(ms), 4478
[OVERALL], Throughput(ops/sec), 223.3139794551139
[TOTAL_GCS_PS_Scavenge], Count, 1
[TOTAL_GC_TIME_PS_Scavenge], Time(ms), 19
[TOTAL_GC_TIME_%_PS_Scavenge], Time(%), 0.4242965609647164
[TOTAL_GCS_PS_MarkSweep], Count, 0
[TOTAL_GC_TIME_PS_MarkSweep], Time(ms), 0
[TOTAL_GC_TIME_%_PS_MarkSweep], Time(%), 0.0
[TOTAL_GCs], Count, 1
[TOTAL_GC_TIME], Time(ms), 19
[TOTAL_GC_TIME_%], Time(%), 0.4242965609647164
[READ], Operations, 527
[READ], AverageLatency(us), 1146.965844402277
[READ], MinLatency(us), 639
[READ], MaxLatency(us), 34751
[READ], 95thPercentileLatency(us), 1988
[READ], 99thPercentileLatency(us), 2927
[READ], Return=OK, 527
[CLEANUP], Operations, 1
[CLEANUP], AverageLatency(us), 2268160.0
[CLEANUP], MinLatency(us), 2267136
[CLEANUP], MaxLatency(us), 2269183
[CLEANUP], 95thPercentileLatency(us), 2269183
[CLEANUP], 99thPercentileLatency(us), 2269183
[UPDATE], Operations, 473
[UPDATE], AverageLatency(us), 1083.169133192389
[UPDATE], MinLatency(us), 599
[UPDATE], MaxLatency(us), 12295
[UPDATE], 95thPercentileLatency(us), 1854
[UPDATE], 99thPercentileLatency(us), 5223
[UPDATE], Return=OK, 473

The default is is workloada, 50% of read and 50% of update. You can specify different workload types by modifying the entrypoint file.

⁠Debug version (v2)

Cassandra client only initializes the keyspace. Any load and run is done through commands

docker run -it -d --cpuset-cpus=$CLIENT_CPUS --name $CLIENT_CONTAINER --net serving_network client "cassandra-server"
docker exec -it $CLIENT_CONTAINER bash -c "/ycsb/bin/ycsb load cassandra-cql -p hosts=$SERVER_CONTAINER -P /ycsb/workloads/workloada -s -target $TARGET -threads $THREADS_LOAD -p recordcount=$RECORDS"
⁠Debug tips

Log files (server):

/apache-cassandra-3.11.2/logs

From the client:

docker exec -it cassandra-client cqlsh cassandra-server 
cqlsh> describe keyspaces; 
cqlsh> describe ycsb; 
cqlsh > select y_id from ycsb.usertable;
cqlsh> tracing on;  // turn on the trace to get an activity log when executing command 
cqlsh > select * from ycsb.usertable where y_id='user...'  // select an y_id 

Expected output (trace for reading data):

                                             Executing single-partition query on usertable [ReadStage-2] | 2019-08-16 13:43:39.103000 | 172.31.0.2 |            560 | 172.31.0.3
                                                              Acquiring sstable references [ReadStage-2] | 2019-08-16 13:43:39.103000 | 172.31.0.2 |            610 | 172.31.0.3
                                                                 Merging memtable contents [ReadStage-2] | 2019-08-16 13:43:39.103000 | 172.31.0.2 |            672 | 172.31.0.3
                                                    Bloom filter allows skipping sstable 3 [ReadStage-2] | 2019-08-16 13:43:39.104000 | 172.31.0.2 |            896 | 172.31.0.3
                                                               Key cache hit for sstable 2 [ReadStage-2] | 2019-08-16 13:43:39.104000 | 172.31.0.2 |           1119 | 172.31.0.3
                                                    Read 1 live rows and 0 tombstone cells [ReadStage-2] | 2019-08-16 13:43:39.104000 | 172.31.0.2 |           1661 | 172.31.0.3
                                                                                        Request complete | 2019-08-16 13:43:39.105042 | 172.31.0.2 |           2042 | 172.31.0.3

From the server:

docker exec -it cassandra-server /bin/bash 
nodetool cfstat ycsb
nodetool getlogginglevels 
nodetool tablehistograms ycsb usertable

Expected output (tablehistograms):

Percentile  SSTables     Write Latency      Read Latency    Partition Size        Cell Count
                              (micros)          (micros)           (bytes)                  
50%             0.00             14.24              0.00              1109                10
75%             0.00             14.24              0.00              1109                10
95%             0.00             17.08              0.00              1109                10
98%             0.00             20.50              0.00              1109                10
99%             0.00             20.50              0.00              1109                10
Min             0.00              3.97              0.00               925                 9
Max             0.00         268650.95              0.00              1109                10

Helpful links to gain a deeper understanding of Cassandra:

Data model of Cassandra: 
https://dzone.com/articles/cassandra-data-modeling-primary-clustering-partiti

How is data read/written: 
https://pandaforme.gitbooks.io/introduction-to-cassandra/content/how_is_data_read.html

Read the code base!
http://prettyprint.me/prettyprint.me/2010/05/02/understanding-cassandra-code-base/index.html

Tag summary

Content type

Image

Digest

Size

554.6 MB

Last updated

over 7 years ago

docker pull zilutian/data-serving:client-amd64-v2