ML framework built on Harp, a Hadoop plug-in for collective-communication on Big Data and Intel DAAL
477
Harp is a Hadoop plug-in, which rewrites the Hadoop MapReduce framework to achieve in-memory communication among nodes of a distributed environment. It has the following two features
Harp has MPI-like collective communication operations that are highly optimized for big data problems. Harp has efficient and innovative computation models for different machine learning proble
See Details: https://github.com/DSC-SPIDAL/harp/tree/master/harp-daal-app
docker run -it lee212/harp-daal /etc/bootstrap.sh -bash
$HADOOP_HOME/bin/hadoop jar $HADOOP_HOME/harp-app-1.0-SNAPSHOT.jar edu.iu.kmeans.regroupallgather.KMeansLauncher 1000 10 100 5 2 2 10 /kmeans /tmp/kmeans
Number of Map Tasks = 2
Generate data.
Generating data.....
Writing 100 vectors to a file
Directory: /tmp/kmeans created
Write file 4
Write file 6
Write file 0
Write file 7
Write file 1
Write file 2
Write file 5
Write file 3
Write file 9
Write file 8
Generate centroid data./kmeans/centroids/init_centroids
Wrote centroids data to file
Starting Job
Start Job 16:01:36.916
Job configure in 10 miliseconds.
Get map collective runner
17/10/12 16:01:36 INFO client.RMProxy: Connecting to ResourceManager at /0.0.0.0:8032
17/10/12 16:01:37 INFO input.FileInputFormat: Total input paths to process : 10
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: NUMBER OF FILES: 10
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: NUMBER OF MAPS: 2
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: Split on host: f2f5b1c9ba71
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: Split on host: f2f5b1c9ba71
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: Total # of splits: 2
17/10/12 16:01:37 INFO mapreduce.JobSubmitter: number of splits:2
17/10/12 16:01:38 INFO mapreduce.JobSubmitter: Submitting tokens for job: job_1507838336882_0001
command: $JAVA_HOME/bin/java -Dlog4j.configuration=container-log4j.properties -Dyarn.app.container.log.dir=<LOG_DIR> -Dyarn.app.container.log.filesize=0 -Dhadoop.root.logger=INFO,CLA -Dhadoop.root.logfile=syslog -Xmx1024m org.apache.hadoop.mapreduce.v2.app.MapCollectiveAppMaster 1><LOG_DIR>/stdout 2><LOG_DIR>/stderr
17/10/12 16:01:38 INFO impl.YarnClientImpl: Submitted application application_1507838336882_0001
17/10/12 16:01:38 INFO mapreduce.Job: The url to track the job: http://f2f5b1c9ba71:8088/proxy/application_1507838336882_0001/
17/10/12 16:01:38 INFO mapreduce.Job: Running job: job_1507838336882_0001
17/10/12 16:01:44 INFO mapreduce.Job: Job job_1507838336882_0001 running in uber mode : false
17/10/12 16:01:44 INFO mapreduce.Job: map 0% reduce 0%
17/10/12 16:01:49 INFO mapreduce.Job: map 100% reduce 0%
17/10/12 16:01:50 INFO mapreduce.Job: Job job_1507838336882_0001 completed successfully
17/10/12 16:01:50 INFO mapreduce.Job: Counters: 30
File System Counters
FILE: Number of bytes read=0
FILE: Number of bytes written=231224
FILE: Number of read operations=0
FILE: Number of large read operations=0
FILE: Number of write operations=0
HDFS: Number of bytes read=818759
HDFS: Number of bytes written=18161
HDFS: Number of read operations=25
HDFS: Number of large read operations=0
HDFS: Number of write operations=6
Job Counters
Launched map tasks=2
Other local map tasks=2
Total time spent by all maps in occupied slots (ms)=6022
Total time spent by all reduces in occupied slots (ms)=0
Total time spent by all map tasks (ms)=6022
Total vcore-seconds taken by all map tasks=6022
Total megabyte-seconds taken by all map tasks=6166528
Map-Reduce Framework
Map input records=10
Map output records=0
Input split bytes=530
Spilled Records=0
Failed Shuffles=0
Merged Map outputs=0
GC time elapsed (ms)=110
CPU time spent (ms)=1740
Physical memory (bytes) snapshot=364224512
Virtual memory (bytes) snapshot=4076158976
Total committed heap usage (bytes)=514850816
File Input Format Counters
Bytes Read=0
File Output Format Counters
Bytes Written=0
end Jod 16:01:50.624
Job finishes in 13708 miliseconds.
Total K-means Execution Time: 13708
$HADOOP_HOME/bin/hdfs dfs -get /kmeans .
source /opt/intel/bin/compilervars.sh intel64
source $HARP_GITHUB_REPO/harp-daal-app/__release__lnx/daal/bin/daalvars.sh intel64
export LIBJARS=${DAALROOT}/lib/daal.jar
cd $HADOOP_HOME
cd $HADOOP_PREFIX
hadoop jar harp-daal-app-1.0-SNAPSHOT.jar edu.iu.daal_kmeans.regroupallgather.KMeansDaalLauncher -D mapred.child.java.opts=-Xmx800M -libjars $LIBJARS 5000 100 100 5 2 4 10 1024 /kmeans-P5000-C100-D100-F5-ITR10-N2 /tmp/kmeans true
Number of Map Tasks = 2
Generate data.
Generating data.....
Writing 500 vectors to a file
Directory: /tmp/kmeans created
Write file 2
Write file 0
Write file 7
Write file 4
Write file 6
Write file 3
Write file 5
Write file 1
Write file 8
Write file 9
File path : /tmp/kmeans and name :data_3
Full Path :/tmp/kmeans/data_3
Deleting file : /tmp/kmeans/data_3
File path : /tmp/kmeans and name :data_7
Full Path :/tmp/kmeans/data_7
Deleting file : /tmp/kmeans/data_7
File path : /tmp/kmeans and name :data_4
Full Path :/tmp/kmeans/data_4
Deleting file : /tmp/kmeans/data_4
File path : /tmp/kmeans and name :data_8
Full Path :/tmp/kmeans/data_8
Deleting file : /tmp/kmeans/data_8
File path : /tmp/kmeans and name :data_2
Full Path :/tmp/kmeans/data_2
Deleting file : /tmp/kmeans/data_2
File path : /tmp/kmeans and name :data_6
Full Path :/tmp/kmeans/data_6
Deleting file : /tmp/kmeans/data_6
File path : /tmp/kmeans and name :data_0
Full Path :/tmp/kmeans/data_0
Deleting file : /tmp/kmeans/data_0
File path : /tmp/kmeans and name :data_5
Full Path :/tmp/kmeans/data_5
Deleting file : /tmp/kmeans/data_5
File path : /tmp/kmeans and name :data_1
Full Path :/tmp/kmeans/data_1
Deleting file : /tmp/kmeans/data_1
File path : /tmp/kmeans and name :data_9
Full Path :/tmp/kmeans/data_9
Deleting file : /tmp/kmeans/data_9
Deleting Directory : /tmp/kmeans
Generate centroid data./kmeans-P5000-C100-D100-F5-ITR10-N2/centroids/init_centroids
Wrote centroids data to file
Starting Job
Start Job 21:56:56.771
Job configure in 10 miliseconds.
Get map collective runner
17/11/02 21:56:56 INFO client.RMProxy: Connecting to ResourceManager at /0.0.0.0:8032
17/11/02 21:56:57 INFO input.FileInputFormat: Total input paths to process : 10
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: NUMBER OF FILES: 10
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: NUMBER OF MAPS: 2
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: Split on host: 112dea2c5eba
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: Split on host: 112dea2c5eba
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: Total # of splits: 2
17/11/02 21:56:57 INFO mapreduce.JobSubmitter: number of splits:2
17/11/02 21:56:57 INFO mapreduce.JobSubmitter: Submitting tokens for job: job_1509673169537_0007
command: $JAVA_HOME/bin/java -Dlog4j.configuration=container-log4j.properties -Dyarn.app.container.log.dir=<LOG_DIR> -Dyarn.app.container.log.filesize=0 -Dhadoop.root.logger=INFO,CLA -Dhadoop.root.logfile=syslog -Xmx1024m org.apache.hadoop.mapreduce.v2.app.MapCollectiveAppMaster 1><LOG_DIR>/stdout 2><LOG_DIR>/stderr
17/11/02 21:56:57 INFO impl.YarnClientImpl: Submitted application application_1509673169537_0007
17/11/02 21:56:57 INFO mapreduce.Job: The url to track the job: http://112dea2c5eba:8088/proxy/application_1509673169537_0007/
17/11/02 21:56:57 INFO mapreduce.Job: Running job: job_1509673169537_0007
17/11/02 21:57:02 INFO mapreduce.Job: Job job_1509673169537_0007 running in uber mode : false
17/11/02 21:57:02 INFO mapreduce.Job: map 0% reduce 0%
17/11/02 21:57:08 INFO mapreduce.Job: map 100% reduce 0%
17/11/02 21:57:08 INFO mapreduce.Job: Job job_1509673169537_0007 completed successfully
17/11/02 21:57:08 INFO mapreduce.Job: Counters: 30
File System Counters
FILE: Number of bytes read=0
FILE: Number of bytes written=236676
FILE: Number of read operations=0
FILE: Number of large read operations=0
FILE: Number of write operations=0
HDFS: Number of bytes read=10066414
HDFS: Number of bytes written=56673
HDFS: Number of read operations=25
HDFS: Number of large read operations=0
HDFS: Number of write operations=6
Job Counters
Launched map tasks=2
Other local map tasks=2
Total time spent by all maps in occupied slots (ms)=6437
Total time spent by all reduces in occupied slots (ms)=0
Total time spent by all map tasks (ms)=6437
Total vcore-seconds taken by all map tasks=6437
Total megabyte-seconds taken by all map tasks=6591488
Map-Reduce Framework
Map input records=10
Map output records=0
Input split bytes=810
Spilled Records=0
Failed Shuffles=0
Merged Map outputs=0
GC time elapsed (ms)=122
CPU time spent (ms)=3350
Physical memory (bytes) snapshot=555610112
Virtual memory (bytes) snapshot=5528924160
Total committed heap usage (bytes)=670040064
File Input Format Counters
Bytes Read=0
File Output Format Counters
Bytes Written=0
end Jod 21:57:08.901
Job finishes in 12130 miliseconds.
Total K-means Execution Time: 12130
Disclaimer: sequenceiq/hadoop-docker is a base image.
Content type
Image
Digest
Size
15.7 GB
Last updated
almost 9 years ago
docker pull lee212/harp-daal