Sign inSign up

lee212/harp-daal

By lee212

•Updated almost 9 years ago

ML framework built on Harp, a Hadoop plug-in for collective-communication on Big Data and Intel DAAL

Image
0

477

lee212/harp-daal repository overview

Harp is a Hadoop plug-in, which rewrites the Hadoop MapReduce framework to achieve in-memory communication among nodes of a distributed environment. It has the following two features

Harp has MPI-like collective communication operations that are highly optimized for big data problems. Harp has efficient and innovative computation models for different machine learning proble

See Details: https://github.com/DSC-SPIDAL/harp/tree/master/harp-daal-app⁠

⁠Staring an interactive docker console

docker run -it lee212/harp-daal /etc/bootstrap.sh -bash

⁠Test Run of Kmeans Map-collective job

$HADOOP_HOME/bin/hadoop jar $HADOOP_HOME/harp-app-1.0-SNAPSHOT.jar edu.iu.kmeans.regroupallgather.KMeansLauncher 1000 10 100 5 2 2 10 /kmeans /tmp/kmeans

  Number of Map Tasks = 2
Generate data.
Generating data.....
Writing 100 vectors to a file
Directory: /tmp/kmeans created
Write file 4
Write file 6
Write file 0
Write file 7
Write file 1
Write file 2
Write file 5
Write file 3
Write file 9
Write file 8
Generate centroid data./kmeans/centroids/init_centroids
Wrote centroids data to file
Starting Job
Start Job 16:01:36.916
Job configure in 10 miliseconds.
Get map collective runner
17/10/12 16:01:36 INFO client.RMProxy: Connecting to ResourceManager at /0.0.0.0:8032
17/10/12 16:01:37 INFO input.FileInputFormat: Total input paths to process : 10
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: NUMBER OF FILES: 10
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: NUMBER OF MAPS: 2
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: Split on host: f2f5b1c9ba71
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: Split on host: f2f5b1c9ba71
17/10/12 16:01:37 INFO fileformat.MultiFileInputFormat: Total # of splits: 2
17/10/12 16:01:37 INFO mapreduce.JobSubmitter: number of splits:2
17/10/12 16:01:38 INFO mapreduce.JobSubmitter: Submitting tokens for job: job_1507838336882_0001
command: $JAVA_HOME/bin/java -Dlog4j.configuration=container-log4j.properties -Dyarn.app.container.log.dir=<LOG_DIR> -Dyarn.app.container.log.filesize=0 -Dhadoop.root.logger=INFO,CLA -Dhadoop.root.logfile=syslog  -Xmx1024m org.apache.hadoop.mapreduce.v2.app.MapCollectiveAppMaster 1><LOG_DIR>/stdout 2><LOG_DIR>/stderr
17/10/12 16:01:38 INFO impl.YarnClientImpl: Submitted application application_1507838336882_0001
17/10/12 16:01:38 INFO mapreduce.Job: The url to track the job: http://f2f5b1c9ba71:8088/proxy/application_1507838336882_0001/
17/10/12 16:01:38 INFO mapreduce.Job: Running job: job_1507838336882_0001
17/10/12 16:01:44 INFO mapreduce.Job: Job job_1507838336882_0001 running in uber mode : false
17/10/12 16:01:44 INFO mapreduce.Job:  map 0% reduce 0%
17/10/12 16:01:49 INFO mapreduce.Job:  map 100% reduce 0%
17/10/12 16:01:50 INFO mapreduce.Job: Job job_1507838336882_0001 completed successfully
17/10/12 16:01:50 INFO mapreduce.Job: Counters: 30
        File System Counters
                FILE: Number of bytes read=0
                FILE: Number of bytes written=231224
                FILE: Number of read operations=0
                FILE: Number of large read operations=0
                FILE: Number of write operations=0
                HDFS: Number of bytes read=818759
                HDFS: Number of bytes written=18161
                HDFS: Number of read operations=25
                HDFS: Number of large read operations=0
                HDFS: Number of write operations=6
        Job Counters
                Launched map tasks=2
                Other local map tasks=2
                Total time spent by all maps in occupied slots (ms)=6022
                Total time spent by all reduces in occupied slots (ms)=0
                Total time spent by all map tasks (ms)=6022
                Total vcore-seconds taken by all map tasks=6022
                Total megabyte-seconds taken by all map tasks=6166528
        Map-Reduce Framework
                Map input records=10
                Map output records=0
                Input split bytes=530
                Spilled Records=0
                Failed Shuffles=0
                Merged Map outputs=0
                GC time elapsed (ms)=110
                CPU time spent (ms)=1740
                Physical memory (bytes) snapshot=364224512
                Virtual memory (bytes) snapshot=4076158976
                Total committed heap usage (bytes)=514850816
        File Input Format Counters
                Bytes Read=0
        File Output Format Counters
                Bytes Written=0
end Jod 16:01:50.624
Job finishes in 13708 miliseconds.
Total K-means Execution Time: 13708

⁠Copying test result from HDFS

$HADOOP_HOME/bin/hdfs dfs -get /kmeans .

⁠Test Run of Kmeans Map-collective job using Harp-DAAL

source /opt/intel/bin/compilervars.sh intel64 source $HARP_GITHUB_REPO/harp-daal-app/__release__lnx/daal/bin/daalvars.sh intel64 export LIBJARS=${DAALROOT}/lib/daal.jar cd $HADOOP_HOME cd $HADOOP_PREFIX hadoop jar harp-daal-app-1.0-SNAPSHOT.jar edu.iu.daal_kmeans.regroupallgather.KMeansDaalLauncher -D mapred.child.java.opts=-Xmx800M -libjars $LIBJARS 5000 100 100 5 2 4 10 1024 /kmeans-P5000-C100-D100-F5-ITR10-N2 /tmp/kmeans true

Number of Map Tasks = 2
Generate data.
Generating data.....
Writing 500 vectors to a file
Directory: /tmp/kmeans created
Write file 2
Write file 0
Write file 7
Write file 4
Write file 6
Write file 3
Write file 5
Write file 1
Write file 8
Write file 9
File path : /tmp/kmeans and name :data_3
Full Path :/tmp/kmeans/data_3
Deleting file : /tmp/kmeans/data_3
File path : /tmp/kmeans and name :data_7
Full Path :/tmp/kmeans/data_7
Deleting file : /tmp/kmeans/data_7
File path : /tmp/kmeans and name :data_4
Full Path :/tmp/kmeans/data_4
Deleting file : /tmp/kmeans/data_4
File path : /tmp/kmeans and name :data_8
Full Path :/tmp/kmeans/data_8
Deleting file : /tmp/kmeans/data_8
File path : /tmp/kmeans and name :data_2
Full Path :/tmp/kmeans/data_2
Deleting file : /tmp/kmeans/data_2
File path : /tmp/kmeans and name :data_6
Full Path :/tmp/kmeans/data_6
Deleting file : /tmp/kmeans/data_6
File path : /tmp/kmeans and name :data_0
Full Path :/tmp/kmeans/data_0
Deleting file : /tmp/kmeans/data_0
File path : /tmp/kmeans and name :data_5
Full Path :/tmp/kmeans/data_5
Deleting file : /tmp/kmeans/data_5
File path : /tmp/kmeans and name :data_1
Full Path :/tmp/kmeans/data_1
Deleting file : /tmp/kmeans/data_1
File path : /tmp/kmeans and name :data_9
Full Path :/tmp/kmeans/data_9
Deleting file : /tmp/kmeans/data_9
Deleting Directory : /tmp/kmeans
Generate centroid data./kmeans-P5000-C100-D100-F5-ITR10-N2/centroids/init_centroids
Wrote centroids data to file
Starting Job
Start Job 21:56:56.771
Job configure in 10 miliseconds.
Get map collective runner
17/11/02 21:56:56 INFO client.RMProxy: Connecting to ResourceManager at /0.0.0.0:8032
17/11/02 21:56:57 INFO input.FileInputFormat: Total input paths to process : 10
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: NUMBER OF FILES: 10
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: NUMBER OF MAPS: 2
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: Split on host: 112dea2c5eba
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: Split on host: 112dea2c5eba
17/11/02 21:56:57 INFO fileformat.MultiFileInputFormat: Total # of splits: 2
17/11/02 21:56:57 INFO mapreduce.JobSubmitter: number of splits:2
17/11/02 21:56:57 INFO mapreduce.JobSubmitter: Submitting tokens for job: job_1509673169537_0007
command: $JAVA_HOME/bin/java -Dlog4j.configuration=container-log4j.properties -Dyarn.app.container.log.dir=<LOG_DIR> -Dyarn.app.container.log.filesize=0 -Dhadoop.root.logger=INFO,CLA -Dhadoop.root.logfile=syslog  -Xmx1024m org.apache.hadoop.mapreduce.v2.app.MapCollectiveAppMaster 1><LOG_DIR>/stdout 2><LOG_DIR>/stderr
17/11/02 21:56:57 INFO impl.YarnClientImpl: Submitted application application_1509673169537_0007
17/11/02 21:56:57 INFO mapreduce.Job: The url to track the job: http://112dea2c5eba:8088/proxy/application_1509673169537_0007/
17/11/02 21:56:57 INFO mapreduce.Job: Running job: job_1509673169537_0007
17/11/02 21:57:02 INFO mapreduce.Job: Job job_1509673169537_0007 running in uber mode : false
17/11/02 21:57:02 INFO mapreduce.Job:  map 0% reduce 0%
17/11/02 21:57:08 INFO mapreduce.Job:  map 100% reduce 0%
17/11/02 21:57:08 INFO mapreduce.Job: Job job_1509673169537_0007 completed successfully
17/11/02 21:57:08 INFO mapreduce.Job: Counters: 30
        File System Counters
                FILE: Number of bytes read=0
                FILE: Number of bytes written=236676
                FILE: Number of read operations=0
                FILE: Number of large read operations=0
                FILE: Number of write operations=0
                HDFS: Number of bytes read=10066414
                HDFS: Number of bytes written=56673
                HDFS: Number of read operations=25
                HDFS: Number of large read operations=0
                HDFS: Number of write operations=6
        Job Counters
                Launched map tasks=2
                Other local map tasks=2
                Total time spent by all maps in occupied slots (ms)=6437
                Total time spent by all reduces in occupied slots (ms)=0
                Total time spent by all map tasks (ms)=6437
                Total vcore-seconds taken by all map tasks=6437
                Total megabyte-seconds taken by all map tasks=6591488
        Map-Reduce Framework
                Map input records=10
                Map output records=0
                Input split bytes=810
                Spilled Records=0
                Failed Shuffles=0
                Merged Map outputs=0
                GC time elapsed (ms)=122
                CPU time spent (ms)=3350
                Physical memory (bytes) snapshot=555610112
                Virtual memory (bytes) snapshot=5528924160
                Total committed heap usage (bytes)=670040064
        File Input Format Counters
                Bytes Read=0
        File Output Format Counters
                Bytes Written=0
end Jod 21:57:08.901
Job finishes in 12130 miliseconds.
Total K-means Execution Time: 12130

Disclaimer: sequenceiq/hadoop-docker is a base image.

Tag summary

Content type

Image

Digest

Size

15.7 GB

Last updated

almost 9 years ago

docker pull lee212/harp-daal