Sign inSign up

zilutian/data-analytics

By zilutian

•Updated over 7 years ago

Image
0

1.4K

zilutian/data-analytics repository overview

⁠TL;DR

Currently support amd64(x86-64) and aarch64 (ARMv8). Default tag "latest" will detect the platform and pull the right image automatically.

docker pull zilutian/hadoop
docker pull zilutian/data-analytics
docker network create hadoop-net
docker run -d --net hadoop-net --name master --hostname master zilutian/data-analytics master
docker run -d --net hadoop-net --name slave01 --hostname slave01 zilutian/hadoop slave
docker run -d --net hadoop-net --name slave02 --hostname slave02 zilutian/hadoop slave
docker exec master benchmark

Expected output:

Mahout: seqwiki
MAHOUT_LOCAL is not set; adding HADOOP_CONF_DIR to classpath.
Running on hadoop, using /opt/hadoop-2.9.1/bin/hadoop and HADOOP_CONF_DIR=/opt/hadoop-2.9.1/etc/hadoop
MAHOUT-JOB: /opt/mahout-0.13.0/mahout-examples-0.13.0-job.jar
19/05/23 12:45:52 INFO WikipediaToSequenceFile: Input: /user/root/wiki Out: /user/root/wiki-seq Categories: /root/categories All Files: false
19/05/23 12:45:53 INFO RMProxy: Connecting to ResourceManager at master/172.18.0.2:8032
19/05/23 12:45:53 WARN JobResourceUploader: Hadoop command-line option parsing not performed. Implement the Tool interface and execute your application with ToolRunner to remedy this.
19/05/23 12:45:53 INFO FileInputFormat: Total input files to process : 1
19/05/23 12:45:53 INFO JobSubmitter: number of splits:5
19/05/23 12:45:53 INFO deprecation: yarn.resourcemanager.system-metrics-publisher.enabled is deprecated. Instead, use yarn.system-metrics-publisher.enabled
19/05/23 12:45:53 INFO JobSubmitter: Submitting tokens for job: job_1558614614151_0001
19/05/23 12:45:54 INFO YarnClientImpl: Submitted application application_1558614614151_0001
19/05/23 12:45:54 INFO Job: The url to track the job: http://master:8088/proxy/application_1558614614151_0001/
19/05/23 12:45:54 INFO Job: Running job: job_1558614614151_0001
19/05/23 12:46:01 INFO Job: Job job_1558614614151_0001 running in uber mode : false
19/05/23 12:46:01 INFO Job:  map 0% reduce 0%
19/05/23 12:46:18 INFO Job:  map 3% reduce 0%
19/05/23 12:46:22 INFO Job:  map 5% reduce 0%
19/05/23 12:46:23 INFO Job:  map 9% reduce 0%
...
19/05/23 12:47:32 INFO Job:  map 100% reduce 100%
19/05/23 12:47:33 INFO Job: Job job_1558614614151_0001 completed successfully
19/05/23 12:47:33 INFO Job: Counters: 49
	File System Counters
		FILE: Number of bytes read=465780686
		FILE: Number of bytes written=848490336
		FILE: Number of read operations=0
		FILE: Number of large read operations=0
		FILE: Number of write operations=0
		HDFS: Number of bytes read=646515909
		HDFS: Number of bytes written=381691842
		HDFS: Number of read operations=18
		HDFS: Number of large read operations=0
		HDFS: Number of write operations=2
	Job Counters 
		Launched map tasks=5
		Launched reduce tasks=1
		Data-local map tasks=5
		Total time spent by all maps in occupied slots (ms)=813960
		Total time spent by all reduces in occupied slots (ms)=34510
		Total time spent by all map tasks (ms)=406980
		Total time spent by all reduce tasks (ms)=17255
		Total vcore-milliseconds taken by all map tasks=406980
		Total vcore-milliseconds taken by all reduce tasks=17255
		Total megabyte-milliseconds taken by all map tasks=833495040
		Total megabyte-milliseconds taken by all reduce tasks=35338240
	Map-Reduce Framework
		Map input records=19821
		Map output records=7532
		Map output bytes=381483448
		Map output materialized bytes=381515657
		Input split bytes=490
		Combine input records=0
		Combine output records=0
		Reduce input groups=7532

Mahout: seq2sparse
...
⁠Detailed (Debug tips)

To begin with, remove the detach option "-d" when starting master and slave containers. After master container is up (Namenode is ready), when adding slave (datanode), you should see the following output acknowledging a datanode has been registered and added:

 * Starting OpenBSD Secure Shell server sshd
   ...done.
... INFO namenode.NameNode: STARTUP_MSG: 
/************************************************************
STARTUP_MSG: Starting NameNode
STARTUP_MSG:   host = master/192.168.96.2
STARTUP_MSG:   args = [-format, cc]
STARTUP_MSG:   version = 2.9.1
...
==> /opt/hadoop-2.9.1/logs/hadoop--namenode-master.log 
2019-05-23 12:43:44,940 INFO org.apache.hadoop.hdfs.StateChange: BLOCK* registerDatanode: from DatanodeRegistration(172.18.0.3:50010, datanodeUuid=05c5c613-1974-4f57-8d35-19586a64960a, infoPort=50075, infoSecurePort=0, ipcPort=50020, storageInfo=lv=-57;cid=CID-be41a62e-61f0-4960-8c18-370970130df6;nsid=2081920441;c=1558614608916) storage 05c5c613-1974-4f57-8d35-19586a64960a
2019-05-23 12:43:44,941 INFO org.apache.hadoop.net.NetworkTopology: Adding a new node: /default-rack/172.18.0.3:50010

You should see the state transitions

2019-05-23 12:46:02,050 INFO org.apache.hadoop.yarn.server.resourcemanager.rmcontainer.RMContainerImpl: container_1558614614151_0001_01_000002 Container Transitioned from ALLOCATED to ACQUIRED
2019-05-23 12:46:02,052 INFO org.apache.hadoop.yarn.server.resourcemanager.rmcontainer.RMContainerImpl: container_1558614614151_0001_01_000003 Container Transitioned from ALLOCATED to ACQUIRED
...
2019-05-23 12:46:02,717 INFO org.apache.hadoop.yarn.server.resourcemanager.rmcontainer.RMContainerImpl: container_1558614614151_0001_01_000005 Container Transitioned from ACQUIRED to RUNNING
2019-05-23 12:46:02,718 INFO org.apache.hadoop.yarn.server.resourcemanager.rmcontainer.RMContainerImpl: container_1558614614151_0001_01_000006 Container Transitioned from ACQUIRED to RUNNING

Please make sure you have enough disk space and watch out warnings as such

2019-05-23 12:49:48,629 INFO org.apache.hadoop.yarn.server.resourcemanager.rmnode.RMNodeImpl: Node slave01:43159 reported UNHEALTHY with details: 1/1 local-dirs usable space is below configured utilization percentage/no more usable space [ /tmp/hadoop-root/nm-local-dir : used space above threshold of 90.0% ] ; 1/1 log-dirs usable space is below configured utilization percentage/no more usable space [ /opt/hadoop-2.9.1/logs/userlogs : used space above threshold of 90.0% ] 

which can lead to following error:

2019-05-23 12:49:48,933 ERROR org.apache.hadoop.yarn.server.resourcemanager.ApplicationMasterService: Application attempt appattempt_1558614614151_0003_000002 doesn't exist in ApplicationMasterService cache.

and will result in benchmark hang without making forward progress

Tag summary

Content type

Image

Digest

Size

874.9 MB

Last updated

over 7 years ago

docker pull zilutian/data-analytics