Sign inSign up

kairen/yarn-spark

By kairen

•Updated about 11 years ago

Apache Spark 1.5.0 on YARN

Image
3

517

kairen/yarn-spark repository overview

⁠Apache Spark on YARN Quick Start

Pull the image from Docker Repository:

docker pull kairen/yarn-spark:1.5

Running the image:

docker run -d -p 8088:8088 -p 50070:50070 -h spark-master  \
-v  <your_dir>:/root/spark-run/  \
--name yarn-spark kairen/yarn-spark:1.5 -d

The <your_dir> is a share directory, you can put a dataset or run file in here.

⁠Running the examples

This image provides Java, Scala, R, Python example, You can perform the following operation.

Enter the container:

Using exec into the docker container:

docker exec -ti <container id> bash

Creating a test data:

$ touch test.txt
$ vi test.txt

# add the following
Absence of evidence
is not evidence 
of absence
Absence of evidence
is not evidence 
of absence
Absence of evidence
is not evidence 
of absence

Upload test data to HDFS:

hadoop fs -mkdir /test
hadoop fs -put test.txt /test
⁠Running Spark Job

Running the spark on standard:

cd /root/python
spark-submit wordcount.py hdfs://spark-master:9000/test/test.txt

Running the spark on YARN:

cd /root/python
spark-submit --master yarn-cluster wordcount.py hdfs://spark-master:9000/test/test.txt

Tag summary

Content type

Image

Digest

sha256:1d8dadd7b…

Size

1.5 GB

Last updated

about 11 years ago

docker pull kairen/yarn-spark:1.5