Sign inSign up

loum/spark-pseudo

By loum

•Updated over 2 years ago

Apache Spark (on Pseudo Distributed Hadoop)

Image
1

2.1K

loum/spark-pseudo repository overview

⁠About

See Apache Spark docs⁠ for more information.

This image allows you to run the Apache Spark compute engine. Spark itself uses YARN as the resource manager under Hadoop Pseudo Distributed mode⁠ which we leverage from the underlying Hadoop install.

See documentation on the underlying Hadoop Pseudo Distributed base Docker image⁠ for details on how Hadoop/YARN has been configured.

⁠Quick reference

⁠Interact with Apache Spark

⁠Start the container
docker run --rm -d\
 --publish 10000:10000\
 --publish 10002:10002\
 --publish 7077:7077\
 --publish 8080:8080\
 --publish 8032:8032\
 --publish 8088:8088\
 --publish 8042:8042\
 --publish 18080:18080\
 --env YARN_SITE__YARN_LOG_AGGREGATION_ENABLE=true\
 --name spark-pseudo loum/spark-pseudo:latest

More information on submitting applications to Spark can be found here⁠.

⁠Sample SparkPi Application

The sample SparkPi application⁠ can be launched with:

docker exec -ti spark-pseudo bash -c\
 "HADOOP_CONF_DIR=/opt/hadoop/etc/hadoop /opt/spark/bin/spark-submit\
 --class org.apache.spark.examples.SparkPi\
 --files /opt/spark/conf/metrics.properties.template\
 --master yarn\
 --deploy-mode cluster\
 --driver-memory 1g\
 --executor-memory 1g\
 --executor-cores 1\
 /opt/spark/examples/jars/spark-examples_2.*-3.*.jar"

Apart from verbose logging displayed on the console it may appear that not much has happened here. However, since the Spark application has been deployed in cluster mode⁠ you will need to dump the associated application ID's log to see meaningful output.

To get a list of Spark application logs (under YARN):

docker exec -ti spark-pseudo bash -c "/opt/hadoop/bin/yarn application -list -appStates ALL"

Then plug in an Application-Id into:

docker exec -ti spark-pseudo bash -c "/opt/hadoop/bin/yarn logs -log_files stdout -applicationId <Application-Id>"

To see something similar to the following:

====================================================================
LogType:stdout
LogLastModifiedTime:Sat Apr 11 21:49:03 +0000 2020
LogLength:33
LogContents:
Pi is roughly 3.1398156990784956

End of LogType:stdout
***********************************************************************
⁠pyspark
docker exec -ti spark-pseudo bash -c "/opt/spark/bin/pyspark"
⁠spark-shell
docker exec -ti spark-pseudo bash -c "/opt/spark/bin/spark-shell"
⁠Terminate the container
docker stop spark-pseudo

⁠Web Interfaces

The following web interfaces are available to view configurations and logs and to track YARN/Spark job submissions:

Only need the Spark standalone cluster?

Tag summary

Content type

Image

Digest

sha256:c50899a4e…

Size

1.9 GB

Last updated

over 2 years ago

docker pull loum/spark-pseudo