The images contains: - Hadoop 3.3.3 - Spark 3.3.3 - Jupyter Notebook
235
#Apache #Hadoop #Spark
Complete installation of Hadoop 3.3.3 with openjdk11 in pseudo-distribted mode and Spark3.3.0.
spark-hadoop directory# docker build -t spark-hadoop .)# docker container run -it --name "sparkc" -h sparkc spark-hadoop)Alternatively you can download this image from my Docker Hub account without needing to build it by the running:
docker run -it --rm --name sparkc -h sparkc -v `pwd`/data:/data asami76/spark-hadoop:latest
# jps inside the container. the output should be the hadoop services running like below:<pid> NodeManager
<pid> SecondaryNameNode
<pid> Jps
<pid> NameNode
<pid> ResourceManager
<pid> DataNode
# ./spark/bin/pyspark to launch PySpark Shell (Spark is installed in /usr/local/spark folder)http://sparkc:4040 to launch the Spark UI page (before that you have to expose the 4040 port or run the hosterr service below).Dockerfile to EXPOSE the required ports, or the container's ip address to use the hadoop services from outside the container/etc/hosts file
in order to be able to access the container by name from the docker hostdocker run -d -v /var/run/docker.sock:/tmp/docker.sock -v /etc/hosts:/tmp/hosts asami76/docker-hosterhttp://sparkc:9870to be able to use JupyterLab to connect to the Spark standalone cluster in the container rather than using the pyspark shell run the following:
jupyter-lab --no-browser --allow-root --ip 0.0.0.0 /data/notebooks/
Then copy the provided link to open Jupyter Notebook from the Docker host's browser
Content type
Image
Digest
sha256:d48a7a296…
Size
3.3 GB
Last updated
about 4 years ago
docker pull asami76/spark-hadoop