Sign inSign up

pratikroy311/bigdata-env

By pratikroy311

•Updated almost 2 years ago

Big Data Environment Image with Hadoop,Apache Spark,Hive, Sqoop,Apache Airflow,Python with PySpark

Image
Machine learning & AI
Data science
0

221

pratikroy311/bigdata-env repository overview

This Docker image contains a complete big data environment with the following tools:

  • Hadoop
  • Apache Spark
  • Hive
  • Sqoop
  • Apache Airflow
  • Python with PySpark, Pandas, and NumPy

⁠Versions Used

  • Base Image: Ubuntu 24.04 LTS
  • Apache Spark: Version 3.5.1
  • Apache Hadoop: Version 3.3.6
  • Python: Version 3.12
  • Jupyter: Notebook and Lab
  • Apache Airflow: Version compatible with Python 3.12

⁠Usage

⁠Pulling the Image

docker pull pratikroy311/bigdata-env:latest

⁠Running the Container

docker run -it --rm -p 8080:8080 -p 8081:8081 -p 4040:4040 -p 10000:10000 -p 50070:50070 -p 50075:50075 -p 9000:9000 9001:9001 -p 8888:8888 pratikroy311/bigdata-env:latest

This will start the container and expose the necessary ports for accessing the web UIs of Spark, Hadoop, Jupyter Notebook, etc.

⁠Ports
• 8080/8081: Spark Master and Worker UI
• 4040: Spark application's web UI
• 10000: Spark Thrift Server
• 50070/50075: Hadoop's NameNode and DataNode UI
• 9000/9001: HDFS and MinIO
• 8888: Jupyter Notebook

Notes Ensure Docker is running on your system before using the commands above.

Tag summary

Content type

Image

Digest

sha256:10148ac80…

Size

3 GB

Last updated

almost 2 years ago

docker pull pratikroy311/bigdata-env