Lightweight PySpark development environment with Python 3.9.18 and Apache Spark 3.5.2.
281
This Docker image provides a lightweight development environment with Python 3.9 and Apache Spark 3.5.2, supporting remote connection via SSH, ideal for data analysis and machine learning projects. It is designed for quickly launching Spark services locally while maintaining a clean environment and isolation from your system dependencies, making it perfect for development, testing, and learning purposes.
For questions or suggestions for improvement, please contact me on GitHub or DockerHub.
docker run -d -p 8822:22 --name spark-container lightcone0204/spark-minimal:latest
docker run -d -p 8822:22 -v /local/path:/app/data --name spark-container lightcone0204/spark-minimal:latest
ssh root@localhost -p 8822
# Password: spark
The image comes with pre-installed database drivers and connectors that are automatically loaded, allowing you to connect to various databases:
You can specify connection details directly in your code:
from pyspark.sql import SparkSession
jar_path = "/opt/spark/jars/db-drivers/mysql-connector.jar"
driver_name = "com.mysql.cj.jdbc.Driver"
# Create SparkSession
spark = SparkSession.builder \
.appName("Database Connection") \
.config("spark.jars", jar_path) \
.getOrCreate()
# Connect to MySQL
jdbc_url = "jdbc:mysql://your-db-server:3306/your_database" # If your are using local db, your db host should be docker0 ip 172.17.0.1, not localhost.
jdbc_properties = {
"user": "username",
"password": "password",
"driver": driver_name
}
# Read data
df = spark.read.jdbc(url=jdbc_url, table="your_table", properties=jdbc_properties)
The following drivers are pre-installed:
docker run -d -p 8822:22 \
-v /path/to/your/drivers:/opt/spark/jars/custom \
--name spark-container lightcone0204/spark-minimal:latest
docker exec -it spark-container pyspark
Save your script in the mounted directory, then execute:
docker exec -it spark-container spark-submit /app/data/your_script.py
/app/data/test.py)from pyspark.sql import SparkSession
# Create SparkSession
spark = SparkSession.builder \
.appName("SimpleExample") \
.getOrCreate()
# Create simple data
data = [("Spark", 2022), ("Python", 3.9), ("Analysis", 100)]
df = spark.createDataFrame(data, ["Name", "Version"])
# Display data
print("Data Preview:")
df.show()
# Stop SparkSession
spark.stop()
Mount Local Data Directory:
docker run -d -p 8822:22 -v /local/data/dir:/app/data --name spark-container lightcone0204/spark-minimal
Place Data Files in Mounted Directory: Copy CSV, JSON, or Parquet files to your local mounted directory
Write and Run Spark Code:
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("DataAnalysis").getOrCreate()
df = spark.read.csv("/app/data/your_file.csv", header=True, inferSchema=True)
df.printSchema()
df.show(5)
# Process data...
result = df.groupBy("category").count()
result.write.parquet("/app/data/results.parquet")
Configure the SSH connection details:
localhost (or your server IP if running on remote machine)8822 (or your custom SSH port)rootspark/usr/bin/pythonIn the path mappings screen, set up the synchronization between your local project and the container:
/path/to/your/local/project/app/data-v /local/path:/app/data, ensure your path mappings reflect thisClick Finish
docker run -d -p 8822:22 -e PYSPARK_DRIVER_MEMORY=2g --name spark-container lightcone0204/spark-minimal:latest
Mount Spark configuration directory:
docker run -d -p 8822:22 -v /local/spark/conf:/opt/spark/conf/custom --name spark-container lightcone0204/spark-minimal:latest
docker run -d -p 8822:22 -p 18080:18080 -v /local/history/dir:/spark-events --name spark-container lightcone0204/spark-minimal:latest
sparkReleased under Apache License 2.0
Content type
Image
Digest
sha256:7a0743792…
Size
641.5 MB
Last updated
over 1 year ago
docker pull lightcone0204/spark-minimal