Docker image with Apache Spark / Hadoop3, compatible with AWS services like S3
2.4K
Apache Spark standalone cluster with AWS services integration (S3, etc.) and comprehensive Data Science environment including PySpark, Pandas, and RDKit for cheminformatics.
github.com/andgineer/spark-aws-rdkit
docker composeLaunch locally with docker compose:
./compose.sh up --build
This starts:
submit)Access points:
spark://localhost:7077
setMaster('spark://localhost:7077')Note: On Linux, change
docker.for.mac.localhosttolocalhostin.envfile.
The submit container demonstrates how to:
For production deployment on AWS Elastic Container Service (ECS):
ecs/ directoryconfig.shDetailed instructions available in ecs/README.md.
Content type
Image
Digest
sha256:ff66675cf…
Size
870.1 MB
Last updated
6 months ago
docker pull andgineer/spark-aws