Sign inSign up

andgineer/spark-aws

By andgineer

•Updated 6 months ago

Docker image with Apache Spark / Hadoop3, compatible with AWS services like S3

Image
0

2.4K

andgineer/spark-aws repository overview

Apache Spark standalone cluster with AWS services integration (S3, etc.) and comprehensive Data Science environment including PySpark, Pandas, and RDKit for cheminformatics.

⁠Dockerfile

github.com/andgineer/spark-aws-rdkit⁠

⁠Features

  • Apache Spark cluster in standalone mode
  • Full AWS services compatibility (S3, etc.)
  • Conda environment with Data Science tools:
  • Deployment options:
    • Local with docker compose
    • Cloud with AWS ECS

    ⁠Quick Start

Launch locally with docker compose:

./compose.sh up --build

This starts:

  • Spark Master
  • Two Spark Workers
  • Example job container (submit)

Access points:

  • Spark Web UI: http://localhost:8080⁠
  • Spark Driver: spark://localhost:7077
    • For PySpark use setMaster('spark://localhost:7077')

Note: On Linux, change docker.for.mac.localhost to localhost in .env file.

⁠Example PySpark Application

The submit container demonstrates how to:

  • Connect to the Spark cluster
  • Submit Spark jobs
  • Process data with PySpark

⁠AWS ECS Deployment

For production deployment on AWS Elastic Container Service (ECS):

  1. Navigate to ecs/ directory
  2. Configure your deployment in config.sh
  3. Run the automated deployment scripts

Detailed instructions available in ecs/README.md.

Tag summary

Content type

Image

Digest

sha256:ff66675cf…

Size

870.1 MB

Last updated

6 months ago

docker pull andgineer/spark-aws