Apache Spark with AWS Glue metastore and Python docker image
8.9K
Dockerfile available at: https://github.com/sebastiandaberdaku/spark-glue-python.git
Apache Spark with AWS Glue metastore and Python docker image
Base Image: Uses sebastiandaberdaku/spark-with-glue-builder:spark-v3.5.0 as the base image.
Python Version: Uses Python 3.10.12-slim-bookworm.
User Setup: Creates a system user 'spark' with a specified UID and GID for running Spark processes.
Java and Other Packages: Installs OpenJDK 17, Tini, Procps, and Gettext-base.
Environment Variables: Sets up environment variables for Java, Spark, and Hadoop.
Spark and Hadoop Installation: Copies Spark and Hadoop binaries from the builder image to the specified directories.
Permissions: Adjusts ownership of Spark and Hadoop directories to the 'spark' user.
Entrypoint Setup: Copies and configures the entrypoint and decommission scripts for Spark.
Python Dependencies: Installs PySpark and other dependencies listed in the requirements.txt file.
The current Docker image inherits a series of JARs from its builder image.
Here is a summary of the JAR files that are included in the Docker image (under /opt/spark/jars):
aws-glue-datacatalog-spark-client-3.5.0.jaraws-java-sdk-bundle-1.12.262.jarhadoop-aws-3.3.4.jarwildfly-openssl-1.0.7.Final.jarpostgresql-42.6.0.jarchecker-qual-3.31.0.jardelta-spark_2.12-3.0.0.jarantlr4-runtime-4.9.3.jardelta-storage-3.0.0.jardelta-storage-s3-dynamodb-3.0.0.jarHadoop native libraries are downloaded and installed in the /opt/hadoop directory.
Content type
Image
Digest
sha256:629908404…
Size
2.7 GB
Last updated
almost 3 years ago
docker pull sebastiandaberdaku/spark-glue-python:spark-v3.5.0-python-v3.10.12