Data Engineering Docker Image with python 3.10.6 and spark 3.3. Azure Data Factory Packages.
174
This is an optimized docker image for data engineerings task. For working with Pyspark and Azure Data Factory Python SDK. It contains python 3.10.6 and spark 3.3. Jupyterlab is also included out of the box for working with jupyter notebooks.
Build Docker image:
docker pull johntorrestensor/data_engineering_adf:latest
Run interactive docker session, where "PWD" is your current working directory in the terminal:
sudo docker run -it --rm -p 8888:8888 -v "${PWD}":/home/ johntorrestensor/data_engineering_adf:latest
Then go to your VsCode and open your working directoy, and press Crtl + Shift + p and select:
Dev containers: Attach to runnig container...
A new VsCode window will open up, now you can start working with jupyter files, python files, debuggers, etc.
For jupyter notebooks install the "Jupyter" extension on the the VsCode window.
Content type
Image
Digest
sha256:3ed59b8d3…
Size
1.1 GB
Last updated
over 3 years ago
docker pull johntorrestensor/data_engineering_adf