Sign inSign up

spiritsree/pyspark-shell

By spiritsree

•Updated over 4 years ago

An interactive pyspark shell in docker

Image
0

284

spiritsree/pyspark-shell repository overview

⁠Pyspark Shell

Docker Cloud Build Status Docker Pulls GitHub tag (latest SemVer)

An interactive pyspark-shell in docker

⁠Quick Start

You can run this by just running the following

$ docker run --rm -ti --name pyspark-shell \
             spiritsree/pyspark-shell:latest

This will give a spark shell to work on.

###########################################################
#                     Spark 2.4.5                         #
#                                                         #
#      Spark session can be accessed using "SPARK"        #
#                                                         #
###########################################################


>>>

⁠Load CSV

For loading csv to work on mount the directory containing CSV files as /data

$ docker run --rm -ti --name pyspark-shell \
             -v /locat/dir:/data \
             spiritsree/pyspark-shell:latest

For changing log level pass the env variable as follows

⁠Custom log levels

$ docker run --rm -ti --name pyspark-shell \
             -v /locat/dir:/data \
             -e LOG_LEVEL=debug \
             spiritsree/pyspark-shell:latest

Supported log levels are all, debug, error, fatal, trace, warn, info, off.

⁠Common Spark Dataframe Functions

Count

Count the number of rows

>>> DF.count()
30

Show

Displays data (will show only first 10 or 20 rows) if not specified.

>>> DF.show(5)
+---+---+---+---+
|_c1|_c2|_c3|_c4|
+---+---+---+---+
|  1| d1|  1|  d|
|  2| d2|  2|  d|
|  3| d3|  3|  d|
|  4| d4|  4|  d|
|  5| d5|  5|  d|
+---+---+---+---+
only showing top 5 rows

Columns

Show the header list.

>>> DF.columns
['_c1', '_c2', '_c3', '_c4']

PrintSchema

Shows the schema of data.

>>> DF.printSchema()
root
 |-- _c1: string (nullable = true)
 |-- _c2: string (nullable = true)
 |-- _c3: string (nullable = true)
 |-- _c4: string (nullable = true)

Select

Select based on expression.

>>> DF.select(DF._c1).show(2)
+---+
|_c1|
+---+
|  1|
|  2|
+---+
only showing top 2 rows

⁠Reference

Pyspark SQL Module⁠

Tag summary

Content type

Image

Digest

Size

505.5 MB

Last updated

over 6 years ago

docker pull spiritsree/pyspark-shell