A Spark image for use with OpenShift
9.7K
Following this example, you will create a functional Apache Spark cluster using OpenShift and Docker.
You will setup a Spark master service and a set of Spark workers using Spark's standalone mode.
For the impatient expert, jump straight to the tl;dr section.
The Docker images are heavily based on https://github.com/mattf/docker-spark. And are curated in https://github.com/nealef/application-images/tree/master/spark
The Spark UI Proxy is taken from https://github.com/nealef/spark-ui-proxy which was almost entirely derived from https://github.com/aseigneurin/spark-ui-proxy.
The PySpark examples are taken from http://stackoverflow.com/questions/4114167/checking-if-a-number-is-a-prime-number-in-python/27946768#27946768
This example assumes
$ oc create -f examples/spark/namespace-spark-cluster.yaml
Now list all namespaces:
$ oc get namespaces
NAME LABELS STATUS
default <none> Active
spark-cluster name=spark-cluster Active
To configure OpenShift to work with our namespace, we will create a new context using our current context as a base:
$ CURRENT_CONTEXT=$(oc config view -o jsonpath='{.current-context}')
$ USER_NAME=$(oc config view -o jsonpath='{.contexts[?(@.name == "'"${CURRENT_CONTEXT}"'")].context.user}')
$ CLUSTER_NAME=$(oc config view -o jsonpath='{.contexts[?(@.name == "'"${CURRENT_CONTEXT}"'")].context.cluster}')
$ oc config set-context spark --namespace=spark-cluster --cluster=${CLUSTER_NAME} --user=${USER_NAME}
$ oc config use-context spark
The Master service is the master service for a Spark cluster.
Use the examples/spark/spark-master-controller.yaml file to create a
replication controller running the Spark Master service.
$ oc create -f examples/spark/spark-master-controller.yaml
replicationcontroller "spark-master-controller" created
Then, use the examples/spark/spark-master-service.yaml file to create a logical service endpoint that Spark workers can use to access the Master pod:
$ oc create -f examples/spark/spark-master-service.yaml
service "spark-master" created
$ oc get pods
NAME READY STATUS RESTARTS AGE
spark-master-controller-5u0q5 1/1 Running 0 8m
Check logs to see the status of the master. (Use the pod retrieved from the previous output.)
$ oc logs spark-master-controller-5u0q5
17/04/18 20:37:32 INFO Master: Registered signal handlers for [TERM, HUP, INT]
17/04/18 20:37:33 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
17/04/18 20:37:33 INFO SecurityManager: Changing view acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: Changing modify acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: SecurityManager: authentication disabled; ui acls disabled; users with view permissions: Set(unknown, dbus); users with modify permissions: Set(unknown, dbus)
17/04/18 20:37:34 INFO Slf4jLogger: Slf4jLogger started
17/04/18 20:37:34 INFO Remoting: Starting remoting
17/04/18 20:37:34 INFO Utils: Successfully started service 'sparkMaster' on port 7077.
17/04/18 20:37:34 INFO Remoting: Remoting started; listening on addresses :[akka.tcp://sparkMaster@spark-master:7077]
17/04/18 20:37:34 INFO Master: Starting Spark master at spark://spark-master:7077
bash-4.2# oc logs spark-master-controller-s84zm | head -30
17/04/18 20:37:32 INFO Master: Registered signal handlers for [TERM, HUP, INT]
17/04/18 20:37:33 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
17/04/18 20:37:33 INFO SecurityManager: Changing view acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: Changing modify acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: SecurityManager: authentication disabled; ui acls disabled; users with view permissions: Set(unknown, dbus); users with modify permissions: Set(unknown, dbus)
17/04/18 20:37:34 INFO Slf4jLogger: Slf4jLogger started
17/04/18 20:37:34 INFO Remoting: Starting remoting
17/04/18 20:37:34 INFO Utils: Successfully started service 'sparkMaster' on port 7077.
17/04/18 20:37:34 INFO Remoting: Remoting started; listening on addresses :[akka.tcp://sparkMaster@spark-master:7077]
17/04/18 20:37:34 INFO Master: Starting Spark master at spark://spark-master:7077
17/04/18 20:37:34 INFO Master: Running Spark version 1.5.2
17/04/18 20:37:35 INFO Utils: Successfully started service 'MasterUI' on port 8080.
17/04/18 20:37:35 INFO MasterWebUI: Started MasterWebUI at http://172.17.0.2:8080
17/04/18 20:37:35 INFO Utils: Successfully started service on port 6066.
17/04/18 20:37:35 INFO StandaloneRestServer: Started REST server for submitting applications on port 6066
17/04/18 20:37:35 INFO Master: I have been elected leader! New state: ALIVE
Once the master is started, we'll want to check the Spark WebUI. In order to access the Spark WebUI, we will deploy a specialized proxy. This proxy is necessary to access worker logs from the Spark UI.
Deploy the proxy controller with examples/spark/spark-ui-proxy-controller.yaml:
$ oc create -f examples/spark/spark-ui-proxy-controller.yaml
replicationcontroller "spark-ui-proxy-controller" created
We'll also need a corresponding Loadbalanced service for our Spark Proxy examples/spark/spark-ui-proxy-service.yaml:
$ oc create -f examples/spark/spark-ui-proxy-service.yaml
service "spark-ui-proxy" created
After creating the service, you should eventually get a loadbalanced endpoint:
$ oc get svc spark-ui-proxy -o wide
NAME CLUSTER-IP EXTERNAL-IP PORT(S) AGE SELECTOR
spark-ui-proxy 172.30.105.176 172.46.224.247,172.46.224.247 8080/TCP 15m component=spark-ui-proxy
You can then create a route to expose the UI to the outside world.
If your OpenShift cluster is not equipped with a Loadbalancer integration
The Spark workers do the heavy lifting in a Spark cluster. They provide execution resources and data cache capabilities for your program.
The Spark workers need the Master service to be running.
Use the examples/spark/spark-worker-controller.yaml file to create a
replication controller that manages the worker pods.
$ oc create -f examples/spark/spark-worker-controller.yaml
replicationcontroller "spark-worker-controller" created
If you launched the Spark WebUI, your workers should just appear in the UI when they're ready. (It may take a little bit to pull the images and launch the pods.) You can also interrogate the status in the following way:
$ oc get pods
NAME READY STATUS RESTARTS AGE
spark-master-controller-5u0q5 1/1 Running 0 25m
spark-worker-controller-e8otp 1/1 Running 0 6m
spark-worker-controller-fiivl 1/1 Running 0 6m
spark-worker-controller-ytc7o 1/1 Running 0 6m
$ oc logs spark-master-controller-5u0q5
[...]
17/04/18 20:39:31 INFO Master: Registering worker 172.17.0.3:35778 with 2 cores, 6.8 GB RAM
17/04/18 20:39:33 INFO Master: Registering worker 172.17.0.5:33797 with 2 cores, 6.8 GB RAM
The Zeppelin UI pod can be used to launch jobs into the Spark cluster either via a web notebook frontend or the traditional Spark command line. See Zeppelin and Spark architecture for more details.
Deploy Zeppelin:
$ oc create -f examples/spark/zeppelin-controller.yaml
replicationcontroller "zeppelin-controller" created
And the corresponding service:
$ oc create -f examples/spark/zeppelin-service.yaml
service "zeppelin" created
Zeppelin needs the spark-master service to be running.
$ oc get pods -l component=zeppelin
NAME READY STATUS RESTARTS AGE
zeppelin-controller-ja09s 1/1 Running 0 53s
Now you have two choices, depending on your predilections. You can do something graphical with the Spark cluster, or you can stay in the CLI.
For both choices, we will be working with this Python snippet:
from math import sqrt; from itertools import count, islice
def isprime(n):
return n > 1 and all(n%i for i in islice(count(2), int(sqrt(n)-1)))
nums = sc.parallelize(xrange(10000000))
print nums.filter(isprime).count()
Simply copy and paste the python snippet into pyspark from within the zeppelin pod:
$ oc exec zeppelin-controller-ja09s -it pyspark
Python 2.7.9 (default, Mar 1 2015, 12:57:24)
[GCC 4.9.2] on linux2
Type "help", "copyright", "credits" or "license" for more information.
Welcome to
____ __
/ __/__ ___ _____/ /__
_\ \/ _ \/ _ `/ __/ '_/
/__ / .__/\_,_/_/ /_/\_\ version 1.5.1
/_/
Using Python version 2.7.9 (default, Mar 1 2015 12:57:24)
SparkContext available as sc, HiveContext available as sqlContext.
>>> from math import sqrt; from itertools import count, islice
>>>
>>> def isprime(n):
... return n > 1 and all(n%i for i in islice(count(2), int(sqrt(n)-1)))
...
>>> nums = sc.parallelize(xrange(10000000))
>>> print nums.filter(isprime).count()
664579
Congratulations, you now know how many prime numbers there are within the first 10 million numbers!
Creating the Zeppelin service should have yielded you a Loadbalancer endpoint:
$ oc get svc zeppelin -o wide
NAME CLUSTER-IP EXTERNAL-IP PORT(S) AGE SELECTOR
zeppelin 172.30.228.174 172.46.134.46,172.46.134.46 80/TCP 1m component=zeppelin
Create a route to expose the Zeppelin UI to the outside world.
Once you've loaded up the Zeppelin UI, create a "New Notebook". In there we will paste our python snippet, but we need to add a %pyspark hint for Zeppelin to understand it:
%pyspark
from math import sqrt; from itertools import count, islice
def isprime(n):
return n > 1 and all(n%i for i in islice(count(2), int(sqrt(n)-1)))
nums = sc.parallelize(xrange(10000000))
print nums.filter(isprime).count()
After pasting in our code, press shift+enter or click the play icon to the right of our snippet. The Spark job will run and once again we'll have our result!
You now have services and replication controllers for the Spark master, Spark workers and Spark driver. You can take this example to the next step and start using the Apache Spark cluster you just created, see Spark documentation for more information.
oc create -f examples/spark
After it's setup:
oc get pods # Make sure everything is running
oc get svc -o wide # Get the Loadbalancer endpoints for spark-ui-proxy and zeppelin
At which point the Master UI and Zeppelin will be available at the URLs under the EXTERNAL-IP field.
You can also interact with the Spark cluster using the traditional spark-shell / spark-subsubmit / pyspark commands by using oc exec against the zeppelin-controller pod.
spark-submit on the zeppelin pod, you will need to provide a way for your clients to get to the examples/spark/spark-master-service.yaml. See Services for more information.The Zeppelin pod is large, so it may take a while to pull depending on your network. The size of the Zeppelin pod is something we're working on, see issue #17231.
Zeppelin may take some time (about a minute) on this pipeline the first time you run it. It seems to take considerable time to load.
Content type
Image
Digest
Size
1.2 GB
Last updated
over 9 years ago
docker pull sinenomine/spark