Sign inSign up

sinenomine/spark

By sinenomine

•Updated over 9 years ago

A Spark image for use with OpenShift

Image
0

9.7K

sinenomine/spark repository overview

⁠Spark example

Following this example, you will create a functional Apache Spark⁠ cluster using OpenShift⁠ and Docker⁠.

You will setup a Spark master service and a set of Spark workers using Spark's standalone mode⁠.

For the impatient expert, jump straight to the tl;dr⁠ section.

⁠Sources

The Docker images are heavily based on https://github.com/mattf/docker-spark⁠. And are curated in https://github.com/nealef/application-images/tree/master/spark⁠

The Spark UI Proxy is taken from https://github.com/nealef/spark-ui-proxy⁠ which was almost entirely derived from https://github.com/aseigneurin/spark-ui-proxy⁠.

The PySpark examples are taken from http://stackoverflow.com/questions/4114167/checking-if-a-number-is-a-prime-number-in-python/27946768#27946768⁠

⁠Step Zero: Prerequisites

This example assumes

  • You have a OpenShift cluster installed and running.
  • That your OpenShift cluster is running SkyDNS⁠ or an equivalent integration.

⁠Step One: Create namespace

$ oc create -f examples/spark/namespace-spark-cluster.yaml

Now list all namespaces:

$ oc get namespaces
NAME          LABELS             STATUS
default       <none>             Active
spark-cluster name=spark-cluster Active

To configure OpenShift to work with our namespace, we will create a new context using our current context as a base:

$ CURRENT_CONTEXT=$(oc config view -o jsonpath='{.current-context}')
$ USER_NAME=$(oc config view -o jsonpath='{.contexts[?(@.name == "'"${CURRENT_CONTEXT}"'")].context.user}')
$ CLUSTER_NAME=$(oc config view -o jsonpath='{.contexts[?(@.name == "'"${CURRENT_CONTEXT}"'")].context.cluster}')
$ oc config set-context spark --namespace=spark-cluster --cluster=${CLUSTER_NAME} --user=${USER_NAME}
$ oc config use-context spark

⁠Step Two: Start your Master service

The Master service⁠ is the master service for a Spark cluster.

Use the examples/spark/spark-master-controller.yaml⁠ file to create a replication controller⁠ running the Spark Master service.

$ oc create -f examples/spark/spark-master-controller.yaml
replicationcontroller "spark-master-controller" created

Then, use the examples/spark/spark-master-service.yaml⁠ file to create a logical service endpoint that Spark workers can use to access the Master pod:

$ oc create -f examples/spark/spark-master-service.yaml
service "spark-master" created
⁠Check to see if Master is running and accessible
$ oc get pods
NAME                            READY     STATUS    RESTARTS   AGE
spark-master-controller-5u0q5   1/1       Running   0          8m

Check logs to see the status of the master. (Use the pod retrieved from the previous output.)

$ oc logs spark-master-controller-5u0q5
17/04/18 20:37:32 INFO Master: Registered signal handlers for [TERM, HUP, INT]
17/04/18 20:37:33 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
17/04/18 20:37:33 INFO SecurityManager: Changing view acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: Changing modify acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: SecurityManager: authentication disabled; ui acls disabled; users with view permissions: Set(unknown, dbus); users with modify permissions: Set(unknown, dbus)
17/04/18 20:37:34 INFO Slf4jLogger: Slf4jLogger started
17/04/18 20:37:34 INFO Remoting: Starting remoting
17/04/18 20:37:34 INFO Utils: Successfully started service 'sparkMaster' on port 7077.
17/04/18 20:37:34 INFO Remoting: Remoting started; listening on addresses :[akka.tcp://sparkMaster@spark-master:7077]
17/04/18 20:37:34 INFO Master: Starting Spark master at spark://spark-master:7077
bash-4.2# oc logs spark-master-controller-s84zm | head -30
17/04/18 20:37:32 INFO Master: Registered signal handlers for [TERM, HUP, INT]
17/04/18 20:37:33 WARN NativeCodeLoader: Unable to load native-hadoop library for your platform... using builtin-java classes where applicable
17/04/18 20:37:33 INFO SecurityManager: Changing view acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: Changing modify acls to: unknown,dbus
17/04/18 20:37:33 INFO SecurityManager: SecurityManager: authentication disabled; ui acls disabled; users with view permissions: Set(unknown, dbus); users with modify permissions: Set(unknown, dbus)
17/04/18 20:37:34 INFO Slf4jLogger: Slf4jLogger started
17/04/18 20:37:34 INFO Remoting: Starting remoting
17/04/18 20:37:34 INFO Utils: Successfully started service 'sparkMaster' on port 7077.
17/04/18 20:37:34 INFO Remoting: Remoting started; listening on addresses :[akka.tcp://sparkMaster@spark-master:7077]
17/04/18 20:37:34 INFO Master: Starting Spark master at spark://spark-master:7077
17/04/18 20:37:34 INFO Master: Running Spark version 1.5.2
17/04/18 20:37:35 INFO Utils: Successfully started service 'MasterUI' on port 8080.
17/04/18 20:37:35 INFO MasterWebUI: Started MasterWebUI at http://172.17.0.2:8080
17/04/18 20:37:35 INFO Utils: Successfully started service on port 6066.
17/04/18 20:37:35 INFO StandaloneRestServer: Started REST server for submitting applications on port 6066
17/04/18 20:37:35 INFO Master: I have been elected leader! New state: ALIVE

Once the master is started, we'll want to check the Spark WebUI. In order to access the Spark WebUI, we will deploy a specialized proxy⁠. This proxy is necessary to access worker logs from the Spark UI.

Deploy the proxy controller with examples/spark/spark-ui-proxy-controller.yaml⁠:

$ oc create -f examples/spark/spark-ui-proxy-controller.yaml
replicationcontroller "spark-ui-proxy-controller" created

We'll also need a corresponding Loadbalanced service for our Spark Proxy examples/spark/spark-ui-proxy-service.yaml⁠:

$ oc create -f examples/spark/spark-ui-proxy-service.yaml
service "spark-ui-proxy" created

After creating the service, you should eventually get a loadbalanced endpoint:

$ oc get svc spark-ui-proxy -o wide
NAME             CLUSTER-IP       EXTERNAL-IP                     PORT(S)    AGE       SELECTOR
spark-ui-proxy   172.30.105.176   172.46.224.247,172.46.224.247   8080/TCP   15m       component=spark-ui-proxy

You can then create a route to expose the UI to the outside world.

If your OpenShift cluster is not equipped with a Loadbalancer integration

⁠Step Three: Start your Spark workers

The Spark workers do the heavy lifting in a Spark cluster. They provide execution resources and data cache capabilities for your program.

The Spark workers need the Master service to be running.

Use the examples/spark/spark-worker-controller.yaml⁠ file to create a replication controller⁠ that manages the worker pods.

$ oc create -f examples/spark/spark-worker-controller.yaml
replicationcontroller "spark-worker-controller" created
⁠Check to see if the workers are running

If you launched the Spark WebUI, your workers should just appear in the UI when they're ready. (It may take a little bit to pull the images and launch the pods.) You can also interrogate the status in the following way:

$ oc get pods
NAME                            READY     STATUS    RESTARTS   AGE
spark-master-controller-5u0q5   1/1       Running   0          25m
spark-worker-controller-e8otp   1/1       Running   0          6m
spark-worker-controller-fiivl   1/1       Running   0          6m
spark-worker-controller-ytc7o   1/1       Running   0          6m

$ oc logs spark-master-controller-5u0q5
[...]
17/04/18 20:39:31 INFO Master: Registering worker 172.17.0.3:35778 with 2 cores, 6.8 GB RAM
17/04/18 20:39:33 INFO Master: Registering worker 172.17.0.5:33797 with 2 cores, 6.8 GB RAM

⁠Step Four: Start the Zeppelin UI to launch jobs on your Spark cluster

The Zeppelin UI pod can be used to launch jobs into the Spark cluster either via a web notebook frontend or the traditional Spark command line. See Zeppelin⁠ and Spark architecture⁠ for more details.

Deploy Zeppelin:

$ oc create -f examples/spark/zeppelin-controller.yaml
replicationcontroller "zeppelin-controller" created

And the corresponding service:

$ oc create -f examples/spark/zeppelin-service.yaml
service "zeppelin" created

Zeppelin needs the spark-master service to be running.

⁠Check to see if Zeppelin is running
$ oc get pods -l component=zeppelin
NAME                        READY     STATUS    RESTARTS   AGE
zeppelin-controller-ja09s   1/1       Running   0          53s

⁠Step Five: Do something with the cluster

Now you have two choices, depending on your predilections. You can do something graphical with the Spark cluster, or you can stay in the CLI.

For both choices, we will be working with this Python snippet:

from math import sqrt; from itertools import count, islice

def isprime(n):
    return n > 1 and all(n%i for i in islice(count(2), int(sqrt(n)-1)))

nums = sc.parallelize(xrange(10000000))
print nums.filter(isprime).count()
⁠Do something fast with pyspark!

Simply copy and paste the python snippet into pyspark from within the zeppelin pod:

$ oc exec zeppelin-controller-ja09s -it pyspark
Python 2.7.9 (default, Mar  1 2015, 12:57:24)
[GCC 4.9.2] on linux2
Type "help", "copyright", "credits" or "license" for more information.
Welcome to
      ____              __
     / __/__  ___ _____/ /__
    _\ \/ _ \/ _ `/ __/  '_/
   /__ / .__/\_,_/_/ /_/\_\   version 1.5.1
      /_/

Using Python version 2.7.9 (default, Mar  1 2015 12:57:24)
SparkContext available as sc, HiveContext available as sqlContext.
>>> from math import sqrt; from itertools import count, islice
>>>
>>> def isprime(n):
...     return n > 1 and all(n%i for i in islice(count(2), int(sqrt(n)-1)))
...
>>> nums = sc.parallelize(xrange(10000000))

>>> print nums.filter(isprime).count()
664579

Congratulations, you now know how many prime numbers there are within the first 10 million numbers!

⁠Do something graphical and shiny!

Creating the Zeppelin service should have yielded you a Loadbalancer endpoint:

$ oc get svc zeppelin -o wide
 NAME       CLUSTER-IP       EXTERNAL-IP                   PORT(S)   AGE       SELECTOR
zeppelin   172.30.228.174   172.46.134.46,172.46.134.46   80/TCP    1m        component=zeppelin

Create a route to expose the Zeppelin UI to the outside world.

Once you've loaded up the Zeppelin UI, create a "New Notebook". In there we will paste our python snippet, but we need to add a %pyspark hint for Zeppelin to understand it:

%pyspark
from math import sqrt; from itertools import count, islice

def isprime(n):
    return n > 1 and all(n%i for i in islice(count(2), int(sqrt(n)-1)))

nums = sc.parallelize(xrange(10000000))
print nums.filter(isprime).count()

After pasting in our code, press shift+enter or click the play icon to the right of our snippet. The Spark job will run and once again we'll have our result!

⁠Result

You now have services and replication controllers for the Spark master, Spark workers and Spark driver. You can take this example to the next step and start using the Apache Spark cluster you just created, see Spark documentation⁠ for more information.

⁠tl;dr

oc create -f examples/spark

After it's setup:

oc get pods # Make sure everything is running
oc get svc -o wide # Get the Loadbalancer endpoints for spark-ui-proxy and zeppelin

At which point the Master UI and Zeppelin will be available at the URLs under the EXTERNAL-IP field.

You can also interact with the Spark cluster using the traditional spark-shell / spark-subsubmit / pyspark commands by using oc exec against the zeppelin-controller pod.

⁠Known Issues With Spark

  • This provides a Spark configuration that is restricted to the cluster network, meaning the Spark master is only available as a cluster service. If you need to submit jobs using external client other than Zeppelin or spark-submit on the zeppelin pod, you will need to provide a way for your clients to get to the examples/spark/spark-master-service.yaml⁠. See Services⁠ for more information.

⁠Known Issues With Zeppelin

  • The Zeppelin pod is large, so it may take a while to pull depending on your network. The size of the Zeppelin pod is something we're working on, see issue #17231.

  • Zeppelin may take some time (about a minute) on this pipeline the first time you run it. It seems to take considerable time to load.

Tag summary

Content type

Image

Digest

Size

1.2 GB

Last updated

over 9 years ago

docker pull sinenomine/spark