Sign inSign up

fullstackml/oneboxml-spark-shell

By fullstackml

•Updated about 9 years ago

OneboxML Apache Spark Shell

Image
0

1.1K

fullstackml/oneboxml-spark-shell repository overview

⁠OneBoxML introduction

OneBoxML is all in one data scientists tool box. OneBoxML extends your data science strengths by the power of cloud computing. Simply speaking, with OneBoxML you can run your data science scripts (Python, R, Apache Spark) in your local machine and EC2 instances of any types (how about an instance with 2TB of RAM or an instance with 16GB NVIDIA GRID GPUs?) and simply synchronize data from your machine to the cloud instances and back.

⁠How it works?

⁠Dockerizing data science tools

OneBoxML uses Docker⁠ under the hood and provides set of dockerized data scientists tools:

  1. OneBoxML Python⁠ - dockerized Python with pre-installed machine learning libraries such as NumPy, SciPy, TensorFlow.
  2. [OneBoxML Rscript] (https://hub.docker.com/r/fullstackml/oneboxml-rscript/⁠) - R with pre-installed libraries.
  3. OneBoxML Spark Shell⁠ - dockerized standalone Apache Spark⁠. It is usefull for out-of-memory data wrangling and out-of-memory machine learning.

You don't need any docker skills or knowladge. OneBoxML Shell commands hide Docker from users. Examples:

oneboxml-python test/pytest.py                    # Run your script by Python container in local machine.
oneboxml-run-instance --instance-type g2.8xlarge  # Launch a GPU instance in EC2.
oneboxml-cloud-python test/pytest.py              # Run python script in EC2 instance.
oneboxml-terminate-instances all                  # Stop the instance.

The only requirement - you should have docker installed⁠ in your machine and set up maximum amount of memory that docker can use. We recomend to use 80% of your RAM memory.

⁠OneBoxML directory

An important concept in OneBoxML is OneBoxML directory which is a directory in your local machine. This directory is the only visiable directory from docker containers in your local machine as well as from remote EC2 instances. All your code and data has to be in this directory.

Use ONEBOXML_DIR variable to set up the directory:

export ONEBOXML_DIR=~/src/oneboxml

⁠Tutorial

⁠Setup

You should have Bash command line and installed Docker⁠.

Let's clone the project source code in your ~/src directory:

cd ~/src
git clone https://github.com/dmpetrov/oneboxml.git

Export OneBoxML commands:

cd oneboxml
source oneboxml.rc

Export OneBoxML directory. All your code and data has to be in the box directory. This directory will be used to sync you data to cloud machines and back - from cloud servers:

export ONEBOXML_DIR=~/src/oneboxml

⁠Run in local machine

Let's run a python script with numpy and scipy in your local machine. You don't need numpy or scipy libraries. The dockerized OneBoxML Python will be launched with pre-installed libraries.

oneboxml-python test/pytest.R test/pyinput.csv test/pyoutput.csv

The script reads test/pyinput.csv file and creates output file test/pyoutput.csv.

⁠Set up AWS credentials

The power of OneBoxML is running script in the EC2 cloud. To do that you have to set up AWS credentials. AWS access key and security access kay have to be set up in OneBoxML config file: ~/src/oneboxml/oneboxml.conf

AccessKeyID = AKIAJ123456789012346 SecretAccessKey = ASDFGHJKLasdfghjklASDFGHJKLasdfghjklASDF

You can find an instruction how to get the AWS keys here: http://docs.aws.amazon.com/general/latest/gr/aws-sec-cred-types.html⁠

⁠Run code in EC2 instance

First, let's launch a regular EC2 instance with 7.5Gb of memory for our experiments. Type m3.large. See EC2 instances types⁠.

oneboxml-run-instance --instance-type m3.large

Output:

New m3.large instance i-53c8d962 was selected as active
Waiting for a running status.
...............

Now one instance is running and active (See the first column in the next command output).

oneboxml-describe-instances

Output:

Active  Id          Type        State       Storage         IP public       IP private      
 ***    i-53c8d962  m3.large    running                     54.166.78.104   172.31.59.211

Many instances might be created but only one of them can be active. Multi-instance environment will be described later. Now we are working with one single instance. Show a list of runing instances:

Synchronize your local OneBoxML directory to the active instance:

oneboxml-sync-to-cloud

Run the script in cloud:

oneboxml-cloud-python test/pytest.py

Sync data from your remote directory to OneBoxML directory:

oneboxml-sync-to-local

Now the output file is synced from the EC2 instance to the local directory test/pyoutput.csv.

The EC2 instance could be stoped:

oneboxml-terminate-instances all

⁠Experiments

⁠Git semantic

OPEN QUESTION:

  1. can we avoid syncing data directly? Let's use only git and S3!

Show list of environments which is empty by default.

$ oneboxml env
$ oneboxml sync-status
No remote server set up
Nothing to sync

Run new instance:

$ oneboxml run-instance --instance-type g2.8xlarge --name gpu-server
$ oneboxml env  # Show the environment
* gpu-server

Run a script remoutly in the new instance:

$ oneboxml sync-to gpu-server   # Push all data from local host to the instance
$ oneboxml run --remote python test/pytest.py test/pyinput.csv output/pyoutput.csv
$ oneboxml sync-status
Unsynced remoute files:
    output/pyoutput.csv
$ oneboxml sync-from gpu-server # Change env to the local
$ cat output/pyoutput.csv         # Output the result.
$ oneboxml sync-status          # Everything is synced

Simple version of the same scenario:

$ oneboxml run --remote --full-sync python test/pytest.py test/pyinput.csv output/pyoutput.csv
$ cat output/pyoutput.csv

Verbose version of the simple script:

$ oneboxml run -v --remote --full-sync python test/pytest.py test/pyinput.csv output/pyoutput.csv
Sync to server:
    test/pytest.py
    test/rtest.R
    test/pyinput.csv
    test/Rinput.csv
Run command [gpu-server]: python test/pytest.py test/pyinput.csv output/pyoutput.csv
Sync from server:
    output/pyoutput.csv
$ cat test/pyoutput.csv

⁠Persist reproducible results

$ git checkout --track -b origin/py_experiment # create tracking branch (local and upstream)
Switched to a new branch 'py_experiment'
Branch origin/py_experiment set up to track local branch master.
Switched to a new branch 'origin/py_experiment'
$ vim test/pytest.py # Edit the source code
$ git commit -am 'New script version'
$ git push # push change to the 
$ oneboxml run --remote --repro --output output/pyoutput.csv --input test/pyinput.csv python test/pytest.py test/pyinput.csv output/pyoutput.csv
$ oneboxml sync-status
Unsynced remoute files:
    output/pyoutput.csv
    repro/622278b4ea910275ec572268c6492449c41fd5e4-output____pyoutput.csv
$ oneboxml sync-from gpu-server
$ cat output/pyoutput.csv    # Output the result.
$ oneboxml sync-status       # Everything is synced

The running command could be simplified. However, a script developer has to maintain this command semantics.

$ oneboxml run --remote --repro python test/pytest.py -i test/pyinput.csv -o output/pyoutput.csv

1

2

3

Tag summary

Content type

Image

Digest

Size

618.3 MB

Last updated

about 9 years ago

docker pull fullstackml/oneboxml-spark-shell