Sign inSign up

joegagliardo/bigdata

By joegagliardo

•Updated over 8 years ago

Big Data (Hadoop 2.9, Pig, Hive, Spark, Cassandra, Mongo, HBase)

Image
3

1.2K

joegagliardo/bigdata repository overview

This is a Hadoop 2.9 build and will be the last one I do for this project. I will instead maintain the hadoop3 project going forward.

Notes: After updating Java 8, Cassandra is not working. I have to fix that soon. I also have to validate that all of my tests are working after major updates to a lot of these releases.

This is a docker image that has every tool I need to develop big data apps. It is not a lean image at all and is very big. It is meant as an easy way for a developer to have a test environment to try out features instead of having to install them on a base OS or build a virtual machine. So please, no comments about it's size. It's really more like a VM than a production docker.

I use the base Ubuntu image I built which has most of the basic linux tools I need as well as languages (MySQL, SQLite3, Java 8, Python 2 & 3, NodeJS and Scala). Then on top of that I add Hadoop, Pig, Hive, Spark, HBase, Cassandra and Mongo, Postgresql and CockroachDB.

This is a work in progress and I still need to get the Spark connectors for the HBase to work and also to make Hive work with Tea and Spark instead of just MapReduce. Ultimately though, my plan is to make Hive, Pig and Spark be able to connect to as many data sources as possible. Also, as I get them to work, I am including sample code and datasets in the /examples folder to test that it works and show others how to do it. So check back in from time to time to pull new updates as I get more and more features working. You can run /scripts/test-all.sh to see if everything works.

Once you have an image that you have either pulled or built yourself, you can create a container based on that image by using the following command:

There is a folder called /home but it is not the same as ~ which goes to the folder /root. So avoid using ~ and instead use /home when referring to the home folder. In order to map the /home folder to a folder on your host machine to allow the container to get to outside resources use the -v switch.

⁠docker run --name bigdata --hostname bigdata -p 50070:50070 -p 8088:8088 -p 10020:10020 -p 9042:9042 -p 10000:10000 -p 10001:10001 -p 10002:10002 -v "$HOME:/home" -it joegagliardo/bigdata /etc/bootstrap.sh -bash

if you don't want to map the /home to a folder on your host, just leave off the -v option as follows:

⁠docker run --name bigdata --hostname bigdata -p 50070:50070 -p 8088:8088 -p 10020:10020 -p 9042:9042 -p 10000:10000 -p 10001:10001 -p 10002:10002 -it joegagliardo/bigdata /etc/bootstrap.sh -bash

That will open the standard ports so you can use your host computer browser to get to localhost:50070 or localhost:8088 to browse the cluster. 9042 is to allow Eclipse to talk to the cluster. 10000 - 10002 are for HiveServer.

By default all data for the various clusters will be stored in /data and MySQL data will be in /data/mysql. It's easier to store the data inside the docker container file system, but if for any reason you wanted to move it, you could use -v switch to map the /data folder to a folder on your host machine. If you do so it will mean you have to recreate all the folder structure and reinitialize the data and reformat the cluster. I don't recommend it, but it is possible. Remember the Hive metastore is stored in there, so whenever you move these things around, or things get corrupt you may need to run the scripts /scripts/format-namenode.sh and/or /scripts/init-schema-mysql.sh. I put a lot of handy scripts in the /scripts folder, for starting and stopping various daemons. The only one missing is start-hbase.sh because it's already in the path.

The hidden docker file gets really big and doesn't reclaim space even when you delete images. At least on the Mac, I do a reset under Docker preferences whenever I find that file is getting too big. I also move it out of my library folder so I can keep tabs on it, and I exclude it from time machine backups. Keep in mind, when you do this you lose all containers and images, so you may want to export them before you reset and import them back afterwards.

After the first launch of the container you need to format the name node and initialize the hive metastore. I have set it up to do this automatically, whenever the folder /data/hdfs/name does not exists. To make this easy I built a few scripts in the /scripts folder. You can use init-schema-mysql.sh or init-schema-postgres.sh to use either SQL implementation you want for the Hive metastore. By default I am using MySQL

Run:

⁠/scripts/format-namenode.sh

⁠/scripts/init-schema-mysql.sh

After you exit a container you could restart it with (remember these are commands you run on the host not inside the running docker):

⁠docker start bigdata

You can attach to a running container with:

⁠docker attach bigdata

Or combine them into one line:

⁠docker start bigdata && docker attach bigdata

You can rename an image like this:

⁠docker tag d583c3ac45fd myname/newimagename:latest

You can delete just a particular container by name or id number

⁠docker rm -f bigdata

You can stop all containers:

⁠docker stop $(docker ps -a -q)

Delete all containers:

⁠docker rm $(docker ps -a -q) -f

If you want to connect another terminal to an already running container to have two windows to it you can type this:

⁠docker exec -i -t bigdata /bin/bash

Inside the container you will find normal pseudo distributed versions of Hadoop, Pig, Hive and Spark. MySql root user has a password of rootpassword and is setup so you can just type mysql to enter the shell. I made it so it only starts up mysql by default, so once you have a bash prompt you can start-dfs.sh and start-yarn.sh and if you need I made other handy scripts to start-cassandra.sh, start-mongo.sh and start-hbase-master.sh. I built a start-everything.sh to start everything.

I highly recommend before exiting from the docker to do a clean shutdown and run the following script, otherwise you may corrupt name node and Hive metastore:

⁠/scripts/stop-everything.sh

To make my life easier I made a few aliases in ~/.bash_profile on my local Mac, so I don't have to type out these long commands every time:

⁠alias new-bd='docker run --name bigdata --hostname bigdata -p 50070:50070 -p 8088:8088 -p 10020:10020 -p 9042:9042 -p 10000:10000 -p 10001:10001 -p 10002:10002 -v "$HOME:/home" -it joegagliardo/bigdata /etc/bootstrap.sh -bash'

⁠alias start-bd='docker start bigdata && docker attach bigdata'

⁠alias ssh-bd='docker exec -i -t bigdata /bin/bash'

Things that don't work yet: I cannot get Hive to run on Spark or Tez yet. It currently only runs in MR mode. The spark-hbase connector is not working yet.

I've encapsulated the necessary startup code to make python connect to spark into a module called initSpark.py that can be found in the /examples/spark folder. Look at the examples there, but basically to go interactive, I load python and then past the first few lines that initialize the sc, spark and conf objects. This detail is always left out of online examples, and people can never figure out how to make there code run through spark-submit.

If you want to build the image yourself, copy the Dockerfile, put it in a folder and type the following command:

⁠docker build -t joegagliardo/bigdata .

Tag summary

Content type

Image

Digest

Size

3.6 GB

Last updated

over 8 years ago

docker pull joegagliardo/bigdata