Sign inSign up

ibmcom/iop-hadoop

By ibmcom

•Updated over 9 years ago

IBM Open Platform with Apache Hadoop

Artifact
Image
12

100K+

ibmcom/iop-hadoop repository overview

⁠IBM Open Platform with Apache Hadoop for non-production environments, v4.2 Docker image README

Welcome to the IBM BigInsights® Quick Start Edition Docker image README for the non-production multinode environment. The purpose of the BigInsights Quick Start Edition with Kitematic is to enable developers and data scientists to begin experimenting with IBM BigInsights in a cloud or virtual environment.

This version of the Quick Start Edition docker image can run with multiple nodes. If you process large volumes of data, handling this data in a reasonable amount of time requires a distributed cluster that works in parallel. With the multinode Quick Start docker you can simulate a distributed cluster environment by adding multiple nodes to the master node.

For more information about the BigInsights Quick Start Edition and the features it includes, see the IBM BigInsights documentation.

⁠What you can install

You can use these instructions to install either of the following software products on either a single-node or multinode configuration:

⁠IBM Open Platform with Apache Spark and Apache Hadoop (IOP) non-production

This installation takes about 25 to 35 minutes depending on the number of nodes.

The software branded in this Docker image is capable of running with a single node or multiple nodes. If you are processing large volumes of data, handling this data in a reasonable amount of time requires a distributed cluster that works in parallel. You can simulate this distributed cluster environment with this docker image by adding multiple nodes to the master node.

You can set up virtual nodes if you do not want to use your own hardware.

⁠System requirements

Before you download, ensure that your system meets the minimum requirements:

⁠Operating system

The following host operating system: Red Hat Enterprise Linux (RHEL) 7.x - 64-bit (Docker 1.10)

⁠Hardware requirements
  • A minimum of 5 nodes with 4 processor cores. To get the best performance, it is recommended that your system have a minimum of 8 processor cores. The product will work with fewer processors, but you might experience a significant difference in performance.
  • At least 16 GB RAM for IBM Open Platform with Apache Spark and Apache Hadoop and the docker image. Therefore, your host machine should have RAM that exceeds these memory levels.
  • A minimum of 50 GB of free space on the host hard disk. The Docker container image size is 3.11 GB.
⁠Authorization

You must have root access to all the nodes in your cluster.

⁠Internet access

All nodes in your cluster must have access to the internet so that the code and docker images can be downloaded.

⁠** Optional: Get nodes on a cloud environment**

If you do not want to use your own hardware, you can get virtual nodes for a nominal cost from IBM SoftLayer and install the image on those nodes. See Optional: Get nodes for a multinode virtual machine on a cloud environment.

⁠Collecting information about your cluster and preparing your cluster environment

Procedure

  1. Designate one of the nodes in your multinode cluster as the management node by using the fully qualified domain name, such as node1.domain.company.com.
  2. Create a file on the management node called hosts.txt that contains all of the host names of the nodes in your cluster. For example, the hosts.txt file for the current example contains the following node information:

For IBM SoftLayer: If you acquired the nodes from IBM SoftLayer, use those node names.

node1.domain.company.com
node2.domain.company.com
node3.domain.company.com
node4.domain.company.com
node5.domain.company.com

Each node is listed on a separate line. The first line must be the designated management node. The management node, or in the case of the example, node1.domain.company.com, is where the installation starts.

  1. If you did not acquire the nodes from Softlayer, which helped you manage the passwordless SSH, you must set up passwordless SSH connections between the Ambari server host and all other cluster hosts so that the Ambari server can install the Ambari agent automatically on each host.

For IBM SoftLayer: If you acquired the nodes from IBM SoftLayer, you can skip this step because you already managed the passwordless SSH.

  1. Log in to your Linux cluster as root or a user with root privileges.

  2. On the Ambari server host, generate the public and private SSH keys with the following command:ssh-keygen. When you are asked to enter a passphrase, press the Enter key to make sure the passphrase is empty. Otherwise, the host registration at Ambari fails with the following error: Permission denied (publickey,gssapi-keyex,gssapi-with-mic,password).

  3. Copy the SSH public key (id_rsa.pub) to the root account on your target hosts, by using the following commands: ssh-copy-id -i ~/.ssh/id_rsa.pub [email protected] ssh-copy-id -i ~/.ssh/id_rsa.pub [email protected] ssh-copy-id -i ~/.ssh/id_rsa.pub [email protected] ... Ensure that permissions on your .ssh directory are set to 700 and the permissions on the authorized_keys file in that directory are set to either 600 or 640.

  4. From the Ambari server host, connect to each host in the cluster using SSH. For example, enter the following commands: ssh [email protected] ssh [email protected] ssh [email protected] ... You might see this warning on your first connection. Are you sure you want to continue connecting (yes/no)? Enter yes.

  5. Save a copy of the SSH private key (id_rsa) on the machine where you will run the Ambari installation wizard. The file is in $HOME/.ssh/, by default. For more information, see Preparing your environment.

⁠Getting and Installing the IBM Open Platform with Apache Hadoop product image

You are now ready to get and install the image that allows you to experience the IBM Open Platform with Apache Hadoop in the Ambari dashboard, or to experience a combination of IBM Open Platform with Apache Hadoop and the BigInsights Quick Start Edition with value-added services. These instructions cover both a manual and automated installation.

Procedure

  1. If you do not already have the latest IBM Open Platform with Apache Hadoop and IBM BigInsights Quick Start Edition Docker image, download it to the first virtual machine (the management node) from the following site:
  1. Extract the TAR file: tar -fxvz tar -fxvz qse-docker-deployer-4.2.0.0.tar.gz. The TAR file contains the following scripts: cleanHosts.py A script to stop and remove any existing docker containers for the iop-hadoop image from all of the nodes that are listed in the hosts.txt file. It also removes the loaded iop-hadoop image and removes the downloaded files run.py and dockerconfigs.py from the /tmp directory on all nodes. Run the script on the master node only:
python cleanHosts.py hosts.txt

deploy.py A script that does the following actions: a. Downloads the image (iop-hadoop) to all nodes that are listed in the hosts.txt file. b. Copies the run.py and dockerconfigs.py scripts to the /tmp directory on all of the nodes that are listed in the hosts.txt file. c. Creates containers on all of the nodes. d. Starts the ambari-server on the management node, and the ambari-agent on all nodes. When the script completes, it provides a list of nodes and the Ambari web interface URL. You use this URL to provision the cluster with the required services. 3. Manual installation with the Ambari wizard Run the script on the master node only: python deploy.py hosts.txt 4. Automated installation with blueprint If the number of nodes in your cluster is 5, then you can take advantage of the automated blueprint installation. The following parameters signal an automated blueprint installation: -i Install the IBM Open Platform with Apache Spark and Apache Hadoop stack by using automated blueprint installation. python deploy.py hosts.txt -i -v Install the IBM Open Platform with Apache Spark and Apache Hadoop and BigInsights value add stacks by using automated blueprint installation. python deploy.py hosts.txt -v When complete, both the Ambari dashboard and the BigInsights Quick Start Edition with the value-add services are installed. Type the following URL in the browser address field to open the BigInsights Home page: https://management_node_hostname:8443/gateway/default/BigInsightsWeb/index.html

⁠Uninstalling IBM Open Platform with Apache Hadoop

Procedure Before you reinstall IBM Open Platform with Apache Hadoop for non-production environments, you must remove any previous installation by running the following command from a command shell from the management node:

python cleanHosts.py hosts.txt
⁠Hints, Tips, and Troubleshooting

Support for the free offerings Support for the free offerings is provided from the following links:

docker run -d --name iop-hadoop --privileged=true \
  --net=host -p 8080:8080 -p 8670:8670 \
  -p 8440:8440 -p 8441:8441 -p 50010:50010 \
  -p 50020:50020 -p 50070:50070 -p 8188:8188 \
  -p 8190:8190 -p 10200:10200 -p 8020:8020 \
  -p 50075:50075 -p 60010:60010 -p 60020:60020 \
  -p 10000:10000 -p 8088:8088 -p 50060:50060 \
  -p 8032:8032 -p 2022:22 -p 80:80 \
  --ulimit nproc=65535 --ulimit nofile=65535 \
  --ulimit core=65535 ibmcom/iop-hadoop
docker exec -it iop-hadoop ambari-server start

Tag summary

Content type

Image

Digest

sha256:b8926eb07…

Size

1.2 GB

Last updated

over 9 years ago

docker pull ibmcom/iop-hadoop