Public | Automated Build

Last pushed: 10 days ago
Short Description
Hadoop namenode of a hadoop cluster
Full Description

Changes

Version 1.1.0 introduces healthchecks for the containers.

Hadoop Docker

To deploy an example HDFS cluster, run:

  docker network create hadoop
  docker-compose up

The configuration parameters can be specified in the hadoop.env file or as environmental variables for specific services (e.g. namenode, datanode etc.):

  CORE_CONF_fs_defaultFS=hdfs://namenode:8020

CORE_CONF corresponds to core-site.xml. fs_defaultFS=hdfs://namenode:8020 will be transformed into:

  <property><name>fs.defaultFS</name><value>hdfs://namenode:8020</value></property>

To define dash inside a configuration parameter, use double underscore, such as YARN_CONF_yarn_logaggregationenable=true (yarn-site.xml):

  <property><name>yarn.log-aggregation-enable</name><value>true</value></property>

The available configurations are:

  • /etc/hadoop/core-site.xml CORE_CONF
  • /etc/hadoop/hdfs-site.xml HDFS_CONF
  • /etc/hadoop/yarn-site.xml YARN_CONF
  • /etc/hadoop/httpfs-site.xml HTTPFS_CONF
  • /etc/hadoop/kms-site.xml KMS_CONF
  • /etc/hadoop/mapred-site.xml MAPRED_CONF

If you need to extend some other configuration file, refer to base/entrypoint.sh bash script.

After starting the example Hadoop cluster, you should be able to access interfaces of all the components (substitute domain names by IP addresses from network inspect hadoop command):

Running example MapReduce job

To run example map reduce job on the Hadoop cluster run:

make wordcount

See Makefile and submit docker image for more details.

Docker Pull Command
Owner
bde2020
Source Repository

Comments (1)
ces241
6 months ago

On starting the namenode after creation, i am getting error "Cluster name not specified" .. Below is the log:

Configuring core

  • Setting fs.defaultFS=hdfs://4477ed5883d0:8020
    Configuring hdfs
  • Setting dfs.namenode.name.dir=file:///hadoop/dfs/name
    Configuring yarn
    Configuring httpfs
    Configuring kms
    Configuring for multihomed network
    Cluster name not specified

Am i missing anything?