Sign inSign up

ascendai/clusterd

By ascendai

•Updated 6 days ago

Image
0

724

ascendai/clusterd repository overview

⁠Cluster Scheduling Component ClusterD

⁠Quick Reference


⁠ClusterD

ClusterD is one of the MindCluster cluster scheduling components, deployed on management nodes. It collects and aggregates cluster job, resource, and fault information along with their impact scope, performs statistical analysis from the dimensions of jobs, chips, and faults, and uniformly determines fault handling levels and policies.

⁠Use Cases

A single node may experience multiple faults. If each node handles faults independently, jobs may simultaneously be subject to multiple recovery strategies. To coordinate job handling levels, MindCluster provides the ClusterD service deployed on management nodes. ClusterD collects and aggregates cluster job, resource, and fault information along with their impact scope, performs statistical analysis from the dimensions of jobs, chips, and faults, and uniformly determines fault handling levels and policies.

⁠Features
  • Obtains chip, node, and network information from Atlas Device Plugin and NodeD components, and retrieves public fault information from ConfigMap or gRPC.
  • Aggregates the above fault information for upper-level cluster scheduling services to query.
  • Establishes connections with training containers to control training processes for recomputation actions.
  • Interacts with out-of-band services to transmit job information.
⁠Upstream and Downstream Dependencies
  1. Obtains chip information from Atlas Device Plugin on each compute node.
  2. Obtains CPU, memory, and disk health status information, DPC shared storage fault information, and Lingqu network fault information from NodeD on each compute node.
  3. Retrieves public fault information from ConfigMap or gRPC.
  4. Aggregates resource information across the entire cluster and reports it to Ascend-volcano-plugin.
  5. Monitors cluster job information and reports job status, resource usage, and other information to CCAE.
  6. Interacts with in-container processes to control training processes for recomputation.

⁠Tag Convention

Starting from version v26.1.0, tags follow the format below:

<version>-<os>
FieldExampleDescription
versionv26.1.1Version Number of ClusterD
osubuntu22.04Operating System for ClusterD Images
⁠ClusterD Latest Version 26.1.1

The following are all images of the latest released 26.1.1 version of ClusterD. For all historical version Tags, please refer to Supported Tags⁠.

TagDockerfileImage Content
v26.1.1-ubuntu22.04Dockerfile.ubuntu⁠ClusterD v26.1.1 (Base Image: Ubuntu 22.04)
v26.1.1-openeuler24.03Dockerfile.openeuler⁠ClusterD v26.1.1 (Base Image: openEuler 24.03)

The tags of versions before v26.1.0 follow the format below:

<version>
FieldExampleDescription
versionv26.0.0Version Number of ClusterD
⁠ClusterD 26.0.0
TagDockerfileImage Content
v26.0.0Dockerfile⁠ClusterD v26.0.0 (Base Image: Ubuntu 22.04)

⁠Quick Start

⁠Prerequisites
⁠Software Dependencies
SoftwareSupported VersionsInstallation LocationDescription
Kubernetes1.17.x~1.34.x (1.19.x or later recommended)All nodesSee Kubernetes Documentation⁠
Atlas Device PluginSame version as ClusterDCompute nodesClusterD depends on Atlas Device Plugin to report chip information
NodeDSame version as ClusterDCompute nodesClusterD depends on NodeD to report node fault information
⁠Hardware Requirements
ResourceUp to 100 Nodes500 Nodes1000 Nodes
CPU1 core2 cores4 cores
Memory1 GB2 GB8 GB
⁠Obtain ClusterD Image Online
  1. Pull the official image

    Pull the ClusterD image from AscendHub, replacing {tag} with the actual version.

    docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/clusterd:{tag}
    
  2. Retag the image

    Retag the official image with a local tag for consistent naming and easier operations management.

    docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/clusterd:{tag} clusterd:{tag}
    
⁠Build Locally (Optional)
⁠Local Build Steps for v26.1.0 and Later Versions

Example: build an ClusterD image of architecture linux-aarch64, version v26.1.1, based on Ubuntu 22.04.

  1. Obtain the target Dockerfile

    Navigate to the chapter Supported Tags and Dockerfile Links, open the Dockerfile.ubuntu link corresponding to your target version, and save the file to a local directory on your aarch64 environment.

  2. Build the Docker image locally (disable cache to ensure a clean build)

    docker build --no-cache -t clusterd:v26.1.1 ./ -f Dockerfile.ubuntu
    

Important Notes If your Docker version is earlier than 18.09 or BuildKit is not manually enabled, the TARGETPLATFORM variable cannot be read during image building, which will cause the image build to fail.

  1. TARGETPLATFORM is a built-in global variable of Docker BuildKit for identifying the target build platform, e.g. linux/amd64, linux/arm64.
  2. This variable is automatically injected only after BuildKit is enabled. It cannot be used in legacy Docker environments or environments where BuildKit is disabled by default.
  3. Run the following command before building to enable BuildKit temporarily:
export DOCKER_BUILDKIT=1
⁠Local Image Build Process for Versions Before v26.1.0

Example: Build an ClusterD image of architecture linux-aarch64, version v26.0.0, based on Ubuntu 22.04.

  1. Download the officially released component package

    wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-clusterd_26.0.0_linux-aarch64.zip
    
  2. Extract the package to a custom directory

    unzip Ascend-mindxdl-clusterd_26.0.0_linux-aarch64.zip -d Ascend-mindxdl-clusterd_26.0.0_linux-aarch64
    
  3. Enter the extracted working directory

    cd Ascend-mindxdl-clusterd_26.0.0_linux-aarch64
    
  4. Build the Docker image locally (disable cache to ensure a clean build)

    docker build --no-cache -t clusterd:v26.0.0 ./ -f Dockerfile
    
⁠Deploy ClusterD
  1. Start ClusterD

    Before deployment, replace the image {tag} in the YAML file with the actual image version.

    kubectl apply -f clusterd-{version}.yaml
    
  2. Verify deployment

    kubectl get pods -A | grep clusterd
    

    Expected result: The clusterd related Pods in the corresponding namespace should be in Running state.


⁠License

View the license information⁠ for the Mind series software contained in these images.

As with all container images, pre-installed software packages (Python, system libraries, etc.) may be subject to their respective license agreements.

Tag summary

Content type

Image

Digest

sha256:a5234898b…

Size

44.7 MB

Last updated

6 days ago

docker pull ascendai/clusterd:v6.0.0.SPC3