Atlas Device Plugin is one of the core components of the MindCluster cluster scheduling suite, deployed on compute nodes to provide resource discovery and reporting strategies tailored for Atlas devices.
Kubernetes needs to be aware of resource information for scheduling. Beyond basic CPU and memory information, the Kubernetes device plugin mechanism allows users to define custom resource types and customize resource discovery and reporting strategies. MindCluster provides the Atlas Device Plugin service deployed on compute nodes to offer resource discovery and reporting strategies suitable for Atlas devices.
Device Discovery: Obtains chip type and model information from the driver and reports it to kubelet and the upper-level ClusterD service. Supports discovering the number of devices from the Atlas device driver and reporting the count to the Kubernetes system. Supports discovering virtual devices split from physical devices and reporting them to the Kubernetes system.
Health Check: Subscribes to chip fault information from the driver, reports chip status to kubelet, and reports chip status along with specific fault details to the upper-level scheduling service. Supports detecting the health status of Atlas devices. When a device is in an unhealthy state, it is reported to the Kubernetes system, which automatically removes the unhealthy device from the available list. The health status of virtual devices is determined by the physical devices from which they are split.
Device Allocation: Supports allocating Atlas devices in the Kubernetes system. Supports NPU device rescheduling — when a device fails, a new container is automatically started, a healthy device is mounted, and the training task is rebuilt. During the resource mounting phase, it retrieves the chip information selected by the cluster scheduler and passes it to Atlas Docker Runtime via environment variables for mounting.
Fault Handling: Configurable fault handling levels, with the ability to escalate fault handling levels when faults recur or persist for extended periods. If a faulty chip is idle and can recover after a restart, a hot reset is performed on the chip.
Network Fault Monitoring: Subscribes to Lingqu network fault information from the Lingqu driver, reports network status to kubelet, and reports Lingqu network status along with specific fault details to the upper-level scheduling service.
Starting from version v26.1.0, tags follow the format below:
<version>-<os>
| Field | Example | Description |
|---|---|---|
version | v26.1.1 | Version Number of Atlas Device Plugin |
os | ubuntu22.04 | Operating System for Atlas Device Plugin Images |
The following are all images of the latest released 26.1.1 version of Atlas Device Plugin. For all historical version Tags, please refer to Supported Tags.
| Tag | Dockerfile | Image Content |
|---|---|---|
v26.1.1-ubuntu22.04 | Dockerfile.ubuntu | Atlas Device Plugin v26.1.1 (Base Image: Ubuntu 22.04) |
v26.1.1-openeuler24.03 | Dockerfile.openeuler | Atlas Device Plugin v26.1.1 (Base Image: openEuler 24.03) |
The tags of versions before v26.1.0 follow the format below:
<version>
| Field | Example | Description |
|---|---|---|
version | v26.0.0 | Version Number of Atlas Device Plugin |
| Tag | Dockerfile | Image Content |
|---|---|---|
v26.0.0 | Dockerfile | Atlas Device Plugin v26.0.0 (Base Image: Ubuntu 22.04) |
| Software | Supported Versions | Installation Location | Description |
|---|---|---|---|
| Kubernetes | 1.17.x~1.34.x (1.19.x or later recommended) | All nodes | See Kubernetes Documentation |
| Docker | 18.09.x~28.5.1 | All nodes | Available from Docker |
| Containerd | 1.4.x~2.1.4 (1.6.x recommended) | All nodes | Available from Containerd |
| Atlas AI Processor Driver and Firmware | See version compatibility table | Compute nodes | See "Installing NPU Driver and Firmware" in the CANN Software Installation Guide |
| UMDK software package | See version compatibility table | Compute nodes | Necessary for Atlas 850、Atlas 950 SuperPod Products |
| Resource | Requirement |
|---|---|
| CPU | 0.5 cores |
| Memory | 0.5 GB |
The host machine must have the driver and firmware installed. For details, see "Installing NPU Driver and Firmware" in the CANN Software Installation Guide (Commercial Edition).
Pull the official image
Pull the Atlas Device Plugin image from AscendHub, replacing {tag} with the actual version.
docker pull swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-k8sdeviceplugin:{tag}
Retag the image
Retag the official image with a local tag for consistent naming and easier operations management.
docker tag swr.cn-south-1.myhuaweicloud.com/ascendhub/ascend-k8sdeviceplugin:{tag} ascend-k8sdeviceplugin:{tag}
Example: build an Atlas Device Plugin image of architecture linux-aarch64, version v26.1.1, based on Ubuntu 22.04.
Obtain the target Dockerfile
Navigate to the chapter Supported Tags and Dockerfile Links, open the Dockerfile.ubuntu link corresponding to your target version, and save the file to a local directory on your aarch64 environment.
Build the Docker image locally (disable cache to ensure a clean build)
docker build --no-cache -t ascend-k8sdeviceplugin:v26.1.1 ./ -f Dockerfile.ubuntu
Important Notes If your Docker version is earlier than 18.09 or BuildKit is not manually enabled, the TARGETPLATFORM variable cannot be read during image building, which will cause the image build to fail.
- TARGETPLATFORM is a built-in global variable of Docker BuildKit for identifying the target build platform, e.g. linux/amd64, linux/arm64.
- This variable is automatically injected only after BuildKit is enabled. It cannot be used in legacy Docker environments or environments where BuildKit is disabled by default.
- Run the following command before building to enable BuildKit temporarily:
export DOCKER_BUILDKIT=1
Example: Build an Atlas Device Plugin image of architecture linux-aarch64, version v26.0.0, based on Ubuntu 22.04.
Download the officially released component package
wget https://gitcode.com/Ascend/mind-cluster/releases/download/v26.0.0/Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64.zip
Extract the package to a custom directory
unzip Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64.zip -d Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64
Enter the extracted working directory
cd Ascend-mindxdl-device-plugin_26.0.0_linux-aarch64
Build the Docker image locally (disable cache to ensure a clean build)
docker build --no-cache -t ascend-k8sdeviceplugin:v26.0.0 ./ -f Dockerfile
Label Kubernetes nodes
Label nodes according to the Atlas processor model for cluster scheduling. Replace <node-name> with the actual node
name.
# Example: Label an Atlas 910 node
kubectl label nodes <node-name> accelerator=huawei-Ascend910
Start Atlas Device Plugin
Select the appropriate YAML resource file based on the device model and scheduling requirements. Replace {tag} in
the YAML file with the actual image version.
# Configuration file for products excluding Atlas 200I SoC A1 core board without Volcano.
kubectl apply -f device-plugin-{version}.yaml
# Configuration file for products excluding Atlas 200I SoC A1 core board with Volcano.
kubectl apply -f device-plugin-volcano-{version}.yaml
Verify deployment
kubectl get pods -A | grep device-plugin
Expected result: The device-plugin related Pods in the corresponding namespace should be in Running state.
Check node resources
kubectl describe node <npu-node-name> | grep "huawei.com/Ascend"
Expected result: The huawei.com/Ascend resource capacity and allocatable resources should be displayed correctly.
View the license information for the Mind series software contained in these images.
As with all container images, pre-installed software packages (Python, system libraries, etc.) may be subject to their respective license agreements.
Content type
Image
Digest
sha256:d9f246add…
Size
54.4 MB
Last updated
11 days ago
docker pull ascendai/ascend-k8sdeviceplugin:v26.1.1-ubuntu22.04