github pod-gpu-metrics-exporter
A simple go http server serving per pod GPU metrics at localhost:9400/gpu/metrics. The exporter connects to kubelet gRPC server (/var/lib/kubelet/pod-resources) to identify the GPUs running on a pod leveraging Kubernetes device assignment feature and appends the GPU device's pod information to metrics collected by dcgm-exporter.
The http server allows Prometheus to scrape GPU metrics directly via a separate endpoint without relying on node-exporter. But if you still want to scrape GPU metrics via node-exporter, follow these instructions.
add /var/run/docker.sock # used to get pod container pid
add hostPID: true # used to check whether the parent porcess of the gpu process is pod container process
use nvml # used to get the gpu process used memory
# TYPE dcgm_process_mem_used gauge
# HELP dcgm_process_mem_used process memory used (in MiB).
dcgm_process_mem_used{gpu="0",uuid="GPU-ad365448-e6c2-68f2-24e4-517b1e56e937",pod_name="test-pod-01",pod_namespace="default",container_name="nvidia-test",process_name="python",process_pid="617",process_type="C"} 847
dcgm_process_mem_used{gpu="0",uuid="GPU-ad365448-e6c2-68f2-24e4-517b1e56e937",pod_name="test-pod-01",pod_namespace="default",container_name="nvidia-test",process_name="python",process_pid="16187",process_type="C"} 587
Content type
Image
Digest
Size
28.9 MB
Last updated
almost 6 years ago
docker pull ruanxingbaozi/pod-gpu-metrics-exporter:v1.0.0-alpha