Sign inSign up

ruanxingbaozi/pod-gpu-metrics-exporter

By ruanxingbaozi

•Updated almost 6 years ago

Image
0

434

ruanxingbaozi/pod-gpu-metrics-exporter repository overview

github pod-gpu-metrics-exporter⁠

⁠Pod GPU Metrics Exporter

A simple go http server serving per pod GPU metrics at localhost:9400/gpu/metrics. The exporter connects to kubelet gRPC server (/var/lib/kubelet/pod-resources) to identify the GPUs running on a pod leveraging Kubernetes device assignment feature⁠ and appends the GPU device's pod information to metrics collected by dcgm-exporter⁠.

The http server allows Prometheus to scrape GPU metrics directly via a separate endpoint without relying on node-exporter. But if you still want to scrape GPU metrics via node-exporter, follow these instructions⁠.

⁠Prerequisites
⁠For GPUShare
⁠Add gpu process memory used metrics
add /var/run/docker.sock    # used to get pod container pid
add hostPID: true           # used to check whether the parent porcess of the gpu process is pod container process 
use nvml                    # used to get the gpu process used memory
⁠metrice sample output
# TYPE dcgm_process_mem_used gauge
# HELP dcgm_process_mem_used process memory used (in MiB).
dcgm_process_mem_used{gpu="0",uuid="GPU-ad365448-e6c2-68f2-24e4-517b1e56e937",pod_name="test-pod-01",pod_namespace="default",container_name="nvidia-test",process_name="python",process_pid="617",process_type="C"} 847
dcgm_process_mem_used{gpu="0",uuid="GPU-ad365448-e6c2-68f2-24e4-517b1e56e937",pod_name="test-pod-01",pod_namespace="default",container_name="nvidia-test",process_name="python",process_pid="16187",process_type="C"} 587

Tag summary

Content type

Image

Digest

Size

28.9 MB

Last updated

almost 6 years ago

docker pull ruanxingbaozi/pod-gpu-metrics-exporter:v1.0.0-alpha