Sign inSign up

lovedheart/lvllm

By lovedheart

•Updated 3 months ago

A vLLM extension for GPU + NUMA dual parallel inference

Image
0

224

lovedheart/lvllm repository overview

⁠Lvllm -- GPU + NUMA Dual Parallel Inference

Lvllm is a vLLM extension for GPU + NUMA dual parallel inference, optimized for large MoE (Mixture-of-Experts) models on multi-socket CPU/GPU systems.

⁠Features

  • GPU+NUMA hybrid inference -- Offloads MoE experts to CPU NUMA nodes while keeping attention layers on GPU
  • Thread binding -- Bind worker threads to CPU cores or NUMA nodes for optimal cache locality
  • Layer-wise expert loading -- Load expert weights per-layer on demand with prefetching
  • NVFP4 support -- Leverages Blackwell FP4 tensor cores for efficient MoE inference
  • CPU power saving -- Throttle idle CPU cores when not processing

⁠Quick Start

docker run --gpus all --privileged \
  -v /path/to/model:/models \
  -p 8000:8000 \
  lovedheart/lvllm:latest \
  --model /models \
  --host 0.0.0.0 \
  --port 8000 \
  --tensor-parallel-size 1 \
  --trust-remote-code

⁠Environment Variables

VariableDefaultDescription
LVLLM_MOE_NUMA_ENABLED1Enable GPU+NUMA hybrid inference
LK_THREAD_BINDINGCPU_COREThread binding policy (CPU_CORE or NUMA_NODE)
LK_THREADSautoNumber of worker threads for CPU experts
OMP_NUM_THREADSautoOpenMP threads
LVLLM_GPU_RESIDENT_MOE_LAYERSunsetGPU-resident expert layer indices (e.g. 0-5)
LVLLM_GPU_PREFETCH_WINDOWunsetExpert prefetch window size
LVLLM_GPU_PREFILL_MIN_BATCH_SIZEunsetMin batch size for GPU prefill
LK_POWER_SAVING0Enable CPU idle core throttling
LVLLM_ENABLE_NUMA_INTERLEAVE1Interleave memory across NUMA nodes
LVLLM_ENABLE_MOE_LAYERWISE_LOAD1Load MoE experts layer-by-layer

⁠Usage Examples

⁠Basic inference
docker run --gpus all --privileged \
  -v /path/to/model:/models \
  -p 8000:8000 \
  -e LK_THREADS=8 \
  -e OMP_NUM_THREADS=8 \
  lovedheart/lvllm:latest \
  --model /models \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.3 \
  --trust-remote-code \
  --enable-prefix-caching
⁠NVIDIA Blackwell NVFP4
docker run --gpus all --privileged \
  -v /path/to/model:/models \
  -p 8000:8000 \
  -e LVLLM_GPU_RESIDENT_MOE_LAYERS=0-5 \
  -e LVLLM_GPU_PREFETCH_WINDOW=1 \
  -e LVLLM_GPU_PREFILL_MIN_BATCH_SIZE=4096 \
  lovedheart/lvllm:latest \
  --model /models \
  --tensor-parallel-size 2 \
  --compilation_config.cudagraph_mode FULL_DECODE_ONLY \
  --compilation_config.mode VLLM_COMPILE

⁠Repository

Tag summary

Content type

Image

Digest

sha256:62344501e…

Size

8 GB

Last updated

3 months ago

docker pull lovedheart/lvllm