Sign inSign up

zxb000/tensorrt_llm

By zxb000

•Updated over 1 year ago

tensorrt_llm commit; 2025-04-28 ad15e45f07b0907; and build Hopper and blackwell(90-real;100-real)

Image
Machine learning & AI
1

3.3K

zxb000/tensorrt_llm repository overview

参考 https://github.com/NVIDIA/TensorRT-LLM/tree/main/examples/models/core/deepseek_v3⁠

nvdia TensorRT-LLM 2025-04-28最新编译,commit id是:ad15e45f07b0907。编译时使用CUDA_ARCHS="90-real;100-real",同时支持Hopper(比如H200)和Blackwell(比如B200)架构。 对于DeepSeek来说,DeepSeek-v3 有671B参数,需要大约671GB的GPU内存来存储FP8权重,并且激活张量和KV缓存需要更多的内存。运行DeepSeek V3/R1 FP8&FP4的最低硬件要求如下。

GPUDeepSeek-V3/R1 FP8DeepSeek-V3/R1 FP4
H100 80GB16N/A
H20 141GB8N/A
H20 96GB8N/A
H2008N/A
B200/GB200Not supported yet, WIP4 (8 GPUs is recommended for best perf)

⁠下载/编译

apt-get update && apt-get -y install git git-lfs
git lfs install

git clone https://github.com/NVIDIA/TensorRT-LLM.git
cd TensorRT-LLM
git submodule update --init --recursive
git lfs pull

编译

# 您可以添加CUDA_ARCHS="<CMake格式的架构列表>"可选参数来指定TensorRT-LLM应支持哪些架构。这限制了支持的GPU架构,但有助于减少编译时间:
# Restrict the compilation to Ada and Hopper architectures. 100是Blackwell(如B200)的算力标记
make -j 60 -C docker release_build CUDA_ARCHS="90-real;100-real" IMAGE_TAG="nv-sm190"
# TensorRT-LLM 的示例安装在 /app/tensorrt_llm/examples 目录中。

⁠启动容器

docker run -itd --privileged --ipc=host --gpus all --network=host --ulimit memlock=-1 --ulimit stack=67108864 -v /deepseek:/root/.cache/huggingface --name tensor zxb000/tensorrt_llm:20250428-ad15e45-nv-sm190 bash

进入容器,后进行以下操作。

⁠快速开始

cd examples/pytorch/

python3 quickstart_advanced.py --model_dir /root/.cache/huggingface/DeepSeek-R1-FP4 --tp_size 8

输出:
Processed requests: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:02<00:00,  1.66it/s]
Prompt: 'Hello, my name is', Generated text: " Dr. Anadale. I teach philosophy at Mount St. Mary's University in Emmitsburg, Maryland. This is the first in a series of videos"
Prompt: 'The president of the United States is', Generated text: ' the head of state and head of government of the United States, indirectly elected to a four-year term via the Electoral College. The officeholder leads the executive branch'
Prompt: 'The capital of France is', Generated text: ' Paris. Paris is one of the most famous and iconic cities in the world, known for its rich history, culture, art, fashion, and cuisine. It'
Prompt: 'The future of AI is', Generated text: ' a topic of great interest and speculation. While it is impossible to predict the future with certainty, there are several trends and possibilities that experts have discussed regarding the future'

⁠运行

cd /root/.cache/huggingface

cat >./extra-llm-api-config.yml <<EOF
pytorch_backend_config:
    use_cuda_graph: true
    cuda_graph_padding_enabled: true
    cuda_graph_batch_sizes:
    - 1
    - 2
    - 4
    - 8
    - 16
    - 32
    - 64
    - 128
    - 256
    - 384
    print_iter_log: true
    enable_overlap_scheduler: true
enable_attention_dp: true
EOF

trtllm-serve \
  /root/.cache/huggingface/DeepSeek-R1-FP4 \
  --host localhost \
  --port 8000 \
  --backend pytorch \
  --max_batch_size 161 \
  --max_num_tokens 8192 \
  --tp_size 8 \
  --ep_size 8 \
  --pp_size 1 \
  --kv_cache_free_gpu_memory_fraction 0.7 \
  --extra_llm_api_options ./extra-llm-api-config.yml

说明

  1. --kv_cache_free_gpu_memory_fraction 分配模型权重和缓冲区后,为KV缓存保留的可用的GPU内存比例。

Tag summary

Content type

Image

Digest

sha256:44ed3133f…

Size

23.7 GB

Last updated

over 1 year ago

docker pull zxb000/tensorrt_llm