tensorrt_llm commit; 2025-04-28 ad15e45f07b0907; and build Hopper and blackwell(90-real;100-real)
3.3K
参考 https://github.com/NVIDIA/TensorRT-LLM/tree/main/examples/models/core/deepseek_v3
nvdia TensorRT-LLM 2025-04-28最新编译,commit id是:ad15e45f07b0907。编译时使用CUDA_ARCHS="90-real;100-real",同时支持Hopper(比如H200)和Blackwell(比如B200)架构。
对于DeepSeek来说,DeepSeek-v3 有671B参数,需要大约671GB的GPU内存来存储FP8权重,并且激活张量和KV缓存需要更多的内存。运行DeepSeek V3/R1 FP8&FP4的最低硬件要求如下。
| GPU | DeepSeek-V3/R1 FP8 | DeepSeek-V3/R1 FP4 |
|---|---|---|
| H100 80GB | 16 | N/A |
| H20 141GB | 8 | N/A |
| H20 96GB | 8 | N/A |
| H200 | 8 | N/A |
| B200/GB200 | Not supported yet, WIP | 4 (8 GPUs is recommended for best perf) |
apt-get update && apt-get -y install git git-lfs
git lfs install
git clone https://github.com/NVIDIA/TensorRT-LLM.git
cd TensorRT-LLM
git submodule update --init --recursive
git lfs pull
编译
# 您可以添加CUDA_ARCHS="<CMake格式的架构列表>"可选参数来指定TensorRT-LLM应支持哪些架构。这限制了支持的GPU架构,但有助于减少编译时间:
# Restrict the compilation to Ada and Hopper architectures. 100是Blackwell(如B200)的算力标记
make -j 60 -C docker release_build CUDA_ARCHS="90-real;100-real" IMAGE_TAG="nv-sm190"
# TensorRT-LLM 的示例安装在 /app/tensorrt_llm/examples 目录中。
docker run -itd --privileged --ipc=host --gpus all --network=host --ulimit memlock=-1 --ulimit stack=67108864 -v /deepseek:/root/.cache/huggingface --name tensor zxb000/tensorrt_llm:20250428-ad15e45-nv-sm190 bash
进入容器,后进行以下操作。
cd examples/pytorch/
python3 quickstart_advanced.py --model_dir /root/.cache/huggingface/DeepSeek-R1-FP4 --tp_size 8
输出:
Processed requests: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████| 4/4 [00:02<00:00, 1.66it/s]
Prompt: 'Hello, my name is', Generated text: " Dr. Anadale. I teach philosophy at Mount St. Mary's University in Emmitsburg, Maryland. This is the first in a series of videos"
Prompt: 'The president of the United States is', Generated text: ' the head of state and head of government of the United States, indirectly elected to a four-year term via the Electoral College. The officeholder leads the executive branch'
Prompt: 'The capital of France is', Generated text: ' Paris. Paris is one of the most famous and iconic cities in the world, known for its rich history, culture, art, fashion, and cuisine. It'
Prompt: 'The future of AI is', Generated text: ' a topic of great interest and speculation. While it is impossible to predict the future with certainty, there are several trends and possibilities that experts have discussed regarding the future'
cd /root/.cache/huggingface
cat >./extra-llm-api-config.yml <<EOF
pytorch_backend_config:
use_cuda_graph: true
cuda_graph_padding_enabled: true
cuda_graph_batch_sizes:
- 1
- 2
- 4
- 8
- 16
- 32
- 64
- 128
- 256
- 384
print_iter_log: true
enable_overlap_scheduler: true
enable_attention_dp: true
EOF
trtllm-serve \
/root/.cache/huggingface/DeepSeek-R1-FP4 \
--host localhost \
--port 8000 \
--backend pytorch \
--max_batch_size 161 \
--max_num_tokens 8192 \
--tp_size 8 \
--ep_size 8 \
--pp_size 1 \
--kv_cache_free_gpu_memory_fraction 0.7 \
--extra_llm_api_options ./extra-llm-api-config.yml
说明
Content type
Image
Digest
sha256:44ed3133f…
Size
23.7 GB
Last updated
over 1 year ago
docker pull zxb000/tensorrt_llm