Sign inSign up

deagwon97/vllm-rocm

By deagwon97

•Updated almost 2 years ago

Image
0

242

deagwon97/vllm-rocm repository overview

⁠Overview

This image, built using the Dockerfile, provides an environment for ROCm 6.2 with PyTorch 2.6.0.dev20240920+rocm6.2 on Ubuntu 20.04. It sets up a user-friendly development environment for machine learning workloads, particularly focusing on VLLM integration. The image includes essential dependencies and configuration for optimal performance on AMD GPUs. (I have tested this image on the AMD MI250 GPU.)

⁠Dockerfile

FROM rocm/pytorch:rocm6.2_ubuntu20.04_py3.9_pytorch_release_2.3.0

ENV PYTORCH_ROCM_ARCH="gfx90a;gfx942"
ENV VLLM_USE_TRITON_FLASH_ATTN=0
# Install dependencies
WORKDIR /

RUN git clone https://github.com/vllm-project/vllm.git
RUN cd vllm && git checkout 260d40b5ea48df9421325388abcc8d907a560fc5
RUN mv vllm vllm-workspace

RUN chmod -R 755 /vllm-workspace

WORKDIR /vllm-workspace

RUN pip uninstall torch -y 
RUN pip install --no-cache-dir torch==2.6.0.dev20240920+rocm6.2 torchvision==0.20.0.dev20240920+rocm6.2  --index-url https://download.pytorch.org/whl/nightly/rocm6.2
RUN pip install --no-cache-dir /opt/rocm/share/amd_smi &&\
    pip install "numpy<2" &&\
    pip install --no-cache-dir -r requirements-rocm.txt &&\
    pip install --upgrade numba scipy huggingface-hub[cli]

RUN wget -N https://github.com/ROCm/vllm/raw/b826cba/rocm_patch/libamdhip64.so.6 -P /opt/rocm/lib && \
    rm -f "$(python3 -c 'import torch; print(torch.__path__[0])')"/lib/libamdhip64.so*

RUN python3 setup.py develop

RUN pip install numpy==1.26.3

RUN pip cache purge

CMD ["bash", "-c"]

⁠how to use

To use this image, follow these steps for running the container and interacting with the ROCm GPUs. Below is an example command to run the container:

> ls /dev/dri/
by-path  card1  card3  card5  card7  renderD128  renderD130  renderD132  renderD134
card0    card2  card4  card6  card8  renderD129  renderD131  renderD133  renderD135

> docker run --rm -it \
  --device=/dev/kfd \
  --device=/dev/dri/renderD134 \
  --device=/dev/dri/renderD135  \
  --security-opt seccomp=unconfined \
  --cap-add=SYS_PTRACE \
  --shm-size=30g \
  --group-add video \
  deagwon97/vllm-rocm:vllm0.6.1_pytorch2.6.0_rocm6.2 \
  bash

root@0031463a128f:~$ rocm-smi
========================================= ROCm System Management Interface =========================================
=================================================== Concise Info ===================================================
Device  Node  IDs              Temp    Power  Partitions          SCLK    MCLK     Fan  Perf  PwrCap  VRAM%  GPU%  
              (DID,     GUID)  (Edge)  (Avg)  (Mem, Compute, ID)                                                   
====================================================================================================================
0       8     0x740c,   27082  33.0°C  91.0W  N/A, N/A, 0         800Mhz  1600Mhz  0%   auto  560.0W  0%     0%    
1       9     0x740c,   23432  35.0°C  N/A    N/A, N/A, 0         800Mhz  1600Mhz  0%   auto  0.0W    0%     0%    
====================================================================================================================
=============================================== End of ROCm SMI Log ================================================
root@0031463a128f:~$ vllm
usage: vllm [-h] {serve,complete,chat} ...
vllm: error: the following arguments are required: {serve,complete,chat}
root@0031463a128f:~$ python
Python 3.9.19 (main, May  6 2024, 19:43:03) 
[GCC 11.2.0] :: Anaconda, Inc. on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import torch
>>> torch.__version__
'2.6.0.dev20240920+rocm6.2'
>>> torch.cuda.is_available()
True
>>> import vllm
>>> vllm.__version__
'0.6.1.post2'

This command will start the container with ROCm-enabled GPUs, allowing you to run Python scripts, including PyTorch and VLLM-related operations.

Tag summary

Content type

Image

Digest

sha256:0d35ee38f…

Size

20.4 GB

Last updated

almost 2 years ago

docker pull deagwon97/vllm-rocm:rocm6.2_torch2.5.0a0gitcedc116_vllm0.6.1.post2rocm624