This image, built using the Dockerfile, provides an environment for ROCm 6.2 with PyTorch 2.6.0.dev20240920+rocm6.2 on Ubuntu 20.04. It sets up a user-friendly development environment for machine learning workloads, particularly focusing on VLLM integration. The image includes essential dependencies and configuration for optimal performance on AMD GPUs. (I have tested this image on the AMD MI250 GPU.)
FROM rocm/pytorch:rocm6.2_ubuntu20.04_py3.9_pytorch_release_2.3.0
ENV PYTORCH_ROCM_ARCH="gfx90a;gfx942"
ENV VLLM_USE_TRITON_FLASH_ATTN=0
# Install dependencies
WORKDIR /
RUN git clone https://github.com/vllm-project/vllm.git
RUN cd vllm && git checkout 260d40b5ea48df9421325388abcc8d907a560fc5
RUN mv vllm vllm-workspace
RUN chmod -R 755 /vllm-workspace
WORKDIR /vllm-workspace
RUN pip uninstall torch -y
RUN pip install --no-cache-dir torch==2.6.0.dev20240920+rocm6.2 torchvision==0.20.0.dev20240920+rocm6.2 --index-url https://download.pytorch.org/whl/nightly/rocm6.2
RUN pip install --no-cache-dir /opt/rocm/share/amd_smi &&\
pip install "numpy<2" &&\
pip install --no-cache-dir -r requirements-rocm.txt &&\
pip install --upgrade numba scipy huggingface-hub[cli]
RUN wget -N https://github.com/ROCm/vllm/raw/b826cba/rocm_patch/libamdhip64.so.6 -P /opt/rocm/lib && \
rm -f "$(python3 -c 'import torch; print(torch.__path__[0])')"/lib/libamdhip64.so*
RUN python3 setup.py develop
RUN pip install numpy==1.26.3
RUN pip cache purge
CMD ["bash", "-c"]
To use this image, follow these steps for running the container and interacting with the ROCm GPUs. Below is an example command to run the container:
> ls /dev/dri/
by-path card1 card3 card5 card7 renderD128 renderD130 renderD132 renderD134
card0 card2 card4 card6 card8 renderD129 renderD131 renderD133 renderD135
> docker run --rm -it \
--device=/dev/kfd \
--device=/dev/dri/renderD134 \
--device=/dev/dri/renderD135 \
--security-opt seccomp=unconfined \
--cap-add=SYS_PTRACE \
--shm-size=30g \
--group-add video \
deagwon97/vllm-rocm:vllm0.6.1_pytorch2.6.0_rocm6.2 \
bash
root@0031463a128f:~$ rocm-smi
========================================= ROCm System Management Interface =========================================
=================================================== Concise Info ===================================================
Device Node IDs Temp Power Partitions SCLK MCLK Fan Perf PwrCap VRAM% GPU%
(DID, GUID) (Edge) (Avg) (Mem, Compute, ID)
====================================================================================================================
0 8 0x740c, 27082 33.0°C 91.0W N/A, N/A, 0 800Mhz 1600Mhz 0% auto 560.0W 0% 0%
1 9 0x740c, 23432 35.0°C N/A N/A, N/A, 0 800Mhz 1600Mhz 0% auto 0.0W 0% 0%
====================================================================================================================
=============================================== End of ROCm SMI Log ================================================
root@0031463a128f:~$ vllm
usage: vllm [-h] {serve,complete,chat} ...
vllm: error: the following arguments are required: {serve,complete,chat}
root@0031463a128f:~$ python
Python 3.9.19 (main, May 6 2024, 19:43:03)
[GCC 11.2.0] :: Anaconda, Inc. on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import torch
>>> torch.__version__
'2.6.0.dev20240920+rocm6.2'
>>> torch.cuda.is_available()
True
>>> import vllm
>>> vllm.__version__
'0.6.1.post2'
This command will start the container with ROCm-enabled GPUs, allowing you to run Python scripts, including PyTorch and VLLM-related operations.
Content type
Image
Digest
sha256:0d35ee38f…
Size
20.4 GB
Last updated
almost 2 years ago
docker pull deagwon97/vllm-rocm:rocm6.2_torch2.5.0a0gitcedc116_vllm0.6.1.post2rocm624