https://github.com/OwenTruong/exllamav3/blob/feat/docker/doc/docker_convert.md
972
https://github.com/OwenTruong/exllamav3/blob/feat/docker/doc/docker_convert.md
Note, your host machine's GPU drivers must be able to run the CUDA version of the container!
Make sure you are at the root directory of the repository, then run the below.
python ./docker/convert/cuda/build.py
Building can takes a while and may require more storage for storing the images mid building.
Because of that, images are also available on Docker Hub.
To use an image from Docker Hub, replace all instance of the exllamav3 image in the create step below with the name of the image on Docker Hub. For example:
docker create \
... \
johndoe/exllamav3:preview-4-17-25--cuda12.8.1
Next, we create the container so that we can reuse it. Replace all instances of /path/to/dir with the root directory where you store all of your llm files.
Note that the -v option allows us to map the host's directory path to the container's directory path (i.e. -v /path/of/host/dir:/path/of/container/dir).
The path provided for /path/of/container/dir should be an absolute path, while /path/of/host/dir can be absolute or relative.
docker create \
--name exl3-container \
--gpus all \
-v /path/to/dir:/path/to/dir \
exllamav3
docker start exl3-container
Example:
docker create \
--name exl3-container \
--gpus all \
-v /mnt/llm:/mnt/llm \
exllamav3
docker start exl3-container
Note, if you need a separate path for cache (i.e. work directory), add another path for the cache directory.
docker create \
--name exl3-container \
--gpus all \
-v /path/to/dir:/path/to/dir \
-v /path/to/cache:/path/to/cache \
exllamav3
docker start exl3-container
To run the container, provide <arg> at the end. Note that it might hang for a few seconds before it outputs anything.
docker exec exl3-container python3 convert.py <arg>
Example:
docker exec exl3-container python3 convert.py \
--in_dir /mnt/llm/full_models/Llama-3.2-1B-Instruct \
--work_dir /mnt/llm/work \
--out_dir /mnt/llm/quants/Llama-3.2-1B-Instruct/4.0 \
--bits 4.0
It is possible to run scripts other than convert.py, but there may be unexpected errors and you may need to install other dependencies like requirements_eval.txt or requirements_examples.txt using docker exec.
To publish an image, username/organization plus tag will be needed. For example:
# Make sure to login with "docker login" first!
docker tag exllamav3 johndoe/exllamav3:latest
docker tag exllamav3 johndoe/exllamav3:preview-4-17-25--cuda12.8.1
docker push johndoe/exllamav3:latest
docker push johndoe/exllamav3:preview-4-17-25--cuda12.8.1
python ./docker/convert/cuda/build.py
docker create \
--name exl3-container \
--gpus all \
-v .\data:/data \ # relative host path
exllamav3
docker start exl3-container
docker exec exl3-container python3 convert.py \
--in_dir /data/full_models/Llama-3.2-1B-Instruct \
--work_dir /data/cache \
--out_dir /data/quants/Llama-3.2-1B-Instruct/4.0 \
--bits 4.0
Content type
Image
Digest
sha256:13319bc7e…
Size
8.6 GB
Last updated
over 1 year ago
docker pull owentruong/exllamav3