LLM collaborator on Discord and Matrix. Running HuggingFace: zephyr, falcon or DialoGPT.
427
Images in this repository require the Nvidia Container Toolkit to run and were designed for Kepler GPUs. The images may run on more modern hardware, but may not take full advantage of later CUDA compute capabilities. The model(s) are loaded with 4 bit quantization via BitsandBytes and require max ~5 GB of GPU memory depending on the specific model used. Containers must also mount a 'credentials' directory which contains keys, tokens etc. for the relevant accounts. The instructions below assume Ubuntu 20.04 and a working Docker (tested with 26.0.0).
First, add the repo:
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg \
&& curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
Then update and install the toolkit:
sudo apt update
sudo apt install nvidia-container-toolkit
Finally, configure docker and restart the daemon:
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
This will make changes to /etc/docker/daemon.json, adding the runtimes stanza:
{
"runtimes": {
"nvidia": {
"args": [],
"path": "nvidia-container-runtime"
}
}
}
Test with:
$ sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
Unable to find image 'ubuntu:latest' locally
latest: Pulling from library/ubuntu
bccd10f490ab: Pull complete
Digest: sha256:77906da86b60585ce12215807090eb327e7386c8fafb5402369e421f44eff17e
Status: Downloaded newer image for ubuntu:latest
Tue Mar 12 21:50:44 2024
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 470.42.01 Driver Version: 470.42.01 CUDA Version: 11.4 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|===============================+======================+======================|
| 0 Tesla K80 On | 00000000:05:00.0 Off | 0 |
| N/A 28C P8 26W / 149W | 0MiB / 11441MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
| 1 Tesla K80 On | 00000000:06:00.0 Off | 0 |
| N/A 33C P8 29W / 149W | 0MiB / 11441MiB | 0% Default |
| | | N/A |
+-------------------------------+----------------------+----------------------+
+-----------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=============================================================================|
| No running processes found |
+-----------------------------------------------------------------------------+
Make a directory for your credentials somewhere safe that the container can mount. This directory needs to contain the files described in the credentials directory of the project GitHub repo.
Pull the image:
docker pull gperdrizet/bartleby:backdrop_launch
Run the image. Replace <CREDENTIALS> with the path to your credentials directory:
docker run --gpus all --mount type=bind,source=<CREDENTIALS>,target=/bartleby/bartleby/credentials --name bartleby -d gperdrizet/bartleby:backdrop_launch
That's it! The first response from bartleby may be slow because the model(s) are pulled from HuggingFace the first time they are used. Any models used are stored persistently via a Docker volume and so will be faster to load the second time.
The container will default to using GPU 0 - if you would like to change this on a multi-GPU system you can do so by setting the 'CUDA_VISIBLE_DEVICES' environment variable at container runtime:
docker run --gpus all --mount type=bind,source=<CREDENTIALS>,target=/bartleby/bartleby/credentials -e CUDA_VISIBLE_DEVICES=1 --name bartleby -d gperdrizet/bartleby:backdrop_launch
Content type
Image
Digest
sha256:11fd111d5…
Size
5 GB
Last updated
over 2 years ago
docker pull gperdrizet/bartleby