Author: 魏文应 Date: 2024-01-11
本容器安装了miniconda,同时要求你的服务器配置有GPU。
我的物理机配置如下:
| 条目 | 配置 | 备注 |
|---|---|---|
| CPU | x86-64架构 | |
| GPU | A100 | NVIDIA-SMI 530.30.02 Driver Version: 530.30.02 CUDA Version: 12.1 |
| 基础镜像 | nvidia/cuda:12.1.0-cudnn8-devel-ubuntu20.04 |
启动容器:
container_name=miniconda # 自定义容器名称
commit_image=nuvic/miniconda:24.4-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
# --gpus '"device=0,1"'
docker run --network host --gpus all -it --name ${container_name} ${commit_image} bash
执行上面命令后,此时已经进入了容器,在容器执行 nvidia-smi,显示GPU正常
# nvidia-smi
Fri Apr 12 17:30:55 2024
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 530.30.02 Driver Version: 530.30.02 CUDA Version: 12.1 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|===============================+======================+======================|
| 0 A100 80GB PCIe Off | 00000000:4F:00.0 Off | 0 |
| N/A 33C P0 44W / 300W | 9MiB / 81251MiB | 0% Default |
| | | Disabled |
+-------------------------------+----------------------+----------------------+
| 1 A100 80GB PCIe Off | 00000000:52:00.0 Off | 0 |
| N/A 34C P0 45W / 300W | 9MiB / 81251MiB | 0% Default |
| | | Disabled |
+-------------------------------+----------------------+----------------------+
执行 conda -V命令,查看conda版本:
conda 24.4
至此表明容器启动正常。
本章节给出本镜像是如何构建的:
container_name=miniconda # 自定义容器名称
image_base=nvidia/cuda:12.1.1-cudnn8-devel-ubuntu20.04
commit_image=nuvic/miniconda:24.4-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -it --name ${container_name} ${image_base} bash
然后在容器内执行:
apt update
apt install -y vim wget git curl libgl1 libglib2.0-0 # 安装过程需要手动选择时区
apt install -y iputils-ping tsocks
mkdir -p /root/workspace && cd /root/workspace
vim /etc/tsocks.conf # 代开科学上网代理工具配置文件
在 /etc/tsocks.conf 添加相应内容(根据你的代理工具而定):
# 如果提示下面错误,这是子网掩码设置错误导致的,255.255.255.0改为255.255.0.0即可:
# libtsocks(1845): SOCKS server 192.168.10.67 (192.168.10.67) is not on a local subnet!
# 子网掩码
local = 192.168.0.0/255.255.0.0
local = 10.0.0.0/255.0.0.0
path {
reaches = 150.0.0.0/255.255.0.0
reaches = 150.1.0.0:80/255.255.0.0
server = 10.1.7.25
server_type = 5
default_user = delius
default_pass = hello
}
# 你的代理服务器的ip地址
server = 192.168.10.232
# socks5协议
server_type = 5
# 代理服务器的端口
server_port = 10808
然后执行下面命令:
# 有内容返回,则说明代理工具正常
tsocks curl www.google.com
然后安装miniconda:
# 下载安装miniconda
wget -c https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
# 退出容器
exit
上述操作退出容器后,重新登录容器:
docker start ${container_name}
docker exec -it ${container_name} bash
然后再容器内执行:
conda -V
显示类似如下,则正常:
conda 24.4.0
然后删除安装包,并退出容器:
apt autoclean && apt clean && apt autoremove
rm /root/workspace/Miniconda3-latest-Linux-x86_64.sh && exit
将其提交为镜像,并上传:
# docker commit -m "说明" 容器名称 镜像名称
docker commit ${container_name} ${commit_image}
# docker login
docker push ${commit_image}
# 删除容器
docker stop ${container_name}
docker rm ${container_name}
在启动命令中,使用 python 和 pip 命令,会抛出如下异常:
sh: 1: pip: not found
sh: 1: python: not found
因此,我们需要将python路径,添加到系统环境变量中,首先你需要创建 Dockerfile 文件,写入如下内容:
# 你的镜像
from nuvic/miniconda:24.4-cuda12.1.1-ubuntu20.04
# 容器内conda安装路径
ENV PATH="/root/miniconda3/bin:/root/miniconda3/condabin:$PATH"
然后在Dockerfile相同目录下,执行如下命令,创建新的镜像:
docker build -t nuvic/miniconda:24.4-cuda12.1.1-ubuntu20.04-env ./
人工智能智算平台,一般通过jupyter启动notebook代码调试窗口,以此python依赖中需要安装jupyter:
docker run --network host --gpus all -it --name miniconda-k8s nuvic/miniconda:24.4-cuda12.1.1-ubuntu20.04-env bash
在容器内执行:
conda install jupyterlab=4.2.5
# 成功执行会显示 successfully
jupyter-lab --no-browser
# 清除缓存
apt autoclean && apt clean && apt autoremove
pip cache purge && conda clean -y --all
exit
提交镜像:
# docker commit -m "说明" 容器名称 镜像名称
docker commit miniconda-k8s nuvic/miniconda:cuda12.1-ubuntu20.04-jupyterlab4.2.5
我的物理机配置如下:
| 条目 | 配置 | 备注 |
|---|---|---|
| CPU | x86-64架构 | |
| GPU | A100 | NVIDIA-SMI 460.106.00 Driver Version: 460.106.00 CUDA Version: 11.2 |
| 基础镜像 | nvidia/cuda:11.2.2-cudnn8-devel-ubuntu20.04 |
启动容器:
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
# docker run --network host --gpus '"device=0,1"' -it --name miniconda nuvic/miniconda:24.1.2-cuda11.2.2-ubuntu20.04 bash
docker run --network host --gpus all -it --name miniconda nuvic/miniconda:24.1.2-cuda11.2.2-ubuntu20.04 bash
执行上面命令后,此时已经进入了容器,在容器执行 nvidia-smi,显示GPU正常
# nvidia-smi
Fri Apr 12 17:30:55 2024
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 460.106.00 Driver Version: 460.106.00 CUDA Version: 11.2 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|===============================+======================+======================|
| 0 A100 80GB PCIe Off | 00000000:4F:00.0 Off | 0 |
| N/A 33C P0 44W / 300W | 9MiB / 81251MiB | 0% Default |
| | | Disabled |
+-------------------------------+----------------------+----------------------+
| 1 A100 80GB PCIe Off | 00000000:52:00.0 Off | 0 |
| N/A 34C P0 45W / 300W | 9MiB / 81251MiB | 0% Default |
| | | Disabled |
+-------------------------------+----------------------+----------------------+
执行 conda -V命令,查看conda版本:
conda 24.1.2
至此表明容器启动正常。
本章节给出本镜像是如何构建的:
container_name=conda # 自定义容器名称
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus '"device=2,3"' -it --name ${container_name} nvidia/cuda:11.2.2-cudnn8-devel-ubuntu20.04 bash
然后在容器内执行:
apt update
apt install -y vim wget git curl libgl1 libglib2.0-0 # 安装过程需要手动选择时区
apt install -y iputils-ping tsocks
mkdir -p /root/workspace && cd /root/workspace
vim /etc/tsocks.conf # 代开科学上网代理工具配置文件
在 /etc/tsocks.conf 添加相应内容(根据你的代理工具而定):
# 如果提示下面错误,这是子网掩码设置错误导致的,255.255.255.0改为255.255.0.0即可:
# libtsocks(1845): SOCKS server 192.168.10.67 (192.168.10.67) is not on a local subnet!
# 子网掩码
local = 192.168.0.0/255.255.0.0
local = 10.0.0.0/255.0.0.0
path {
reaches = 150.0.0.0/255.255.0.0
reaches = 150.1.0.0:80/255.255.0.0
server = 10.1.7.25
server_type = 5
default_user = delius
default_pass = hello
}
# 你的代理服务器的ip地址
server = 192.168.10.232
# socks5协议
server_type = 5
# 代理服务器的端口
server_port = 10808
然后执行下面命令:
# 有内容返回,则说明代理工具正常
tsocks curl www.google.com
然后安装miniconda:
# 下载安装miniconda
wget -c https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
# 退出容器
exit
上述操作退出容器后,重新登录容器:
docker start ${container_name}
docker exec -it ${container_name} bash
然后再容器内执行:
conda -V
显示类似如下,则正常:
conda 24.1.2
然后删除安装包,并退出容器:
apt autoclean && apt clean && apt autoremove
rm /root/workspace/Miniconda3-latest-Linux-x86_64.sh && exit
将其提交为镜像,并上传:
# docker commit -m "说明" 容器名称 镜像名称
docker commit -m 'miniconda is available.' conda nuvic/miniconda:24.1.2-cuda11.2.2-ubuntu20.04
# docker login
docker push nuvic/miniconda:24.1.2-cuda11.2.2-ubuntu20.04
Content type
Image
Digest
sha256:3ed028bff…
Size
5.5 GB
Last updated
almost 2 years ago
docker pull nuvic/miniconda:cuda12.1-ubuntu20.04-jupyterlab4.2.5