stable-diffusion-webui基础镜像。
273
Author: 魏文应 Date: 2024-04-15
AUTOMATIC1111/stable-diffusion-webui 项目在NVIDIA GPU设备上的基本镜像。
我的物理机配置如下:
| 条目 | 配置 | 备注 |
|---|---|---|
| CPU | x86-64架构 | |
| GPU | A100 | NVIDIA-SMI 530.30.02 Driver Version: 530.30.02 CUDA Version: 12.1 |
| OS | ubuntu20.04 | |
| 镜像 | stable diffusion webui | pytorch2.2.0,python3.10, cuda12.1.0, ubuntu20.04 |
启动容器:
container_name=sdwebdui
commit_image=nuvic/sdwebui:v1.9.0-pytorch2.2.0-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -it --name ${container_name} ${commit_image} bash
执行上面命令后,此时已经进入了容器,在容器执行:
# export CUDA_VISIBLE_DEVICES=3 && python launch.py --skip-version-check --xformers --share
# --share是生成一个临时公网链接
# python launch.py --skip-version-check --xformers --share
cd /root/workspace/stable-diffusion-webui/
# --listen是允许局域网,通过IP:端口方式访问服务
python launch.py --skip-version-check --xformers --listen
创建容器:
container_name=sdwebdui
image_base=nuvic/pytorch:2.2.0-python3.10-cuda12.1.1-ubuntu20.04
commit_image=nuvic/sdwebui:v1.9.0-pytorch2.2.0-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -it --name ${container_name} ${image_base} bash
在容器内下载 stable-diffusion-webui 项目:
cd /root/workspace
git clone https://github.com/AUTOMATIC1111/stable-diffusion-webui.git
cd stable-diffusion-webui/
git checkout v1.9.0 -b runtime
然后 vim requirements_versions.txt 打开文件,删除里面的 torch ,然后执行下面命令安装Python依赖:
# 如果抛出如下异常:
# setuptools.installer and fetch_build_eggs are deprecated.
# 请使用下面命令解决:
# pip install --use-pep517 --verbose basicsr==1.4.2
pip install basicsr==1.4.2
# https://github.com/facebookresearch/xformers
# xformers的版本要求比较严格,安装时要注意cuda相关库是否在下载,如果在下载,说明版本和当前Pytorch版本不匹配
conda install xformers -c xformers
pip install httpx[socks] # 代理工具需要
pip install -r requirements_versions.txt
执行下面命令,配置一下gradio:
# 这里配置你的科学上网代理工具
vim /etc/tsocks.conf
# 配置完成后,下载一个文件
tsocks wget https://cdn-media.huggingface.co/frpc-gradio-0.2/frpc_linux_amd64 && \
mv frpc_linux_amd64 frpc_linux_amd64_v0.2 && \
chmod +x frpc_linux_amd64_v0.2 && \
mv frpc_linux_amd64_v0.2 /root/miniconda3/lib/python3.10/site-packages/gradio/
然后清除缓存,并退出容器:
pip cache purge && conda clean -y --all && exit
并将容器提交为镜像:
# docker commit -m "说明" 容器名称 镜像名称
docker commit ${container_name} ${commit_image}
# docker login
docker push ${commit_image}
# 删除容器
docker stop ${container_name}
docker rm ${container_name}
我的物理机配置如下:
| 条目 | 配置 | 备注 |
|---|---|---|
| CPU | x86-64架构 | |
| GPU | A100 | NVIDIA-SMI 460.106.00 Driver Version: 460.106.00 CUDA Version: 11.2 |
| OS | ubuntu20.04 | |
| 镜像 | stable diffusion webui | pytorch2.2.0,python3.10, cuda11.2.2, ubuntu20.04 |
启动容器:
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus '"device=0,1"' -it --name sdwebui nuvic/sdwebui:v1.9.0-pytorch2.2.0-ubuntu20.04 bash
执行上面命令后,此时已经进入了容器,在容器执行:
# export CUDA_VISIBLE_DEVICES=3 && python launch.py --skip-version-check --xformers --share
# --share是生成一个临时公网链接
# python launch.py --skip-version-check --xformers --share
# --listen是允许局域网,通过IP:端口方式访问服务
python launch.py --skip-version-check --xformers --listen
创建容器:
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus '"device=0,1"' -it --name webui nuvic/pytorch:2.2.0-python3.10-cuda11.2.2-ubuntu20.04 bash
在容器内下载 stable-diffusion-webui 项目:
cd /root/workspace
git clone https://github.com/AUTOMATIC1111/stable-diffusion-webui.git
cd stable-diffusion-webui/
git checkout v1.9.0 -b runtime
然后 vim requirements_versions.txt 打开文件,删除里面的 torch ,然后执行下面命令安装Python依赖:
# 如果抛出如下异常:
# setuptools.installer and fetch_build_eggs are deprecated.
# 请使用下面命令解决:
# pip install --use-pep517 --verbose basicsr==1.4.2
pip install basicsr==1.4.2
# https://github.com/facebookresearch/xformers
# xformers的版本要求比较严格,安装时要注意cuda相关库是否在下载,如果在下载,说明版本和当前Pytorch版本不匹配
conda install xformers -c xformers
pip install httpx[socks] # 代理工具需要
pip install -r requirements_versions.txt
执行下面命令,配置一下gradio:
# 这里配置你的科学上网代理工具
vim /etc/tsocks.conf
# 配置完成后,下载一个文件
tsocks wget https://cdn-media.huggingface.co/frpc-gradio-0.2/frpc_linux_amd64 && \
mv frpc_linux_amd64 frpc_linux_amd64_v0.2 && \
chmod +x frpc_linux_amd64_v0.2 && \
mv frpc_linux_amd64_v0.2 /root/miniconda3/lib/python3.10/site-packages/gradio/
然后清除缓存,并退出容器:
pip cache purge && conda clean -y --all && exit
并将容器提交为镜像:
# docker commit 容器名称 镜像名称
docker commit webui nuvic/sdwebui:v1.9.0-pytorch2.2.0-ubuntu20.04
# docker login
docker push nuvic/sdwebui:v1.9.0-pytorch2.2.0-ubuntu20.04
我的物理机配置如下:
| 条目 | 配置 | 备注 |
|---|---|---|
| CPU | x86-64架构 | |
| GPU | A100 | NVIDIA-SMI 530.30.02 Driver Version: 530.30.02 CUDA Version: 12.1 |
| OS | ubuntu20.04 | |
| 镜像 | stable diffusion webui | pytorch2.2.0,python3.10, cuda12.1.0, ubuntu20.04 |
启动容器:
container_name=loraui
commit_image=nuvic/sdwebui:lora-pytorch2.2.0-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -v huggingface:/root/.cache/huggingface -it --name ${container_name} ${commit_image} bash
执行上面命令后,此时已经进入了容器,在容器执行:
# export CUDA_VISIBLE_DEVICES=3 && python launch.py --skip-version-check --xformers --share
# --share是生成一个临时公网链接
# python launch.py --skip-version-check --xformers --share
cd /root/workspace/stable-diffusion-webui/
export TF_ENABLE_ONEDNN_OPTS=0
# --listen是允许局域网,通过IP:端口方式访问服务
python launch.py --skip-version-check --xformers --listen --port 12345
下载自己希望的底模:
# 可以在civitai上直接下载,但推荐huggingface比较快
# https://civitai.com/models/139562/realvisxl-v40
# huggingface-cli download 项目名称 模型文件名称
huggingface-cli download SG161222/RealVisXL_V4.0 RealVisXL_V4.0.safetensors
ln -s /root/.cache/huggingface/hub/models--SG161222--RealVisXL_V4.0/snapshots/49740684ab2d8f4f5dcf6c644df2b33388a8ba85/RealVisXL_V4.0.safetensors \
/root/workspace/stable-diffusion-webui/models/Stable-diffusion/RealVisXL_V4.0.safetensors
创建容器:
container_name=lora
image_base=nuvic/pytorch:2.2.0-python3.10-cuda12.1.1-ubuntu20.04
commit_image=nuvic/sdwebui:lora-pytorch2.2.0-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -v 数据卷:容器挂载路径 -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -v huggingface:/root/.cache/huggingface -it --name ${container_name} ${image_base} bash
在容器内下载 stable-diffusion-webui 项目:
cd /root/workspace
git clone https://github.com/AUTOMATIC1111/stable-diffusion-webui.git
cd stable-diffusion-webui/
git checkout v1.9.0 -b runtime
然后 vim requirements_versions.txt 打开文件,删除里面的 torch ,然后执行下面命令安装Python依赖:
# 如果抛出如下异常:
# setuptools.installer and fetch_build_eggs are deprecated.
# 请使用下面命令解决:
# pip install --use-pep517 --verbose basicsr==1.4.2
pip install basicsr==1.4.2
# https://github.com/facebookresearch/xformers
# xformers的版本要求比较严格,安装时要注意cuda相关库是否在下载,如果在下载,说明版本和当前Pytorch版本不匹配
conda install xformers -c xformers
pip install httpx[socks] # 代理工具需要
pip install -r requirements_versions.txt
执行下面命令,配置一下gradio:
# 这里配置你的科学上网代理工具
# vim /etc/tsocks.conf
# 配置完成后,下载一个文件
wget https://cdn-media.huggingface.co/frpc-gradio-0.2/frpc_linux_amd64 && \
mv frpc_linux_amd64 frpc_linux_amd64_v0.2 && \
chmod +x frpc_linux_amd64_v0.2 && \
mv frpc_linux_amd64_v0.2 /root/miniconda3/lib/python3.10/site-packages/gradio/
先下载默认的模型到数据卷中:
# huggingface-cli download 项目名称 模型文件名称
huggingface-cli download runwayml/stable-diffusion-v1-5 v1-5-pruned-emaonly.safetensors
ln -s /root/.cache/huggingface/hub/models--runwayml--stable-diffusion-v1-5/blobs/6ce0161689b3853acaa03779ec93eafe75a02f4ced659bee03f50797806fa2fa \
/root/workspace/stable-diffusion-webui/models/Stable-diffusion/v1-5-pruned-emaonly.safetensors
然后启动webui:
python launch.py --skip-version-check --xformers --listen --port 12345
在浏览器中打开如下链接:
# http://你的服务器IP:对应端口
http://192.168.10.144:12345
然后到 Extensions -> Available 中,安装简体中文插件 zh_Hans Localization 、用于数据标注的插件 WD 1.4 Tagger :

插件安装后,到 Settings -> User Interface -> User interface 中的Localization选择zh-Hans(Stable),点击Apply settings,最后到 Extensions ->Installed 页面点击 Apply and quit 应用设置并退出:

启动webui,安装汉化插件和数据标注插件:
python launch.py --skip-version-check --xformers --listen --port 12345
然后执行下面命令:
# 容器内执行
mkdir -p /root/testdata/william
# 物理机内执行
# 更好的数据集准备参考:https://www.bilibili.com/video/BV1Z841137ft/?vd_source=3864c9f21aada39c764f51eff1ef53c6
# docker cp 测试图片数据集 lora:容器内目录
docker cp data/80T/lora训练数据集/pick/ lora:/root/testdata/william
然后在容器内处理一下图片文件名称:
# 安装处理工具
pip install aigcfile
# 处理
aigcfile rename -s /root/testdata/william/pick
接着使用 webui 中的WD1.4标签器,输入路径 /root/testdata/william/pick ,点击 反推:

这里为了测试,我为每个标签文件添加类似的内容:
# 写入内容
aigcfile write -s /root/testdata/william/pick --string "weiwenying, chinese man, chinese william, china, 30 year old, middle-aged man, "
然后清除缓存,并退出容器:
pip cache purge && conda clean -y --all && exit
并将容器提交为镜像:
# docker commit -m "说明" 容器名称 镜像名称
docker commit ${container_name} ${commit_image}
# docker login
docker push ${commit_image}
# 删除容器
docker stop ${container_name}
docker rm ${container_name}
启动容器:
container_name=kohya
commit_image=nuvic/sdwebui:kohya-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -v huggingface:/root/.cache/huggingface -it --shm-size=1g --ulimit memlock=-1 --name ${container_name} ${commit_image} bash
执行上面命令后,此时已经进入了容器,在容器执行:
# export CUDA_VISIBLE_DEVICES=3 && python launch.py --skip-version-check --xformers --share
# --share是生成一个临时公网链接
# python launch.py --skip-version-check --xformers --share
# export TF_ENABLE_ONEDNN_OPTS=0
# --listen是允许局域网,通过IP:端口方式访问服务
# 1.数据标注; 3.模型生成
cd /root/workspace/stable-diffusion-webui/
python launch.py --skip-version-check --xformers --listen --port 12345
# 2. lora训练
cd /root/workspace/kohya_ss
python kohya_gui.py --listen 0.0.0.0 --server_port 12346
启动基础镜像:
# 下面这个webui内的插件太老了,不推荐使用,应该要将训练环境分开。
# kohya_ss LoRA extension for webui: https://github.com/kohya-ss/sd-webui-additional-networks
container_name=kohya
image_base=nuvic/sdwebui:lora-pytorch2.2.0-cuda12.1.1-ubuntu20.04
commit_image=nuvic/sdwebui:kohya-cuda12.1.1-ubuntu20.04
# docker run --gpus 指定物理GPU -it --name 自定义容器名称 镜像名称:版本(相应版本要求物理机nvidia driver支持) bash
docker run --network host --gpus all -v huggingface:/root/.cache/huggingface -it --shm-size=1g --ulimit memlock=-1 --name ${container_name} ${image_base} bash
在容器内执行下面步骤:
cd /root/workspace/
git clone --recursive https://github.com/bmaltais/kohya_ss.git
cd /root/workspace/kohya_ss
conda create -n kohya python=3.10
conda activate kohya
conda install pytorch==2.2.0 torchvision==0.17.0 torchaudio==2.2.0 pytorch-cuda=12.1 -c pytorch -c nvidia
conda install xformers -c xformers
pip install bitsandbytes
# pip install -e ./sd-scripts
# pip install gradio toml easygui psutil transformer accelerate diffusers einops imagesize opencv-python
# pip install voluptuous scipy wandb timm pytorch-lightning prodigyopt open-clip-torch protobuf omegaconf lycoris_lora
打开 requirements.txt ,将版本号去掉:
accelerate
aiofiles
altair
dadaptation
diffusers
easygui
einops
fairscale
ftfy
gradio
huggingface-hub
imagesize
invisible-watermark
lion-pytorch
lycoris_lora
omegaconf
onnx
prodigyopt
protobuf
open-clip-torch
opencv-python
prodigyopt
pytorch-lightning
rich
safetensors
scipy
timm
tk
toml
transformers
voluptuous
wandb
scipy
# for kohya_ss library
-e ./sd-scripts
执行安装依赖:
pip install -r requirements.txt
环境依赖安装完成后,执行下面操作:
# 准备模型
rm -rf /root/workspace/kohya_ss/models && ln -s /root/workspace/stable-diffusion-webui/models/Stable-diffusion models
# 准备一个测试集, 数据集的文件夹得有下划线
mv /root/testdata/william/pick/ /root/testdata/william/3_wil
ln -s /root/testdata/william/ /root/workspace/kohya_ss/datasets
# huggingface-cli download 项目名称 模型文件名称
huggingface-cli download SG161222/RealVisXL_V4.0 RealVisXL_V4.0.safetensors
ln -s /root/.cache/huggingface/hub/models--SG161222--RealVisXL_V4.0/snapshots/49740684ab2d8f4f5dcf6c644df2b33388a8ba85/RealVisXL_V4.0.safetensors \
/root/workspace/stable-diffusion-webui/models/Stable-diffusion/RealVisXL_V4.0.safetensors
conda activate kohya
python kohya_gui.py --listen 0.0.0.0 --server_port 12346
然后在浏览器打开:
http://192.168.10.144:12346/
先在 Accelerate lauch 里,设置 Number of processes(8表示使用8个GPU),以及启动 Multi GPU, Distributed GPUs:

然后选择模型、选择数据集、选择输出目录:

最后拉到底部,点击start training即可开启训练。训练完成后,打开stable diffusion webui:
# 替换一下webui的Lora路径
rm -rf /root/workspace/stable-diffusion-webui/models/Lora && ln -s /root/workspace/kohya_ss/outputs/ /root/workspace/stable-diffusion-webui/models/Lora
conda activate bash
cd /root/workspace/kohya_ss
python kohya_gui.py --listen 0.0.0.0 --server_port 12346

点击生成后,在生成栏显示如下:
weiwenying, chinese man, chinese william, china, 30 year old, middle-aged man, william was standing in the street. <lora:last:1>

Tips: docker 容器启动,没有设置
--shm-size=1g --ulimit memlock=-1的时候,当GPU并行开启超过4个,就会抛出如下异常:torch.distributed.DistBackendError: NCCL error in: /opt/conda/conda-bld/pytorch_1704987394225/work/torch/csrc/distributed/c10d/ProcessGroupNCCL.cpp:1691, unhandled system error (run with NCCL_DEBUG=INFO for details), NCCL version 2.19.3共享内存耗尽导致的,因此docker启动时,将
--shm-size设置大一些即可,默认只有64MB.
确认无误后,清理容器:
# 下面目录的文件,适当清理一下, 我只留下一个最后的文件last.safetensors用于方便测试,其它删除
# /root/workspace/stable-diffusion-webui/models/Lora
# config_lora-20240702-222539.toml last.safetensors last_20240702-222539.json
# 图像生成输出目录,也清理一下
rm -rf /root/workspace/stable-diffusion-webui/outputs/*
然后清除缓存,并退出容器:
pip cache purge && conda clean -y --all && exit
并将容器提交为镜像:
# docker commit -m "说明" 容器名称 镜像名称
docker commit ${container_name} ${commit_image}
# docker login
docker push ${commit_image}
# 删除容器
docker stop ${container_name}
docker rm ${container_name}
Content type
Image
Digest
sha256:9edea8c8b…
Size
9.6 GB
Last updated
almost 2 years ago
docker pull nuvic/sdwebui:v1.0-cuda12.1-base