Blackwell-first LTX 2.5 inference on ComfyUI and RunPod, with models, caches and Python environment.
2.6K
Blackwell-first LTX 2.5 inference on ComfyUI and RunPod, with persistent models, caches, ComfyUI, and Python environment under /workspace.
Set:
PERSIST_WORKSPACE=true
RUN_MODE=worker
LTX25_PRELOAD_VARIANT=distilled-int8
LTX25_PRELOAD_PROMPT_ENHANCER=true
HUGGINGFACE_ACCESS_TOKEN=hf_xxx
Send the checked-in LTX 2.5 I2V API workflow to /run or /runsync using the workflow request contract.
The model bootstrap downloads weights to the persistent model root. Model weights are never baked into Docker layers.
| Target | CUDA | Model profile |
|---|---|---|
base | 13.0.2 | No startup preload |
base-cuda12-8-1 | 12.8.1 | No startup preload |
ltx2-5-distilled-int8-cu130 | 13.0.2 | Distilled INT8 ConvRot, recommended |
ltx2-5-distilled-int8-cu128 | 12.8.1 | Distilled INT8 ConvRot fallback |
worker: ComfyUI, bundled frontend, and RunPod serverless handlerlocal-api: ComfyUI, frontend, and local RunPod-compatible API on port 8000pod: ComfyUI and frontend without the serverless handlerThe frontend is served on port 7777; ComfyUI uses 8188.
The preferred request shape is below. workflow must be an API-format ComfyUI workflow; {} is only a structural placeholder.
{
"input": {
"workflow": {},
"images": [
{
"name": "source.png",
"image": "data:image/png;base64,..."
}
]
}
}
The worker returns output.images[] and/or output.videos[]. Artifacts are inline base64 unless S3 is configured. The older input.prompt + input.image_url request remains available for existing clients.
The default INT8 workflow preloads:
models/diffusion_models/ltx-2.5-22b-distilled-transformer-comfy-int8-convrot.safetensorsmodels/text_encoders/gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensorsmodels/text_encoders/gemma4_e2b_it_bf16.safetensorsmodels/vae/ltx-2.5-video-vae-bf16.safetensorsmodels/vae/ltx-2.5-audio-vae-bf16.safetensorsmodels/latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensorsStart with at least 48 GB VRAM for the distilled INT8 profile and validate the exact workflow before production. Attach at least 100 GB of persistent storage for the default stack and caches.
The test workflow validates shell scripts, JSON, Python syntax, workflow transformation, payload handling, and Docker bake definitions. A release is not GPU-validated until the image has been built for linux/amd64, booted on the target CUDA/GPU class, and completed the checked-in workflow.
Content type
Image
Digest
sha256:73d1621ef…
Size
3.6 GB
Last updated
about 2 months ago
docker pull notrius/ltx-2.5-serverless