RPX model runtimes for image/video depth, tracking, camera pose, and VQA.
2.3K
Run reference models on RPX using task-specific Docker environments for image depth, video depth, object tracking, camera pose, and visual question answering.
Toolkit guide · Code and evaluation tools · Dataset · Browse image tags
Every image lives in irvlutd/rpx. The part after : selects the runtime.
Each task's images are built cumulatively: the final image of a task already contains every earlier model of that task. Pull that one.
| What do you want to run? | Task | Pull irvlutd/rpx:… |
|---|---|---|
| Predict depth from an image | T1 | depth-zipdepth-latest (all ten models) |
| Predict depth from a video | T2 | depth-zipdepth-latest (eight models); depth-dvd-latest for DVD |
| Track objects from a first-frame mask | T3 | tracking-dam4sam-rpx-latest (eight models) |
| Track objects from a box or text prompt | T3 | one image per configuration, see T3 below |
| Estimate relative camera pose | T4 | rcpe-monst3r-rpx-latest (all ten models) |
| Ground a question in one image or with a reference image | T5 / T6 | vqa-vllm (all twelve models) |
Always supply a task tag. docker pull irvlutd/rpx selects the old
latest capture/annotation environment, not one of these benchmark runtimes.
Every other tag is an earlier build stage, kept for provenance; use the selection above.
The images package code and model environments. Download/mount the dataset and model weights separately. No single image contains every task.
You need Linux x86-64 (linux/amd64), Docker, an NVIDIA GPU, NVIDIA Container
Toolkit, and a driver compatible with the selected image. Model memory
requirements vary; some large models need multiple GPUs. Allow disk space for
the image, checkpoints, dataset cache, and results.
For example, inspect the VQA runtime:
docker pull irvlutd/rpx:vqa-vllm
docker run --rm irvlutd/rpx:vqa-vllm help
docker run --rm irvlutd/rpx:vqa-vllm list-models
docker run --rm --gpus device=0 irvlutd/rpx:vqa-vllm verify
verify is an environment/GPU preflight, not a benchmark result. Next, follow
the task's smoke/micro inference → acceptance → full benchmark instructions
linked below. Use the same dataset revision and evaluation protocol when
comparing models.
Keep downloads and predictions on the host. For VQA, these are the persistent mount locations used by its runner:
mkdir -p "$HOME/.cache/huggingface" "$HOME/.cache/rpx-vqa" "$PWD/rpx-results"
# Set HF_TOKEN in your shell if your chosen checkpoint requires authentication.
docker run --rm --gpus device=0 --ipc=host \
-e HF_TOKEN \
-v "$HOME/.cache/huggingface:/cache/huggingface" \
-v "$HOME/.cache/rpx-vqa:/cache/rpx-vqa" \
-v "$PWD/rpx-results:/outputs" \
irvlutd/rpx:vqa-vllm help
This displays the runner's commands. Use the VQA guide for the model-specific inference command and GPU allocation; the example does not launch inference.
depth-zipdepth-latest is the final depth image. It holds the 16-model shared
runtime plus the FE2E and ZipDepth overlays; choose the model through the task runner.
| Task | Model(s) | Tag after irvlutd/rpx: |
|---|---|---|
| Image depth | DA-V2 Large, DA3 Metric-L, Depth Pro, FE2E, HyDen, Lotus-2, Metric3D V2, MoGe-2 ViT-L, UniDepth V2, ZipDepth | depth-zipdepth-latest |
| Video depth | ChronoDepth, DA3 video, DepthCrafter, MonST3R, RollingDepth, VGGT, ViGeo, Video-DA | depth-zipdepth-latest |
| Video depth | DVD v1.1 | depth-dvd-latest |
| Video depth | GemDepth | No prebuilt image; build recipe available |
docker pull irvlutd/rpx:depth-zipdepth-latest
Depth task guide · DVD · GemDepth build recipe
tracking-dam4sam-rpx-latest is the final mask-initialized tracking image and
contains all eight first-frame-mask trackers. Box- and text-prompted runs are
different benchmark inputs and use their own images.
| Initialization | Model(s) | Tag after irvlutd/rpx: |
|---|---|---|
| First-frame mask | SAM 2, EdgeTAM, Cutie, SAM2Long, SAM 2++, MiTS, XMem, DAM4SAM | tracking-dam4sam-rpx-latest |
| First-frame box | SAM 3.1 | tracking-sam3.1-bbox-rpx-latest |
| Text prompt | SAM 3.1 | tracking-sam3.1-text-rpx-latest |
| Text prompt | Grounded SAM 2 | tracking-grounded-sam2-text-rpx-latest |
docker pull irvlutd/rpx:tracking-dam4sam-rpx-latest
Tracking task guide and smoke gates
rcpe-monst3r-rpx-latest is the final pose image and contains all ten models
(VGGT, DA3, CUT3R, DUSt3R, MASt3R, MUSt3R, Reloc3R, Pi3X, Fast3R, MonST3R), each
in its own environment at /opt/rpx-envs/<model>/bin/python.
docker pull irvlutd/rpx:rcpe-monst3r-rpx-latest
Depth and pose images for DA3, MonST3R, and VGGT are task-specific. Use the
rcpe- image for camera pose, even if you already downloaded a depth image.
Camera-pose task guide and smoke gates
Start with irvlutd/rpx:vqa-vllm for normal one-image and in-context
two-image evaluation. Select the model and task in the runner; you do not need
one Docker image per checkpoint size. The runtime uses vLLM and model-specific
native backends where configured.
The reference families include Cosmos Reason2, DeepSeek-VL2, Florence-2,
InternVL 3.5, PaliGemma 2, and Qwen3-VL. Run list-models to see the exact
keys supported by your image. Scoring evaluates the grounded bounding box;
free-form answers are not a substitute for that output contract.
vqa-vqa-incontext-de59a51 and the vqa-vllm-sha-… tags preserve older builds.
Do not select a runtime merely because its name contains “incontext.”
vqa-gemma4 is an additional specialized environment, separate from the shared
reference selection above.
VQA task guide, model backends, and inference commands
| Tag pattern | Meaning |
|---|---|
depth-… | Image/video depth runtime |
tracking-… | Object-tracking runtime |
rcpe-… | Relative camera-pose runtime |
vqa-… | VQA/grounding runtime |
…-latest | A convenient model-specific alias; it may change |
…-sha-… or a revision suffix | A retained build reference; use its digest for immutable identity |
latest with no task prefix | Legacy capture/annotation environment |
For a reproducible experiment, record the image digest, checkpoint revision, dataset revision, model arguments, and GPU configuration:
docker pull irvlutd/rpx:tracking-dam4sam-rpx-latest
docker image inspect irvlutd/rpx:tracking-dam4sam-rpx-latest \
--format '{{index .RepoDigests 0}}'
# Use the printed irvlutd/rpx@sha256:... reference for subsequent runs.
Some linked source guides still show the previous registry names. Translate the full image reference, preserving its original tag:
| Previous task repository | Current image reference |
|---|---|
rpx-depth-smoke:TAG | irvlutd/rpx:depth-TAG |
rpx-tracking-smoke:TAG | irvlutd/rpx:tracking-TAG |
rpx-rcpe-smoke:TAG | irvlutd/rpx:rcpe-TAG |
rpx-vqa-smoke:TAG | irvlutd/rpx:vqa-TAG |
This mapping applies to both previous personal namespaces. If a build script constructs tags internally, changing only its repository variable is not enough: pass the complete image reference where supported or adjust its tag construction too.
On 9 October 2026, 144 task tags / 102 distinct image manifests were copied and verified against their source digests. Every copied tag was independently resolved through the public registry without a saved Docker login. Historical builds are retained for reproducibility; they are not 102 distinct models.
GEMDepth remains a prebuilt-image gap. NVS is excluded. Publication checks establish image identity and public access, not a new GPU inference pass for each model. Run the selected model's smoke gate before a full benchmark.
To evaluate your own model or extend the benchmark, start with the RPX toolkit documentation.
Content type
Image
Digest
sha256:424aecd70…
Size
10.2 GB
Last updated
1 day ago
docker pull irvlutd/rpx:vqa-vllm-sha-ee4070b0726b