Sign inSign up

irvlutd/rpx

By irvlutd

•Updated 1 day ago

RPX model runtimes for image/video depth, tracking, camera pose, and VQA.

Image
0

2.3K

irvlutd/rpx repository overview

⁠RPX — benchmark model environments

Run reference models on RPX using task-specific Docker environments for image depth, video depth, object tracking, camera pose, and visual question answering.

Toolkit guide⁠ · Code and evaluation tools⁠ · Dataset⁠ · Browse image tags⁠

⁠Start here: choose your task

Every image lives in irvlutd/rpx. The part after : selects the runtime.

Each task's images are built cumulatively: the final image of a task already contains every earlier model of that task. Pull that one.

What do you want to run?TaskPull irvlutd/rpx:…
Predict depth from an imageT1depth-zipdepth-latest (all ten models)
Predict depth from a videoT2depth-zipdepth-latest (eight models); depth-dvd-latest for DVD
Track objects from a first-frame maskT3tracking-dam4sam-rpx-latest (eight models)
Track objects from a box or text promptT3one image per configuration, see T3 below
Estimate relative camera poseT4rcpe-monst3r-rpx-latest (all ten models)
Ground a question in one image or with a reference imageT5 / T6vqa-vllm (all twelve models)

Always supply a task tag. docker pull irvlutd/rpx selects the old latest capture/annotation environment, not one of these benchmark runtimes. Every other tag is an earlier build stage, kept for provenance; use the selection above.

The images package code and model environments. Download/mount the dataset and model weights separately. No single image contains every task.

⁠Quick start

You need Linux x86-64 (linux/amd64), Docker, an NVIDIA GPU, NVIDIA Container Toolkit, and a driver compatible with the selected image. Model memory requirements vary; some large models need multiple GPUs. Allow disk space for the image, checkpoints, dataset cache, and results.

For example, inspect the VQA runtime:

docker pull irvlutd/rpx:vqa-vllm
docker run --rm irvlutd/rpx:vqa-vllm help
docker run --rm irvlutd/rpx:vqa-vllm list-models
docker run --rm --gpus device=0 irvlutd/rpx:vqa-vllm verify

verify is an environment/GPU preflight, not a benchmark result. Next, follow the task's smoke/micro inference → acceptance → full benchmark instructions linked below. Use the same dataset revision and evaluation protocol when comparing models.

Keep downloads and predictions on the host. For VQA, these are the persistent mount locations used by its runner:

mkdir -p "$HOME/.cache/huggingface" "$HOME/.cache/rpx-vqa" "$PWD/rpx-results"

# Set HF_TOKEN in your shell if your chosen checkpoint requires authentication.
docker run --rm --gpus device=0 --ipc=host \
  -e HF_TOKEN \
  -v "$HOME/.cache/huggingface:/cache/huggingface" \
  -v "$HOME/.cache/rpx-vqa:/cache/rpx-vqa" \
  -v "$PWD/rpx-results:/outputs" \
  irvlutd/rpx:vqa-vllm help

This displays the runner's commands. Use the VQA guide for the model-specific inference command and GPU allocation; the example does not launch inference.

⁠T1 / T2 — image and video depth

depth-zipdepth-latest is the final depth image. It holds the 16-model shared runtime plus the FE2E and ZipDepth overlays; choose the model through the task runner.

TaskModel(s)Tag after irvlutd/rpx:
Image depthDA-V2 Large, DA3 Metric-L, Depth Pro, FE2E, HyDen, Lotus-2, Metric3D V2, MoGe-2 ViT-L, UniDepth V2, ZipDepthdepth-zipdepth-latest
Video depthChronoDepth, DA3 video, DepthCrafter, MonST3R, RollingDepth, VGGT, ViGeo, Video-DAdepth-zipdepth-latest
Video depthDVD v1.1depth-dvd-latest
Video depthGemDepthNo prebuilt image; build recipe available
docker pull irvlutd/rpx:depth-zipdepth-latest

Depth task guide⁠ · DVD⁠ · GemDepth build recipe⁠

⁠T3 — object tracking

tracking-dam4sam-rpx-latest is the final mask-initialized tracking image and contains all eight first-frame-mask trackers. Box- and text-prompted runs are different benchmark inputs and use their own images.

InitializationModel(s)Tag after irvlutd/rpx:
First-frame maskSAM 2, EdgeTAM, Cutie, SAM2Long, SAM 2++, MiTS, XMem, DAM4SAMtracking-dam4sam-rpx-latest
First-frame boxSAM 3.1tracking-sam3.1-bbox-rpx-latest
Text promptSAM 3.1tracking-sam3.1-text-rpx-latest
Text promptGrounded SAM 2tracking-grounded-sam2-text-rpx-latest
docker pull irvlutd/rpx:tracking-dam4sam-rpx-latest

Tracking task guide and smoke gates⁠

⁠T4 — relative camera pose

rcpe-monst3r-rpx-latest is the final pose image and contains all ten models (VGGT, DA3, CUT3R, DUSt3R, MASt3R, MUSt3R, Reloc3R, Pi3X, Fast3R, MonST3R), each in its own environment at /opt/rpx-envs/<model>/bin/python.

docker pull irvlutd/rpx:rcpe-monst3r-rpx-latest

Depth and pose images for DA3, MonST3R, and VGGT are task-specific. Use the rcpe- image for camera pose, even if you already downloaded a depth image.

Camera-pose task guide and smoke gates⁠

⁠T5 / T6 — VQA and visual grounding

Start with irvlutd/rpx:vqa-vllm for normal one-image and in-context two-image evaluation. Select the model and task in the runner; you do not need one Docker image per checkpoint size. The runtime uses vLLM and model-specific native backends where configured.

The reference families include Cosmos Reason2, DeepSeek-VL2, Florence-2, InternVL 3.5, PaliGemma 2, and Qwen3-VL. Run list-models to see the exact keys supported by your image. Scoring evaluates the grounded bounding box; free-form answers are not a substitute for that output contract.

vqa-vqa-incontext-de59a51 and the vqa-vllm-sha-… tags preserve older builds. Do not select a runtime merely because its name contains “incontext.” vqa-gemma4 is an additional specialized environment, separate from the shared reference selection above.

VQA task guide, model backends, and inference commands⁠

⁠Understanding tags and reproducing a run

Tag patternMeaning
depth-…Image/video depth runtime
tracking-…Object-tracking runtime
rcpe-…Relative camera-pose runtime
vqa-…VQA/grounding runtime
…-latestA convenient model-specific alias; it may change
…-sha-… or a revision suffixA retained build reference; use its digest for immutable identity
latest with no task prefixLegacy capture/annotation environment

For a reproducible experiment, record the image digest, checkpoint revision, dataset revision, model arguments, and GPU configuration:

docker pull irvlutd/rpx:tracking-dam4sam-rpx-latest
docker image inspect irvlutd/rpx:tracking-dam4sam-rpx-latest \
  --format '{{index .RepoDigests 0}}'
# Use the printed irvlutd/rpx@sha256:... reference for subsequent runs.

Some linked source guides still show the previous registry names. Translate the full image reference, preserving its original tag:

Previous task repositoryCurrent image reference
rpx-depth-smoke:TAGirvlutd/rpx:depth-TAG
rpx-tracking-smoke:TAGirvlutd/rpx:tracking-TAG
rpx-rcpe-smoke:TAGirvlutd/rpx:rcpe-TAG
rpx-vqa-smoke:TAGirvlutd/rpx:vqa-TAG

This mapping applies to both previous personal namespaces. If a build script constructs tags internally, changing only its repository variable is not enough: pass the complete image reference where supported or adjust its tag construction too.

⁠Publication status

On 9 October 2026, 144 task tags / 102 distinct image manifests were copied and verified against their source digests. Every copied tag was independently resolved through the public registry without a saved Docker login. Historical builds are retained for reproducibility; they are not 102 distinct models.

GEMDepth remains a prebuilt-image gap. NVS is excluded. Publication checks establish image identity and public access, not a new GPU inference pass for each model. Run the selected model's smoke gate before a full benchmark.

To evaluate your own model or extend the benchmark, start with the RPX toolkit documentation⁠.

Tag summary

Content type

Image

Digest

sha256:424aecd70…

Size

10.2 GB

Last updated

1 day ago

docker pull irvlutd/rpx:vqa-vllm-sha-ee4070b0726b