Zero-shot text-prompted 3D pose estimation on ROS 2 - SAM 3 + RGB-D. No training, no CAD models.
444
Zero-shot object detection, tracking, and 6-DoF pose estimation on ROS 2. Type a text prompt ("thermos", "book"...) - get 3D position, size, orientation and a confidence covariance for each object on a conveyor. No training, no CAD models.
vision_msgs/Detection3DArray (+ position covariance), RViz markers, debug video, CSVgit clone https://github.com/ingon1026/pose-anything.git && cd pose-anything
export HF_TOKEN=hf_xxxx # HF account with gated facebook/sam3 access
docker compose run --rm perception # live RealSense D455
docker compose run --rm perception ./run.sh bags/mybag --prompts "book" # rosbag
Requires NVIDIA GPU (>= 6 GB VRAM) + nvidia-container-toolkit. Model weights are not in the image (Meta gated license) - downloaded once into a mounted HF cache volume.
| Tag | Contents |
|---|---|
latest, 1.1.1 | probabilistic fusion filter, occlusion handling, RViz world TF |
1.1.0 | first fusion-filter build |
1.0.0 | initial release (binary quality gates) |
Full documentation, measured accuracy numbers and source: https://github.com/ingon1026/pose-anything
Content type
Image
Digest
sha256:96657bbd7…
Size
7.2 GB
Last updated
about 1 month ago
docker pull ingon1026/pose-anything