Sign inSign up

ingon1026/pose-anything

By ingon1026

•Updated about 1 month ago

Zero-shot text-prompted 3D pose estimation on ROS 2 - SAM 3 + RGB-D. No training, no CAD models.

Image
Machine learning & AI
0

444

ingon1026/pose-anything repository overview

⁠pose-anything

Zero-shot object detection, tracking, and 6-DoF pose estimation on ROS 2. Type a text prompt ("thermos", "book"...) - get 3D position, size, orientation and a confidence covariance for each object on a conveyor. No training, no CAD models.

  • Perception: SAM 3 (Meta, open-vocabulary) keyframes + optical-flow tracking, 9-13 FPS on RTX 4070 Ti
  • Pose: RGB-D back-projection -> Open3D OBB, per-track Kalman fusion with chi-square observation gating
  • Robust to occlusion: corrupted observations are rejected; unreliable poses are withheld, never published
  • Output: vision_msgs/Detection3DArray (+ position covariance), RViz markers, debug video, CSV

⁠Quick start

git clone https://github.com/ingon1026/pose-anything.git && cd pose-anything
export HF_TOKEN=hf_xxxx     # HF account with gated facebook/sam3 access
docker compose run --rm perception                                  # live RealSense D455
docker compose run --rm perception ./run.sh bags/mybag --prompts "book"   # rosbag

Requires NVIDIA GPU (>= 6 GB VRAM) + nvidia-container-toolkit. Model weights are not in the image (Meta gated license) - downloaded once into a mounted HF cache volume.

⁠Tags

TagContents
latest, 1.1.1probabilistic fusion filter, occlusion handling, RViz world TF
1.1.0first fusion-filter build
1.0.0initial release (binary quality gates)

Full documentation, measured accuracy numbers and source: https://github.com/ingon1026/pose-anything⁠

Tag summary

Content type

Image

Digest

sha256:96657bbd7…

Size

7.2 GB

Last updated

about 1 month ago

docker pull ingon1026/pose-anything