Sign inSign up

humanoidsctu/vitpose

By humanoidsctu

•Updated 9 days ago

Image
0

686

humanoidsctu/vitpose repository overview

This image, with the "paper" tag has been used for the paper: https://link.springer.com/article/10.3758/s13428-025-02816-x⁠ (free access to the pdf: https://rdcu.be/eFwac⁠) The Dockerfile can be found at the osf repository: https://osf.io/vaexp/⁠
Note: the internal_pipeline tagged image is updated for our own purposes for convenience, the commands and instructions listed in the readme below might not work as-is for it. (the image contains its own readme, but be warned that the image can be changed at any moment, without notice, and without description of the changes.)

⁠0. Prerequisite

The host computer must have the appropriate packages installed to allow GPUs to be used by the Docker containers, e.g., for Nvidia GPUs, it requires to install the Nvidia container toolkit: https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html⁠

All ready-to-use Docker containers (including this one), are constrained to the dependencies software versions they had when built. The main cause for compatibility issues and the container not working properly comes from the graphic card and the versions of CUDA/cudnn/Pytorch. In particular, any recent enough GPU that is not handled by the pre-installed version would lead to errors, or the images/videos not being processed. In such a case, the fix is to use the Dockerfile shared on our osf and modify the appropriate lines to use versions of the software. For example, using an Nvidia GPU RTX A4000 or A5000 requires to modify the Dockerfile to use CUDA version 11.3, and cudnn 8; then to rebuild a docker container using the updated Dockerfile.

⁠1. Docker build

Dockerfile based on the official Dockerfile from April 2023 https://github.com/open-mmlab/mmpose/blob/main/master/docker⁠ with added MMDetection and MMClassification installation

Also see: https://mmdetection.readthedocs.io/en/stable/get_started.html⁠ https://mmpose.readthedocs.io/en/v0.28.0/install.html⁠ https://mmpose.readthedocs.io/en/latest/model_zoo_papers/datasets.html?highlight=vitpose#topdown-heatmap-vitpose-on-coco⁠

docker build -t vitpose:paper .

⁠2. Docker startup

Default command: sudo docker run --rm --gpus all -it -e="DISPLAY" -v=/tmp/.X11-unix:/tmp/.X11-unix:rw -v /home/$user:/home/$user vitpose:paper /bin/bash

Replace $user by the session's username.

Add -v /media:/media to give access to external hard drives

If needed, limit the amount of RAM shared with docker with the parameter: --shm-size=8gb

⁠3. Running ViTPose (Top-Down)

Replace $data_path by the path to the data folder and $result_path by the path to the result folder.

Memory profiling can be done by adding the following command before the processing commands described below, depending on the method: mprof run --include-children --multiprocess --output $result_path/vitpose/mprofile.dat

To run ViTPose, run the following command: python demo/top_down_mmdet.py demo/mmdetection_cfg/faster_rcnn_r50_fpn_coco.py vitpose_configs/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth configs/body_2d_keypoint/topdown_heatmap/coco/td-hm_ViTPose-huge_8xb64-210e_coco-256x192.py vitpose_configs/td-hm_ViTPose-huge_8xb64-210e_coco-256x192-e32adcd4_20230314.pth --input $data_path/$image_video.png --output-root $result_path

This will create a folder $result_path/keypoints_coco17 in which the predictions output will be saved in .json files.

$image_video.png should be the name of the video or the first image in the folder being processed (it will find and process all the images in the folder).

Add the following to output images with keypoints and skeleton, it will create a $result_path/vis folder inside which the images will be: --save-visuals

To avoid potential changes or unavailability of the files, and ensure reproduction of results, the configuration files were downloaded on 2023.06.14 and 2023.01.11 respectively from:

  1. https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pthas file vitpose_configs/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth
  2. https://download.openmmlab.com/mmpose/v1/body_2d_keypoint/topdown_heatmap/coco/td-hm_ViTPose-huge_8xb64-210e_coco-256x192-e32adcd4_20230314.pthas file vitpose_configs/td-hm_ViTPose-huge_8xb64-210e_coco-256x192-e32adcd4_20230314.pth

⁠4. Changes from the official github:

Modifications made to the script top_down_mmdet.py:

  • Row 4 added import glob
  • Rows 72-81 added function to custom save results to .json
  • Rows 106-110 added argument for saving of visualisation
  • Rows 113-114 changed default of saving prediction as True
  • Rows 175-176 commented asserts for visualisation
  • Rows 225-226 added loop to go through all images
  • Rows 233-235 added custom saving of predictions
  • Rows 238,274 changed condition for outputting visualisation
  • Rows 239, 240 added lines to update output file name
  • Row 254 deleted for being redundant
  • Rows 270-273 added custom saving of predictions
  • Row 271 frame_idx increment moved here
  • Row 308 deleted redundant saving of predictions

Tag summary

Content type

Image

Digest

sha256:4eabb7ad5…

Size

9.6 GB

Last updated

9 days ago

docker pull humanoidsctu/vitpose:internal_pipeline_gui_v1.0