RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025 Highlight

Raktim Gautam Goswami1, Prashanth Krishnamurthy1, Yann LeCun2,3, Farshad Khorrami1

1 New York University Tandon School of Engineering
2 New York University Courant Institute of Mathematical Sciences
3 Meta-FAIR
Fig. 1: Overview of our RoboPEPP framework.

RoboPEPP Contributions

Video

Results

Table 1: Comparison of robot pose estimation using AUC on the ADD metric. Best values among methods using unknown joint angles and bounding boxes during evaluation are bolded. HPE∗ denotes HPE [4] evaluated with the same off-the-shelf bounding box detector as RoboPEPP.
Fig. 2: Qualitative Comparison on Panda Photo (Example 1) and Occlusion (Example 2 and 3) datasets: Predicted poses and joint angles are used to generate a mesh overlaid on the original image, where closer alignment indicates greater accuracy. Highlighted rectangles indicate regions where other methods’ meshes misalign, while RoboPEPP achieves high precision.

Citation

@inproceedings{goswami2025robopepp,
  title={Robopepp: Vision-based robot pose and joint angle estimation through embedding predictive pre-training},
  author={Goswami, Raktim Gautam and Krishnamurthy, Prashanth and LeCun, Yann and Khorrami, Farshad},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={6930--6939},
  year={2025}
}