Video by OpenCV via YouTube

The 3D-Object Perception Transformer (3PT) unifies detection, segmentation, and 6DoF pose estimation into two multi-view, RGB-only transformers. Demonstrating exceptional accuracy and cross-domain robustness, it placed first by significant margins in both the Industrial Robotics and AR/VR tracks of the BOP 2025 challenge at ICCV. Today, the 3PT architecture is actively deployed in real-world industrial robotic workcells across the world as the Intrinsic Vision Model (IVM). Join Intrinsic (Google) engineer Agastya Kalra, one of the 3PT authors, as we take a look at this acclaimed paper which was highlighted at CVPR 2026.
Read the paper: https://www.intrinsic.ai/publications/3pt-cvpr2026
OpenCV is a 501(c)(3) registered non-profit in the United States. See how you can support open source CV & AI: http://opencv.org/support/
Watch along for your chance to win during our live trivia segment, and participate in the live Q&A session with questions from you in the audience.
Become a paid member of the channel to help us make more episodes https://www.youtube.com/channel/UCkrcW82Y2kbgU-U9RaYfgxw/join
Got a cool project of your own? Send it to us and you may be featured https://www.jotform.com/form/233105358823151
Thumbnail by Natalia de la Rosa natdlrs.com