NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model
AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.



