REVIEW 10 cited by
MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a region of interest displaying the object in the image and (ii) a CAD model of the observed object. The contributions of this work are threefold. First, we present a 6D pose refiner based on a render&compare strategy which can be applied to novel objects. The shape and coordinate system of the novel object are provided as inputs to the network by rendering multiple synthetic views of the object's CAD model. Second, we introduce a novel approach for coarse pose estimation which leverages a network trained to classify whether the pose error between a synthetic rendering and an observed image of the same object can be corrected by the refiner. Third, we introduce a large-scale synthetic dataset of photorealistic images of thousands of objects with diverse visual and shape properties and show that this diversity is crucial to obtain good generalization performance on novel objects. We train our approach on this large synthetic dataset and apply it without retraining to hundreds of novel objects in real images from several pose estimation benchmarks. Our approach achieves state-of-the-art performance on the ModelNet and YCB-Video datasets. An extensive evaluation on the 7 core datasets of the BOP challenge demonstrates that our approach achieves performance competitive with existing approaches that require access to the target objects during training. Code, dataset and trained models are available on the project page: https://megapose6d.github.io/.
Forward citations
Cited by 10 Pith papers
-
Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects
OneViewAll reports 92.5% ADD-0.1 pose accuracy on LINEMOD from a single real reference RGB-D view, using projection-based refinement with mirror-fusion symmetry priors rather than CAD rendering.
-
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.
-
RadGS-Reg: Registering Spine CT with Biplanar X-rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration
A joint 3D Gaussian-splatting-based reconstruction and registration network registers spine CT to two X-rays with 1.14 mm mean error in 0.82 seconds on a small in-house set.
-
Computer Vision for Objects used in Group Work: Challenges and Opportunities
FiboSB is a new 6D pose dataset for blocks in group work; all four state-of-the-art 6D pose methods failed on it, while a fine-tuned YOLO11-x detector achieved 0.898 mAP50.
-
GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation
Adding a category-level semantic shape prior, reconstructed from partial RGB-D input, improves 6D pose and size estimation for unseen objects on HouseCat6D and NOCS-REAL275.
-
PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects
A training-free geometry-only pipeline matches RGB images against rendered depth/normal maps to estimate 6D poses of unseen, textureless, and slightly defective objects from one image.
-
DynamicPose: Real-time and Robust 6D Object Pose Tracking for Fast-Moving Cameras and Objects
DynamicPose maintains real-time 6D object pose during fast camera and object motion by combining VIO-based ROI compensation, depth-informed 2D tracking, and VIO-guided Kalman-filter pose prediction in a closed loop.
-
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
The claimed IDC-Net framework is absent; the body text is an unrelated instance-segmentation paper.
-
Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data
PLANTPose estimates a banana's 6D pose and per-instance shape deformation from RGB images using a lattice-based mesh warp, trained on synthetic scenes refined by Stable Diffusion, and reports large gains over MegaPose...
- RGBTrack: Fast, Robust Depth-Free 6D Pose Estimation and Tracking
Discussion (0). Continue with ORCID to comment.