Pith. sign in

REVIEW 10 cited by

MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.06870 v1 pith:ERMVCOZY submitted 2022-12-13 cs.CV cs.RO

classification cs.CVcs.RO
keywords objectsnovelposeobjectapproachsyntheticdatasetestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a region of interest displaying the object in the image and (ii) a CAD model of the observed object. The contributions of this work are threefold. First, we present a 6D pose refiner based on a render&compare strategy which can be applied to novel objects. The shape and coordinate system of the novel object are provided as inputs to the network by rendering multiple synthetic views of the object's CAD model. Second, we introduce a novel approach for coarse pose estimation which leverages a network trained to classify whether the pose error between a synthetic rendering and an observed image of the same object can be corrected by the refiner. Third, we introduce a large-scale synthetic dataset of photorealistic images of thousands of objects with diverse visual and shape properties and show that this diversity is crucial to obtain good generalization performance on novel objects. We train our approach on this large synthetic dataset and apply it without retraining to hundreds of novel objects in real images from several pose estimation benchmarks. Our approach achieves state-of-the-art performance on the ModelNet and YCB-Video datasets. An extensive evaluation on the 7 core datasets of the BOP challenge demonstrates that our approach achieves performance competitive with existing approaches that require access to the target objects during training. Code, dataset and trained models are available on the project page: https://megapose6d.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    OneViewAll reports 92.5% ADD-0.1 pose accuracy on LINEMOD from a single real reference RGB-D view, using projection-based refinement with mirror-fusion symmetry priors rather than CAD rendering.

  2. One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.

  3. RadGS-Reg: Registering Spine CT with Biplanar X-rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A joint 3D Gaussian-splatting-based reconstruction and registration network registers spine CT to two X-rays with 1.14 mm mean error in 0.82 seconds on a small in-house set.

  4. Computer Vision for Objects used in Group Work: Challenges and Opportunities

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FiboSB is a new 6D pose dataset for blocks in group work; all four state-of-the-art 6D pose methods failed on it, while a fine-tuned YOLO11-x detector achieved 0.898 mAP50.

  5. GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Adding a category-level semantic shape prior, reconstructed from partial RGB-D input, improves 6D pose and size estimation for unseen objects on HouseCat6D and NOCS-REAL275.

  6. PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free geometry-only pipeline matches RGB images against rendered depth/normal maps to estimate 6D poses of unseen, textureless, and slightly defective objects from one image.

  7. DynamicPose: Real-time and Robust 6D Object Pose Tracking for Fast-Moving Cameras and Objects

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DynamicPose maintains real-time 6D object pose during fast camera and object motion by combining VIO-based ROI compensation, depth-informed 2D tracking, and VIO-guided Kalman-filter pose prediction in a closed loop.

  8. IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The claimed IDC-Net framework is absent; the body text is an unrelated instance-segmentation paper.

  9. Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data

    cs.CV 2025-05 conditional novelty 5.0 of 10

    PLANTPose estimates a banana's 6D pose and per-instance shape deformation from RGB images using a lattice-based mesh warp, trained on synthetic scenes refined by Stable Diffusion, and reports large gains over MegaPose...

  10. RGBTrack: Fast, Robust Depth-Free 6D Pose Estimation and Tracking

    cs.CV 2025-06

Pith tools