REVIEW 15 cited by
Grounding Image Matching in 3D with MASt3R
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Image Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D problem. This makes sense as the goal of matching is to establish correspondences between 2D pixel fields, but also seems like a potentially hazardous choice. In this work, we take a different stance and propose to cast matching as a 3D task with DUSt3R, a recent and powerful 3D reconstruction framework based on Transformers. Based on pointmaps regression, this method displayed impressive robustness in matching views with extreme viewpoint changes, yet with limited accuracy. We aim here to improve the matching capabilities of such an approach while preserving its robustness. We thus propose to augment the DUSt3R network with a new head that outputs dense local features, trained with an additional matching loss. We further address the issue of quadratic complexity of dense matching, which becomes prohibitively slow for downstream applications if not carefully treated. We introduce a fast reciprocal matching scheme that not only accelerates matching by orders of magnitude, but also comes with theoretical guarantees and, lastly, yields improved results. Extensive experiments show that our approach, coined MASt3R, significantly outperforms the state of the art on multiple matching tasks. In particular, it beats the best published methods by 30% (absolute improvement) in VCRE AUC on the extremely challenging Map-free localization dataset.
Forward citations
Cited by 15 Pith papers
-
Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
Rig3R conditions learned 3D reconstruction on optional rig metadata and predicts rig-relative raymaps, enabling state-of-the-art pose estimation and rig calibration discovery from images.
-
Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
Past tokens that the model can predict from their future context are evicted from the KV cache, judged by a counter-causal attention pass that reuses cached keys and values.
-
GeoWorldAD: Geometry World Action Model for Autonomous Driving
Grounding an autonomous-driving action model in ego-aligned multi-scale 3D geometry and latent future-geometry tokens improves NAVSIM closed-loop PDMS/EPDMS over prior geometry- and world-model-based planners.
-
PACE: Polar Axis-Conditioned Estimation for PairUAV Relative Localization
A shared image-pair network beats a single-head baseline by giving heading and range their own decoder readouts—PACE's raw model scores 0.002460 on the PairUAV hidden test.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.
-
MACRO: Training-free Multi-plane Attention for Closeup Render Optimization
Training-free multi-plane attention with image-space scale-matched reference crops restores correct close-up detail from 3DGS without retraining the enhancer.
-
G3Splat: Geometrically Consistent Generalizable Gaussian Splatting
Adding ray-alignment and local-normal orientation losses to generalizable Gaussian splatting fixes geometrically degenerate splats and improves zero-shot depth, mesh, and pose estimation.
-
UFV-Splatter: Pose-Free Feed-Forward 3D Gaussian Splatting Adapted to Unfavorable Views
Re-centering inputs, LoRA fine-tuning, and a Gaussian refiner let a pretrained pose-free 3D Gaussian Splatting model reconstruct objects from off-center, unknown camera views without any unfavorable-view training data.
-
DCHM: Depth-Consistent Human Modeling for Multiview Detection
DCHM uses superpixel-based Gaussian Splatting to make monocular depth estimates multiview-consistent, producing point clouds that yield state-of-the-art label-free pedestrian detection on Wildtrack, Terrace, and MultiviewX.
-
SteerPose: Simultaneous Extrinsic Camera Calibration and Matching from Articulation
A learned mental rotation of 2D articulated poses estimates relative camera rotations and cross-view instance correspondences at the same time, enabling calibration and 3D reconstruction in multi-camera setups.
-
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.
-
Quo Vadis, World Modeling?
An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.
-
Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian Splatting
NIRSplat fuses NIR imagery and vegetation-index metadata with 3D Gaussian Splatting via cross-attention and positional encoding, outperforming 3DGS, CoR-GS, and InstantSplat on the new multimodal agriculture dataset NIRPlant.
-
Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps
S3PO-GS uses the 3DGS-rendered pointmap as the scale anchor for pose estimation and patch-aligns pretrained pointmaps to map scale, improving outdoor monocular 3DGS SLAM tracking and rendering.
-
Surf3R: Rapid Surface Reconstruction from Sparse RGB Views in Seconds
Surf3R reconstructs 3D surfaces from sparse unposed RGB views in under 10 seconds using a multi-branch feedforward network with Gaussian-based depth-normal regularization.
Discussion (0). Sign in to comment.