Pith. sign in

REVIEW 7 cited by

FoundationStereo: Zero-Shot Stereo Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09898 v4 pith:AQSGC5YZ submitted 2025-01-17 cs.CV cs.LGcs.RO

classification cs.CVcs.LGcs.RO
keywords stereozero-shotfoundationfoundationstereomatchingstrongcomponentsdepth
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization - a hallmark of foundation models in other computer vision tasks - remains challenging for stereo matching. We introduce FoundationStereo, a foundation model for stereo depth estimation designed to achieve strong zero-shot generalization. To this end, we first construct a large-scale (1M stereo pairs) synthetic training dataset featuring large diversity and high photorealism, followed by an automatic self-curation pipeline to remove ambiguous samples. We then design a number of network architecture components to enhance scalability, including a side-tuning feature backbone that adapts rich monocular priors from vision foundation models to mitigate the sim-to-real gap, and long-range context reasoning for effective cost volume filtering. Together, these components lead to strong robustness and accuracy across domains, establishing a new standard in zero-shot stereo depth estimation. Project page: https://nvlabs.github.io/FoundationStereo/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans

    cs.RO 2026-03 conditional novelty 6.0 of 10

    TiPToP, a zero-training modular planner using pretrained vision-language models and GPU-accelerated TAMP, achieves 74.6% success over 165 trials versus 52.4% for the 350-hour-trained pi0.5-DROID baseline across 28 man...

  2. 3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.

  3. ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ParticleFormer uses a Transformer over point-cloud particles and a hybrid Chamfer-Hausdorff loss to predict multi-material object dynamics, and it reports lower errors than GNN and image-based baselines in simulation ...

  4. Repurposing Marigold for Zero-Shot Metric Depth Estimation via Defocus Blur Cues

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Two differently blurred images plus a pretrained diffusion depth prior are optimized together at inference time to recover metric depth without retraining.

  5. Diving into the Fusion of Monocular Priors for Generalized Stereo Matching

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Local ordering maps and per-pixel affine registration of monocular depth improve zero-shot generalization of iterative stereo matching on ill-posed regions.

  6. Performance of universal machine-learned potentials with explicit long-range interactions in biomolecular simulations

    physics.chem-ph 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims a systematic benchmark of universal machine-learned potentials on biomolecular simulations, but the body text supplied is an unrelated stereo-vision paper, leaving the claim unverifiable.

  7. BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A single network that iteratively aligns monocular features with stereo hypotheses reduces zero-shot stereo depth error by over 40% on Middlebury and ETH3D.

Pith tools