Pith. sign in

REVIEW 12 cited by

VGGT: Visual Geometry Grounded Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.11651 v1 pith:55L4SHDZ submitted 2025-03-14 cs.CV

classification cs.CV
keywords pointvggttaskscameradepthestimationfeed-forwardgeometry
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typically been constrained to and specialized for single tasks. It is also simple and efficient, reconstructing images in under one second, and still outperforming alternatives that require post-processing with visual geometry optimization techniques. The network achieves state-of-the-art results in multiple 3D tasks, including camera parameter estimation, multi-view depth estimation, dense point cloud reconstruction, and 3D point tracking. We also show that using pretrained VGGT as a feature backbone significantly enhances downstream tasks, such as non-rigid point tracking and feed-forward novel view synthesis. Code and models are publicly available at https://github.com/facebookresearch/vggt.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RayOcc: Occlusion-Aware Ray Occupancy Estimation via Gaussian Mixture Intensity

    cs.CV 2026-07 conditional novelty 6.0 of 10

    RayOcc models each camera ray as a non-normalized Gaussian mixture with Poisson-based occupancy probabilities, allowing multiple depth hypotheses per ray and improving Gaussian-initialized 3D occupancy prediction on nuScenes.

  2. PACE: Polar Axis-Conditioned Estimation for PairUAV Relative Localization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A shared image-pair network beats a single-head baseline by giving heading and range their own decoder readouts—PACE's raw model scores 0.002460 on the PairUAV hidden test.

  3. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  4. UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A diffusion-prior pipeline that reconstructs and tracks novel objects from single RGB-D frames, using pixel-wise uncertainty to guide 3D Gaussian Splatting and pose graph optimization.

  5. Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.

  6. BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A dual-stream diffusion model trained with Blender-render conditioning, source masking, and object jittering performs 3D-grounded multi-object editing and compositing better than existing baselines on three video datasets.

  7. Test3R: Learning to Reconstruct 3D at Test Time

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Test3R improves 3D reconstruction by optimizing visual prompts at test time so that pointmaps from different image pairs are geometrically consistent.

  8. PointGS: Point Attention-Aware Sparse View Synthesis with Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PointGS improves few-shot 3D Gaussian splatting by fusing multi-view image features per 3D point and refining them with a neighbor-attention network before decoding Gaussian colors.

  9. SalientGS: Unified SfM-to-3DGS with Importance-Guided MCMC Gaussian Allocation

    cs.CV 2026-07 conditional novelty 5.5 of 10

    SalientGS integrates fast first-order SfM, joint pose refinement, and importance-guided MCMC Gaussian birth/relocation to reach 27.65 dB macro-average PSNR at 1.5M Gaussians in about 10 minutes end-to-end.

  10. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5 of 10

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  11. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  12. StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation

    cs.RO 2026-02 reject novelty 4.0 of 10

    StemVLA supervises a GPT-2-based VLA with predicted future 3D-geometry features (VGGT) and temporally aggregated history, reporting 86.0% on LIBERO-Long - but its CALVIN results and equations are placeholders.

Pith tools