Pith. sign in

REVIEW 17 cited by

Continuous 3D Perception Model with Persistent State

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.12387 v1 pith:HFDRWFJ4 submitted 2025-01-21 cs.CV

classification cs.CV
keywords imagesmodelpointmapsstatecontinuouscut3rmethodreconstruction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolving state can be used to generate metric-scale pointmaps (per-pixel 3D points) for each new input in an online fashion. These pointmaps reside within a common coordinate system, and can be accumulated into a coherent, dense scene reconstruction that updates as new images arrive. Our model, called CUT3R (Continuous Updating Transformer for 3D Reconstruction), captures rich priors of real-world scenes: not only can it predict accurate pointmaps from image observations, but it can also infer unseen regions of the scene by probing at virtual, unobserved views. Our method is simple yet highly flexible, naturally accepting varying lengths of images that may be either video streams or unordered photo collections, containing both static and dynamic content. We evaluate our method on various 3D/4D tasks and demonstrate competitive or state-of-the-art performance in each. Project Page: https://cut3r.github.io/

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Rig3R conditions learned 3D reconstruction on optional rig metadata and predicts rig-relative raymaps, enabling state-of-the-art pose estimation and rig calibration discovery from images.

  2. Syn4D: A Multiview Synthetic 4D Dataset

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Syn4D supplies multiview synthetic dynamic scenes with dense geometric, tracking and pose ground truth that lets any pixel be unprojected to any time and camera.

  3. LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An incremental 3D Gaussian Splatting pipeline that jointly optimizes camera poses and scene geometry using MASt3R priors and density-adaptive octree anchors achieves state-of-the-art novel view synthesis on casual lon...

  4. STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A decoder-only Transformer with causal attention and cached past-frame features performs incremental 3D reconstruction from streaming images, beating the RNN-based CUT3R on several benchmark metrics.

  5. LongSplat: Online Generalizable 3D Gaussian Splatting from Long Sequence Images

    cs.CV 2025-07 reject novelty 6.0 of 10

    A feed-forward 3D Gaussian Splatting pipeline that incrementally fuses and compresses historical Gaussians using a 2D image-like representation.

  6. SpatialTrackerV2: 3D Point Tracking Made Easy

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.

  7. Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.

  8. Test3R: Learning to Reconstruct 3D at Test Time

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Test3R improves 3D reconstruction by optimizing visual prompts at test time so that pointmaps from different image pairs are geometrically consistent.

  9. EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EX-4D uses a depth watertight mesh and simulated occlusion masks to condition a video diffusion model for extreme-viewpoint 4D video synthesis from monocular input.

  10. RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A ray-conditioned transformer turns single-image 3D shape completion into novel-view depth prediction, achieving state-of-the-art chamfer distance on synthetic and real benchmarks.

  11. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fine-tuning a pretrained video diffusion transformer to predict geometry in one shared global frame produces consistent, camera-free surface normals and coordinates across entire video clips.

  12. X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography

    eess.IV 2025-05 conditional novelty 6.0 of 10

    A large transformer with fixed-voxel Gaussian splatting reconstructs CT volumes from 6-10 X-ray projections in under a second, substantially beating prior sparse-view methods in simulation.

  13. RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

    cs.CV 2026-07 conditional novelty 5.5 of 10

    A feed-forward Gaussian head on OmniVGGT plus road-plane grid fusion and structure-aware grouping reconstructs compact road surfaces that beat RoGS and AnySplat on Waymo and zero-shot nuScenes.

  14. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  15. InstantSfM: Towards GPU-Native SfM for the Deep Learning Era

    cs.CV 2025-10 conditional novelty 5.0 of 10

    A fully GPU-native, PyTorch-based global Structure-from-Motion pipeline using sparse-aware Levenberg-Marquardt with optional metric depth priors reports ~8-40× speedups over COLMAP at comparable accuracy on several be...

  16. UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.

  17. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

Pith tools