Pith. sign in

REVIEW 11 cited by

FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15259 v3 pith:RXCN3P3C submitted 2024-04-23 cs.CV

classification cs.CV
keywords methoddepthcameraintrinsicsdifferentiablegradient-descentposesdegree
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces FlowMap, an end-to-end differentiable method that solves for precise camera poses, camera intrinsics, and per-frame dense depth of a video sequence. Our method performs per-video gradient-descent minimization of a simple least-squares objective that compares the optical flow induced by depth, intrinsics, and poses against correspondences obtained via off-the-shelf optical flow and point tracking. Alongside the use of point tracks to encourage long-term geometric consistency, we introduce differentiable re-parameterizations of depth, intrinsics, and pose that are amenable to first-order optimization. We empirically show that camera parameters and dense depth recovered by our method enable photo-realistic novel view synthesis on 360-degree trajectories using Gaussian Splatting. Our method not only far outperforms prior gradient-descent based bundle adjustment methods, but surprisingly performs on par with COLMAP, the state-of-the-art SfM method, on the downstream task of 360-degree novel view synthesis (even though our method is purely gradient-descent based, fully differentiable, and presents a complete departure from conventional SfM).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Light3R-SfM: Towards Feed-forward Structure-from-Motion

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Light3R-SfM replaces global bundle adjustment in Structure-from-Motion with a learned attention module and a sparse tree of image pairs, cutting runtime by up to two orders of magnitude while keeping competitive pose ...

  2. Glob3R: Global Structure-from-Motion with 3D Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A frozen Pi3X backbone plus dense warping tracks and keyframe sliding-window global optimization yields more accurate, scalable SfM than feed-forward or classical baselines alone.

  3. CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    CGGS generates viewpoint-consistent, text-aligned ego-centric 3D scenes via consistency-augmented multi-view diffusion, flow-guided layout initialization, and mutual-information depth-refined Gaussian optimization.

  4. 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    4D-Animal fits SMAL animal models to video using silhouette, part, pixel, and tracking losses from off-the-shelf 2D models, removing the need for sparse keypoint annotations.

  5. E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    E3D-Bench compares 16 3D geometric foundation models on depth, reconstruction, pose, and view-synthesis tasks with a unified evaluation toolkit.

  6. Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single-pass transformer generalizes DUSt3R's pointmap regression from two views to all-to-all multi-view attention, reconstructing 1000+ images and estimating camera poses in one forward pass.

  7. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.

  8. LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LoRA3D specializes pretrained 3D foundation models to target scenes via confidence-calibrated pseudo-labels from multi-view robust optimization and LoRA fine-tuning, improving performance by up to 88%.

  9. MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MATE tracks any point from event cameras alone, using motion vectors extracted from time surfaces to guide matching, and reports higher accuracy and survival than video- and event-based baselines on four benchmarks.

  10. Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EgoMono4D estimates depth, camera intrinsics and poses from unlabeled egocentric videos in a single feed-forward pass, reconstructing dense per-frame point clouds better than baseline methods on in-domain and zero-sho...

  11. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

Pith tools