Pith. sign in

REVIEW 8 cited by

Flash3D: Feed-Forward Generalisable 3D Scene Reconstruction from a Single Image

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04343 v2 pith:SFBD33OQ submitted 2024-06-06 cs.CV

classification cs.CV
keywords flash3dreconstructionsinglewhenachievesdepthefficientfeed-forward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a "foundation" model for monocular depth estimation and extend it to a full 3D shape and appearance reconstructor. For efficiency, we base this extension on feed-forward Gaussian Splatting. Specifically, we predict a first layer of 3D Gaussians at the predicted depth, and then add additional layers of Gaussians that are offset in space, allowing the model to complete the reconstruction behind occlusions and truncations. Flash3D is very efficient, trainable on a single GPU in a day, and thus accessible to most researchers. It achieves state-of-the-art results when trained and tested on RealEstate10k. When transferred to unseen datasets like NYU it outperforms competitors by a large margin. More impressively, when transferred to KITTI, Flash3D achieves better PSNR than methods trained specifically on that dataset. In some instances, it even outperforms recent methods that use multiple views as input. Code, models, demo, and more results are available at https://www.robots.ox.ac.uk/~vgg/research/flash3d/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TinySplat: Feedforward Approach for Generating Compact 3D Scene Representation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    TinySplat compresses feedforward 3D Gaussian scenes by 105-199x on two-view benchmarks (about 50x on DL3DV) while keeping rendered quality close to the uncompressed model.

  2. GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    GeoWorld improves image-to-3D scene generation by conditioning a video-diffusion model on full-frame geometry features extracted by a multi-view geometry model, yielding higher PSNR/SSIM/LPIPS than prior methods.

  3. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.

  4. iLRM: An Iterative Large 3D Reconstruction Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    iLRM reconstructs 3D Gaussian scenes from multiple photos through iterative refinement of viewpoint tokens, achieving higher quality and speed than prior feed-forward models.

  5. VoluMe -- Authentic 3D Video Calls from Live Gaussian Splat Prediction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    VoluMe predicts real-time 3D Gaussian head reconstructions from a single webcam feed, preserving the input view while allowing realistic novel viewpoints for 3D video calls.

  6. BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.

  7. Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SceneDINO performs semantic scene completion from a single image in a fully unsupervised way by lifting self-supervised DINO features into a 3D feature field trained with multi-view consistency.

  8. Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A plug-and-play 3D Chamfer loss, using pointmaps from a pretrained transformer as pseudo-ground truth, improves feed-forward 3DGS rendering and geometry across MVSplat and DepthSplat.

Pith tools