Pith. sign in

REVIEW 6 cited by

pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12337 v4 pith:KZ2ZKILR submitted 2023-12-19 cs.CV cs.LG

classification cs.CVcs.LG
keywords gaussiandistributionfieldmodelpairspixelsplatprobabilityradiance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce pixelSplat, a feed-forward model that learns to reconstruct 3D radiance fields parameterized by 3D Gaussian primitives from pairs of images. Our model features real-time and memory-efficient rendering for scalable training as well as fast 3D reconstruction at inference time. To overcome local minima inherent to sparse and locally supported representations, we predict a dense probability distribution over 3D and sample Gaussian means from that probability distribution. We make this sampling operation differentiable via a reparameterization trick, allowing us to back-propagate gradients through the Gaussian splatting representation. We benchmark our method on wide-baseline novel view synthesis on the real-world RealEstate10k and ACID datasets, where we outperform state-of-the-art light field transformers and accelerate rendering by 2.5 orders of magnitude while reconstructing an interpretable and editable 3D radiance field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.

  2. GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.

  3. DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DrivingRecon predicts 4D Gaussians of street scenes from surround-view video in one forward pass, using a novel Prune and Dilate Block to reduce redundant overlapping points.

  4. VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.5 of 10

    VistaVLA lifts multi-view vision-language features into 3D Gaussians, compresses them 99% via Merge-then-Query, and improves real-robot manipulation success by ~23% over baselines.

  5. Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis

    cs.CV 2025-01 reject novelty 5.0 of 10

    Aug3D augments large outdoor datasets with synthetic views rendered from SfM reconstructions, improving PixelNeRF's PSNR from 20.03 to 21.80 on UrbanScene3D Campus, though reducing cluster size alone reached 22.94.

  6. Instructive3D: Editing Large Reconstruction Models with Text Instructions

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A text-conditioned diffusion adapter operating on the triplane latents of a frozen large reconstruction model enables natural-language editing of generated 3D objects.

Pith tools