Pith. sign in

REVIEW 9 cited by

Neural Volumes: Learning Dynamic Renderable Volumes from Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.07751 v1 pith:7PYZZ7FH submitted 2019-06-18 cs.GR cs.CV

classification cs.GRcs.CV
keywords dynamicimagesrepresentationapproachcaptureduringenablesencoder-decoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occlusion, and biological motion. Mesh-based reconstruction and tracking often fail in these cases, and other approaches (e.g., light field video) typically rely on constrained viewing conditions, which limit interactivity. We circumvent these difficulties by presenting a learning-based approach to representing dynamic objects inspired by the integral projection model used in tomographic imaging. The approach is supervised directly from 2D images in a multi-view capture setting and does not require explicit reconstruction or tracking of the object. Our method has two primary components: an encoder-decoder network that transforms input images into a 3D volume representation, and a differentiable ray-marching operation that enables end-to-end training. By virtue of its 3D representation, our construction extrapolates better to novel viewpoints compared to screen-space rendering techniques. The encoder-decoder architecture learns a latent representation of a dynamic scene that enables us to produce novel content sequences not seen during training. To overcome memory limitations of voxel-based representations, we learn a dynamic irregular grid structure implemented with a warp field during ray-marching. This structure greatly improves the apparent resolution and reduces grid-like artifacts and jagged motion. Finally, we demonstrate how to incorporate surface-based representations into our volumetric-learning framework for applications where the highest resolution is required, using facial performance capture as a case in point.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

    cs.GR 2025-06 conditional novelty 7.0 of 10

    A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.

  2. Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Lume-Palette decouples multi-view indoor relighting into diffusion-based distillation of canonical illumination palettes and casting under receiver-centric 3D lighting maps with asymmetric multi-view conditioning.

  3. ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ATLAS decouples skeleton and shape parameters in a parametric human body model, improving fit accuracy and controllability over previous models like SMPL-X.

  4. PixNerd: Pixel Neural Field Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PixNerd is a single-stage pixel-space diffusion transformer that uses predicted neural field weights to decode large patches, reaching 2.15 FID on ImageNet 256 without a VAE.

  5. ViscoReg: Neural Signed Distance Functions via Viscosity Solutions

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A viscosity-regularized Eikonal loss with annealed epsilon improves Neural SDF reconstruction and yields the first generalization bound for SDF learning.

  6. Flow-Anything: Learning Real-World Optical Flow Estimation from Large-Scale Single-view Images

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A single-image-to-3D pipeline that renders realistic image pairs and flow labels at scale, training optical flow models to outperform synthetic-data and unsupervised baselines on KITTI.

  7. LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LocalDyGS reconstructs dynamic scenes by decomposing space into seed-based local regions and generating time-varying Temporal Gaussians, though its claim of being first for large-scale scenes omits the existing Swift4...

  8. Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.

  9. Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting

    cs.GR 2025-08 conditional novelty 4.0 of 10

    A three-stage perceive-sample-compress framework for 3D Gaussian Splatting improves rendering fidelity and storage efficiency across small and large scenes.

Pith tools