REVIEW 24 cited by
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We present a method that achieves state-of-the-art results for synthesizing novel views of complex scenes by optimizing an underlying continuous volumetric scene function using a sparse set of input views. Our algorithm represents a scene using a fully-connected (non-convolutional) deep network, whose input is a single continuous 5D coordinate (spatial location $(x,y,z)$ and viewing direction $(\theta, \phi)$) and whose output is the volume density and view-dependent emitted radiance at that spatial location. We synthesize views by querying 5D coordinates along camera rays and use classic volume rendering techniques to project the output colors and densities into an image. Because volume rendering is naturally differentiable, the only input required to optimize our representation is a set of images with known camera poses. We describe how to effectively optimize neural radiance fields to render photorealistic novel views of scenes with complicated geometry and appearance, and demonstrate results that outperform prior work on neural rendering and view synthesis. View synthesis results are best viewed as videos, so we urge readers to view our supplementary video for convincing comparisons.
Forward citations
Cited by 24 Pith papers
-
Detangled: A Framework for Creating, Editing, and Inferencing Feature Rich Hair Strands
A 5D texture parameterization plus centerline-based canonical space and supervised diffusion enables generation and texture transfer of feature-rich hair strands independent of style.
-
PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur
PRISM3D bootstraps 3D Gaussian Splatting from extreme motion blur via VGGSfM initialization, MCMC densification, and Bézier trajectories, with an event-assisted extension that sets new SOTA.
-
SubdivAR: Autoregressive Next-Scale Prediction for Neural Mesh Subdivision
SubdivAR reformulates neural mesh subdivision as autoregressive next-scale vertex-offset prediction, reporting 18.8% lower Hausdorff and 14.2% lower Chamfer distance than NMR on closed meshes.
-
You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes
YOGO reformulates stochastic 3D Gaussian Splatting into a deterministic budget-aware system and supplies an ultra-dense dataset to enforce physical fidelity over viewpoint interpolation.
-
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.
-
PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations
PokeNet estimates joint types, axes, ranges, and operation order of articulated objects directly from a single-view point cloud video of a human demonstration.
-
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
Flow equivariant world models use a latent memory that shifts with the agent and with inferred object motion, giving stable long-horizon prediction under partial observability.
-
MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
MRD finds physically different 3D scenes that reproduce a target model activation, revealing which shape and material properties vision models are sensitive to.
-
Deformable Medical Image Registration with KAN-based Implicit Neural Representations
KAN-based implicit neural networks with randomized basis sampling outperform existing INR registration methods on three medical imaging datasets at lower computational cost.
-
LuxDiT: Lighting Estimation with Video Diffusion Transformer
A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.
-
SLRTP2025 Sign Language Production Challenge: Methodology, Results, and Future Work
A first competitive benchmark for sign language production, with a hidden test set, retrieval-based winning systems, and a released evaluation network.
-
Variational volume reconstruction with the Deep Ritz Method
A Deep Ritz variational method with a modified Cahn-Hilliard regularizer reconstructs volumes from sparse noisy slices without segmentation.
-
GraphBrep: Learning B-Rep in Graph Structure for Efficient CAD Generation
GraphBrep replaces the redundant tree-based topology of prior B-Rep generators with an explicit graph adjacency representation, cutting training and inference cost while preserving generation quality.
-
RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction
A feed-forward Gaussian head on OmniVGGT plus road-plane grid fusion and structure-aware grouping reconstructs compact road surfaces that beat RoGS and AnySplat on Waymo and zero-shot nuScenes.
-
Quo Vadis, World Modeling?
An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.
-
Towards optimal photometric calibration of digital astronomical plates with deep learning
A deep network that jointly models magnitude, color, and position dependence improves photometric calibration of digitized photographic plates, roughly halving bright-star errors versus the separable MYX25 method.
-
Quantifying and Attributing Power Flexibility from GPU-Heavy Data Centers
Energy-aware scheduling yields latent GPU-data-center power flexibility via cooling shifts (~$30/MWh) and job movement/reordering ($30–$3000+/MWh), larger with perfect queue foresight.
-
DiskChunGS: Large-Scale 3D Gaussian SLAM Through Chunk-Based Memory Management
Storing inactive spatial chunks of a 3D Gaussian map on disk and loading only camera-visible chunks into GPU memory lets DiskChunGS map all 11 KITTI sequences on a 24 GB GPU without memory failures.
-
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.
-
HairGS: Hair Strand Reconstruction based on 3D Gaussian Splatting
HairGS reconstructs 3D hair strands from multi-view images in about one hour by fitting 3D Gaussians, merging them into strands with distance and direction rules, and refining them against the photos.
-
Construction of Digital Terrain Maps from Multi-view Satellite Imagery using Neural Volume Rendering
Neural terrain maps reconstruct digital elevation models from multi-view satellite imagery alone, reaching near image-resolution accuracy.
-
Sequential Neural Operator Transformer for High-Fidelity Surrogates of Time-Dependent Non-linear Partial Differential Equations
S-NOT, a GRU-transformer hybrid, predicts full-field solutions of time-dependent nonlinear PDEs with lower error than Sequential DeepONet on steel solidification, 3D lug, and dogbone benchmarks.
-
Real-Time Scene Reconstruction using Light Field Probes
A probe-based renderer built from laser point clouds reconstructs a room-scale scene in real time with constant per-frame cost.
-
From images to properties: a NeRF-driven framework for granular material parameter inversion
A pipeline combining NeRF 3D reconstruction, MPM simulation, and Bayesian optimization recovers sand friction angle from rendered images with mean absolute errors between 0.64 and 1.38 degrees in synthetic tests.
Discussion (0). Sign in to comment.