REVIEW 9 cited by
DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Understanding and modeling lighting effects are fundamental tasks in computer vision and graphics. Classic physically-based rendering (PBR) accurately simulates the light transport, but relies on precise scene representations--explicit 3D geometry, high-quality material properties, and lighting conditions--that are often impractical to obtain in real-world scenarios. Therefore, we introduce DiffusionRenderer, a neural approach that addresses the dual problem of inverse and forward rendering within a holistic framework. Leveraging powerful video diffusion model priors, the inverse rendering model accurately estimates G-buffers from real-world videos, providing an interface for image editing tasks, and training data for the rendering model. Conversely, our rendering model generates photorealistic images from G-buffers without explicit light transport simulation. Experiments demonstrate that DiffusionRenderer effectively approximates inverse and forwards rendering, consistently outperforming the state-of-the-art. Our model enables practical applications from a single video input--including relighting, material editing, and realistic object insertion.
Forward citations
Cited by 9 Pith papers
-
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Tracking plus world-position maps in a neural G-buffer outperform depth as a geometric condition for reference-guided video diffusion rendering on a 68-clip synthetic benchmark.
-
Video Generation Models are General-Purpose Vision Learners
A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.
-
LuxDiT: Lighting Estimation with Video Diffusion Transformer
A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.
-
InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance Modeling
InvRGB+L jointly estimates visible and LiDAR albedo with a physics-based specular LiDAR model and cross-modal consistency losses, improving inverse rendering and LiDAR intensity simulation for urban and indoor scenes.
-
UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting
Jointly predicting albedo and relit appearance with one video-diffusion pass improves relighting fidelity and generalization over two-stage inverse-plus-forward pipelines.
-
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
Post-trained Cosmos world models generate controllable multi-view driving videos and LiDAR; augmenting real AV training data with these synthetic clips improves downstream perception and policy metrics, especially in ...
-
MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation
A feed-forward two-stage compositing framework that harmonizes inserted objects across views using a Hilbert-ordered Gaussian color mapping, trained and evaluated on a new 480k-scene synthetic dataset.
-
Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling
A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.
-
StableIntrinsic: Detail-preserving One-step Diffusion Model for Multi-view Material Estimation
StableIntrinsic estimates albedo, roughness, and metallic maps from multi-view RGB images in a single diffusion step, achieving higher PSNR and lower MSE than prior multi-step diffusion methods.
Discussion (0). Continue with ORCID to comment.