REVIEW 3 major objections 5 minor 116 references
Gaussians on Fire: High-Frequency Reconstruction of Flames
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read A pipeline reconstructs high-frequency 3D flame dynamics from just three calibrated, hardware-synchronized cameras, using Gaussian splats that each carry a lifetime and a linear velocity.
desk verdict Valuable dataset and capture rig; the central 3D claim needs a held-out-view test before it is proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery has three coupled parts. First, the static/dynamic split: the background is reconstructed once with vanilla 3D Gaussian splats whose geometry is initialized by dense stereo depth and regularized by aligned monocular depth, so the under-constrained three-view problem never touches the flames. Second, the flow-fusion initialization: per-view dense optical flow is back-projected into a voxel grid, and at each voxel the unknown 3D velocity F must satisfy (F - u_i) . u_i = 0 for each camera's back-projected 2D flow u_i; because three constraints do not determine a 3D vector, a Tikhonov term with a voxel-dependent weight selects a velocity, and the resulting volume seeds
What would settle it
Run the pipeline on a synthetic fire sequence with known ground-truth 3D velocity, using only three views; measure per-Gaussian velocity error against ground truth while keeping PSNR high. If the recovered trajectories diverge substantially from ground truth even though the images match, the central claim of reconstructing flame dynamics is falsified — the flow-initialized Gaussians would be fitting appearance, not motion.
Extended reading notes
Core claim
The discovery is that a dynamic 3D Gaussian representation can reconstruct fire from extremely sparse views, provided the Gaussians' temporal parameters are initialized from observed motion rather than learned from scratch. The key move is to treat each flame Gaussian as a particle: it has a birth time, a lifetime, and a constant velocity, so its world-space trajectory is x(t) = x0 + (t - t_mu)v. This parameterization is seeded by fusing dense 2D optical flow from all three cameras onto a voxel grid; at each voxel the back-projected 2D flows constrain the unknown 3D velocity to satisfy a projection constraint for each camera, and a regularization term selects the smallest admissible velocity
Load-bearing premise
The whole method stands or falls on the 3D flow field stitched together from three 2D motion estimates: with only three cameras, each voxel is under-constrained, and the paper's own limitation section (Sec. 6) concedes that optical-flow errors in bright or low-texture flame regions propagate directly into the 3D motion field and into every Gaussian seeded from it.
Editorial extensions
If this is right
- Fire capture moves from dense camera arrays or physics simulators to three synchronized consumer cameras, lowering hardware cost and setup complexity.
- The reconstructed model can be rendered from novel viewpoints outside the training rig and at arbitrary times, because each Gaussian's position and opacity are defined for all times by its velocity and lifetime.
- Flame motion becomes an explicit, queryable quantity: every Gaussian's velocity is a number, not a learned latent, so trajectories can be exported for analysis, animation, or reactive systems.
- Sparse-view dynamic Gaussian reconstruction is substantially more stable when velocities are initialized from observed optical flow rather than randomly, per the paper's RandFlow ablation.
- Separating static background from dynamic fire during training yields not only sharper fire rendering but also geometrically plausible depth, whereas joint optimization collapses depth accuracy.
Reading between the lines
- The same split-and-seed recipe should transfer to other semi-transparent volumetric phenomena such as smoke, steam, dust, or spray; the paper only evaluates combustion by-products, not whether the velocity-carrying Gaussian form is the general mechanism.
- The explicit velocity field could be validated against combustion physics (buoyant rise profiles, turbulent spectra) as a way to separate reconstructed motion from photometric artifacts; the paper does not do this comparison.
- A testable scaling prediction follows: if the flow fusion is the active ingredient, reconstruction quality in fire regions should degrade smoothly as cameras are removed from three to two, and the regularization bias toward small velocities should become visible in the recovered trajectories; the paper reports only the three-view case.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Gaussians on Fire, a pipeline for reconstructing 4D fire from three calibrated, hardware-synchronized consumer cameras. The static background is reconstructed with vanilla 3DGS initialized from stereo depth and regularized by aligned DepthAnythingV2 monocular depth; the dynamic fire is represented with FreeTimeGS-style Gaussians carrying linear velocity, lifetime, and temporal center, initialized from a Tikhonov-regularized fusion of per-view dense optical flow into a voxel grid. The authors introduce a real dataset of 17 fire scenes at 400 Hz, a custom ESP32/LED synchronization rig, and a rolling-shutter model. They report PSNR_flame 27.46 ± 1.76 versus 18.24/25.51 for the two 4DGS baselines, RMSE_depth 0.0412 versus ~0.27, and ablations for sync, rolling shutter, flow initialization, and static pre-training.
Significance. If the central claim were firmly established, this would be a useful contribution: it would show that high-frequency dynamic fire can be reconstructed from a minimal three-camera rig and that explicit per-Gaussian velocity and lifetime parameters are a practical representation for emissive, transparent scenes. Strengths include a real captured dataset, a concrete hardware synchronization design with measured temporal precision, a rolling-shutter correction model, and ablations that isolate the main components. The code and data are promised for release, which would help reproducibility. However, the evidence as presented does not yet establish genuine 3D generalization across viewpoints, and the only quantitative geometric metric is partially circular.
major comments (3)
- [Sec. 5, Tables 1–2, Supp. Sec. 8] All quantitative metrics are computed on temporally held-out frames rendered from the same three training viewpoints. The supplementary acknowledges that 4DGS (Wu et al.) degenerates to a separate per-camera 2D video representation, and PSNR_flame/SSIM_flame would still reward such a 2D model on a temporal test. The central claim is '3D reconstruction from three views', so the evaluation must include a held-out viewpoint: for example, train on two views and test on the third, or render a calibrated novel camera path and compare against a held-out view quantitatively. The qualitative novel-view video is a useful start but does not separate true view interpolation from appearance memory.
- [Sec. 3.1 Eq. (3), Sec. 5 Table 1] RMSE_depth is computed against DepthAnythingV2 predictions after linear alignment, while the static reconstruction is trained with λ_depth increased by a factor of 100 against those same aligned predictions. The large gap between 0.0412 and ~0.27 therefore partly measures agreement with the training regularizer, not geometric accuracy. For the transparent/emissive fire region, monocular depth is an unreliable geometric reference. Please report depth against a geometrically grounded reference not used in training, or at least evaluate with the depth regularization switched off and clearly state the change in the comparison.
- [Sec. 3.2 Eqs. (5)–(8), Table 2 (RandFlow)] The fused 3D flow is underdetermined: with m=3 cameras, each voxel provides only three scalar projection constraints on a 3D vector, and closure relies on Tikhonov regularization with hand-set α0. The RandFlow ablation retains PSNR_flame 27.01 versus 27.46, showing that the photometric loss largely compensates for incorrect velocity initialization. Since the paper claims reconstruction of flame dynamics, please provide a direct evaluation of the velocity field—e.g., predicted 3D trajectories projected into a held-out view and compared with tracked flame features, or comparison with a physically motivated vertical-motion prior. The paper's own limitation statement in Sec. 6 concedes that optical-flow errors propagate into the 3D motion field, so an independent validation of the flow component is needed.
minor comments (5)
- [Sec. 5] Typographical issues: 'RSMEdepth' should be 'RMSE_depth', 'archives' should be 'achieves', and 'high offerfitting' should be 'high overfitting'.
- [Eq. (5)] The notation π^{-1}_i(..., x) is nonstandard, and the expression f_i(π(x)) + π_i(x) mixes 2D and 3D operands. Please rewrite with explicit coordinates or provide a reference for this back-projection operation.
- [Secs. 3.3–3.4] The measured readout time is stated as 2.85 μs per line in Sec. 3.3 and t_row ≈ 3 μs in Sec. 3.4; clarify whether these are the same quantity or different measurements.
- [References] References [91] and [92] appear to be the same paper (Wu et al., CVPR 2024) listed twice. Please deduplicate.
- [Supp. Sec. 10] The description 'increased the depth loss by a factor of 100, yielding an exponentially decreasing contribution with an initial value of 100 and a final value of 1' is ambiguous about the scheduling. Specify the exact annealing schedule.
Circularity Check
Depth evaluation is circular: RMSE_depth compares against the same monocular-depth predictor used as the static-scene regularizer; photometric reconstruction itself is tested on held-out frames.
-
fitted input called prediction
[Sec. 3.1 (Eq. 3) and Sec. 5 (Tab. 1, Fig. 9)]
"During training, the aligned monocular depths D′mono are incorporated as a regularization term ... Ldepth = λdepth/wh Σ ||D′mono(i,j)−Drender(i,j)||1 ... We increase the default regularization parameter value by a factor of 100. ... To evaluate the plausibility of the rendered depth, we compare it against a predicted monocular depth after linear alignment and report this score as RSMEdepth."
The reported RMSE_depth metric uses DepthAnythingV2 monocular depth (after linear alignment) as the reference. That is the same signal used as the supervision target in Eq. 3, with λ_depth raised by a factor of 100 during static-scene training. A low RMSE_depth therefore largely indicates agreement with the training regularizer, not agreement with independent 3D geometry. For the dynamic fire region, monocular depth is not a reliable geometric reference, and for the static region the low error is partly by construction. Thus the paper's quantitative evidence for 'geometrically plausible depth' and multi-view 3D consistency is partly circular.
full rationale
The photometric reconstruction claim is not circular: PSNR_flame/SSIM_flame are computed on temporally held-out frames (every 8th frame), which is a legitimate test of temporal generalization, and the flow-initialization is ablated (RandFlow) rather than assumed. The paper also does not rely on a load-bearing self-citation chain: MEMFOF, DepthAnythingV2, and FreeTimeGS are external methods, not author-defined constraints. However, the only quantitative geometry metric, RMSE_depth, is evaluated against the same monocular-depth predictor used to regularize the static scene, making that specific evaluation partially self-referential. The paper's supplementary even notes that one baseline degenerates to a per-camera 2D video, but no quantitative experiment separates the proposed method from a more graceful version of the same failure; this is a validity concern rather than a derivation-level circularity. Overall, the central contribution retains independent content, but the depth evidence is weakened by construction, yielding a moderate circularity score.
Assumptions & free parameters
free parameters (6)
- α0 (Tikhonov regularization strength, Eq. 7) =
not reported
- λ_depth scaling factor =
100 -> 1 (exponential decay)
- Gaussian lifespan initialization =
4 x frame time ~ 10 ms
- Maximum screen-space motion cap =
4 px/ms
- Voxel size s (flow-fusion grid) =
not reported
- Velocity spatial learning-rate reduction =
x10^-2
assumptions (6)
- domain assumption Minimum-intensity projection over time removes fire: flames are locally brighter than their backdrop and the background is static during the sequence.
- domain assumption The 3-view projection constraint (F-u_i)·u_i = 0 together with Tikhonov regularization recovers a usable 3D flow field (Eqs. 5-8).
- domain assumption FreeTimeGS linear-velocity plus Gaussian-lifetime representation is expressive enough for flame dynamics.
- domain assumption DepthAnythingV2 monocular depths, after per-camera affine alignment, are valid regularizers and valid references for evaluating rendered depth.
- domain assumption Rolling-shutter delay is dominated by line readout (2.85 μs/line) and is locally linear; the denominator 1 - ∇t·v ≈ 0.989 can be dropped.
- domain assumption COLMAP camera poses estimated on fire-free static frames remain valid for the dynamic frames.
invented entities (1)
-
ESP32 LED synchronization rig (Gray-code frame counter + 5 toggling COB LED strips)
independent evidence
Cite this review
Pith. "Pith review of Gaussians on Fire: High-Frequency Reconstruction of Flames." pith.science (2026). https://pith.science/paper/45FRVK6L
@misc{pith2026251122459,
author = {Pith},
title = {Pith review of: Gaussians on Fire: High-Frequency Reconstruction of Flames},
year = {2026},
howpublished = {\url{https://pith.science/paper/45FRVK6L}},
note = {Machine review of arXiv:2511.22459}
}
read the original abstract
We propose a method to reconstruct dynamic fire in 3D from a limited set of camera views with a Gaussian-based spatiotemporal representation. Capturing and reconstructing fire and its dynamics is highly challenging due to its volatile nature, transparent quality, and multitude of high-frequency features. Despite these challenges, we aim to reconstruct fire from only three views, which consequently requires solving for under-constrained geometry. We solve this by separating the static background from the dynamic fire region by combining dense multi-view stereo images with monocular depth priors. The fire is initialized as a 3D flow field, obtained by fusing per-view dense optical flow projections. To capture the high frequency features of fire, each 3D Gaussian encodes a lifetime and linear velocity to match the dense optical flow. To ensure sub-frame temporal alignment across cameras we employ a custom hardware synchronization pattern -- allowing us to reconstruct fire with affordable commodity hardware. Our quantitative and qualitative validations across numerous reconstruction experiments demonstrate robust performance for diverse and challenging real fire scenarios.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Per-gaussian embedding- based deformation for deformable 3d gaussian splatting
Jeongmin Bae, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, and Youngjung Uh. Per-gaussian embedding- based deformation for deformable 3d gaussian splatting
-
[2]
Narasimhan
Aayush Bansal, Minh V o, Yaser Sheikh, Deva Ramanan, and Srinivasa G. Narasimhan. 4d visualization of dynamic events from unconstrained multi-view videos. InCVPR,
-
[3]
Memfof: High-resolution training for memory-efficient multi-frame optical flow estimation
Vladislav Bargatin, Egor Chistov, Alexander Yakovenko, and Dmitriy Vatolin. Memfof: High-resolution training for memory-efficient multi-frame optical flow estimation. In ICCV, 2025. 2, 4
2025
-
[4]
SURF: Speeded up robust features
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. SURF: Speeded up robust features. InECCV, 2006. 2
2006
-
[5]
Black and Padmanabhan Anandan
Michael J. Black and Padmanabhan Anandan. A framework for the robust estimation of optical flow. InICCV, 1993. 2
1993
-
[6]
Patchmatch stereo-stereo matching with slanted support windows
Michael Bleyer, Christoph Rhemann, and Carsten Rother. Patchmatch stereo-stereo matching with slanted support windows. InBmvc, pages 1–11, 2011. 4
2011
-
[7]
Depth pro: Sharp monocular metric depth in less than a second
Alexey Bochkovskiy, Ama ¨el Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. InICLR, pages 75602–75637, 2025. 2
2025
-
[8]
Deepdeform: Learning non-rigid RGB- D reconstruction with semi-supervised data
Aljaz Bozic, Michael Zollh ¨ofer, Christian Theobalt, and Matthias Nießner. Deepdeform: Learning non-rigid RGB- D reconstruction with semi-supervised data. InCVPR,
Show all 116 references
-
[9]
Recovering non-rigid 3d shape from image streams
Christoph Bregler, Aaron Hertzmann, and Henning Bier- mann. Recovering non-rigid 3d shape from image streams. InCVPR, 2000. 3
2000
-
[10]
High accuracy optical flow estimation based on a theory for warping
Thomas Brox, Andr ´es Bruhn, Nils Papenberg, and Joachim Weickert. High accuracy optical flow estimation based on a theory for warping. InECCV, 2004. 2
2004
-
[11]
Large displacement optical flow
Thomas Brox, Christoph Bregler, and Jitendra Malik. Large displacement optical flow. InCVPR, 2009. 2
2009
-
[12]
Immersive light field video with a layered mesh representation.ACM TOG, 39(4), 2020
Michael Broxton, John Flynn, Ryan Overbeck, Daniel Er- ickson, Peter Hedman, Matthew Duvall, Jason Dourgarian, Jay Busch, Matt Whalen, and Paul Debevec. Immersive light field video with a layered mesh representation.ACM TOG, 39(4), 2020. 3
2020
-
[13]
Hexplane: A fast representa- tion for dynamic scenes
Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes. InCVPR, pages 130–141, 2023. 3
2023
-
[14]
Black, Otmar Hilliges, and Andreas Geiger
Xu Chen, Yufeng Zheng, Michael J. Black, Otmar Hilliges, and Andreas Geiger. SNARF: Differentiable forward skin- ning for animating non-rigid neural implicit shapes. In ICCV, 2021. 3
2021
-
[15]
Physics informed neural fields for smoke reconstruction with sparse data.ACM TOG, 41(4):119:1–119:14, 2022
Mengyu Chu, Lingjie Liu, Quan Zheng, Erik Franz, Hans- Peter Seidel, Christian Theobalt, and Rhaleb Zayer. Physics informed neural fields for smoke reconstruction with sparse data.ACM TOG, 41(4):119:1–119:14, 2022. 1, 2, 3
2022
-
[16]
A simple prior- free method for non-rigid structure-from-motion factoriza- tion.IJCV, 107(2), 2014
Yuchao Dai, Hongdong Li, and Mingyi He. A simple prior- free method for non-rigid structure-from-motion factoriza- tion.IJCV, 107(2), 2014. 3
2014
-
[17]
Superpoint: Self-supervised interest point detection and description
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. Superpoint: Self-supervised interest point detection and description. InCVPRW, 2018. 2
2018
-
[18]
Tap-vid: A benchmark for track- ing any point in a video
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adri `a Re- casens, Lucas Smaira, Yusuf Aytar, Jo˜ao Carreira, Andrew Zisserman, and Yi Yang. Tap-vid: A benchmark for track- ing any point in a video. InNeurIPS, 2022. 2
2022
-
[19]
Tapir: Tracking any point with per-frame initialization and temporal refinement
Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Jo˜ao Carreira, and Andrew Zisserman. Tapir: Tracking any point with per-frame initialization and temporal refinement. InICCV, 2023
2023
-
[20]
Bootstap: Bootstrapped training for tracking-any-point
Carl Doersch, Pauline Luc, Yi Yang, Dilara Gokay, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ignacio Rocco, Ross Goroshin, Joao Carreira, et al. Bootstap: Bootstrapped training for tracking-any-point. InACCV, 2024. 2
2024
-
[21]
Degtyarev, Philip L
Mingsong Dou, Sameh Khamis, Yury G. Degtyarev, Philip L. Davidson, Sean Fanello, Adarsh Kowdle, Ser- gio Orts, Christoph Rhemann, David Kim, Jonathan Taylor, Pushmeet Kohli, Vladimir Tankovich, and Shahram Izadi. Fusion4d: Real-time volumetric capture of dynamic scenes. ACM TO...
2016
-
[22]
Tenen- baum, and Jiajun Wu
Yilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenen- baum, and Jiajun Wu. Neural radiance flow for 4d view synthesis and video processing. InICCV, 2021. 3
2021
-
[23]
4d-rotor gaussian splat- ting: Towards efficient novel view synthesis for dynamic scenes
Yuanxing Duan, Fangyin Wei, Qiyu Dai, Yuhang He, Wen- zheng Chen, and Baoquan Chen. 4d-rotor gaussian splat- ting: Towards efficient novel view synthesis for dynamic scenes. InSIGGRAPH 2024 Conference Papers, 2024. 3
2024
-
[24]
MASt3r-sfm: a fully-integrated solution for unconstrained structure-from-motion
Bardienus Pieter Duisterhof, Lojze Zust, Philippe Weinza- epfel, Vincent Leroy, Yohann Cabon, and Jerome Revaud. MASt3r-sfm: a fully-integrated solution for unconstrained structure-from-motion. InInternational Conference on 3D Vision 2025, 2025. 2, 3
2025
-
[25]
Black, Trevor Darrell, and Angjoo Kanazawa
Haiwen Feng*, Junyi Zhang*, Qianqian Wang, Yufei Ye, Pengcheng Yu, Michael J. Black, Trevor Darrell, and Angjoo Kanazawa. St4rtrack: Simultaneous 4d reconstruc- tion and tracking in the world. InICCV, 2025. 2, 3
2025
-
[26]
K- planes: Explicit radiance fields in space, time, and appear- ance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warbæk, Benjamin Recht, and Angjoo Kanazawa. K- planes: Explicit radiance fields in space, time, and appear- ance. InCVPR, 2023. 3
2023
-
[27]
Dynamic view synthesis from dynamic monocu- lar video
Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocu- lar video. InICCV, 2021. 3
2021
-
[28]
Fluid- nexus: 3d fluid reconstruction and prediction from a single video
Yue Gao, Hong-Xing Yu, Bo Zhu, and Jiajun Wu. Fluid- nexus: 3d fluid reconstruction and prediction from a single video. InCVPR, 2025. 2, 3 9
2025
-
[29]
Fleet, Saurabh Sax- ena, and Andrea Tagliasacchi
Lily Goli, Sara Sabour, Mark Matthews, Marcus Brubaker, Dmitry Lagun, Alec Jacobson, David J. Fleet, Saurabh Sax- ena, and Andrea Tagliasacchi. RoMo: Robust motion seg- mentation improves structure from motion.ICCV, 2025. 2, 3
2025
-
[30]
Harley, Zhaoyuan Fang, and Katerina Fragki- adaki
Adam W. Harley, Zhaoyuan Fang, and Katerina Fragki- adaki. Particle video revisited: Tracking through occlusions using point trajectories. InECCV, 2022. 2
2022
-
[31]
Berthold K. P. Horn and Brian G. Schunck. Determining optical flow.Artificial Intelligence, 1981. 2
1981
-
[32]
2d gaussian splatting for geometrically ac- curate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2d gaussian splatting for geometrically ac- curate radiance fields. InSIGGRAPH 2024 Conference Pa- pers, 2024. 3
2024
-
[33]
Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes
Yi-Hua Huang, Yang-Tian Sun, Ziyi Yang, Xiaoyang Lyu, Yan-Pei Cao, and Xiaojuan Qi. Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes. InCVPR, pages 4220–4230, 2024. 2, 3
2024
-
[34]
Self-supervised monocular scene flow estimation
Junhwa Hur and Stefan Roth. Self-supervised monocular scene flow estimation. InCVPR, 2020. 2, 3
2020
-
[35]
Ivo Ihrke and Marcus A. Magnor. Image-based tomo- graphic reconstruction of flames. InProceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation (SCA), pages 365–373, 2004. 2, 3
2004
-
[36]
V olumedeform: Real-time volumetric non-rigid reconstruction
Matthias Innmann, Michael Zollh ¨ofer, Matthias Nießner, Christian Theobalt, and Marc Stamminger. V olumedeform: Real-time volumetric non-rigid reconstruction. InECCV,
-
[37]
Humanrf: High-fidelity neural radiance fields for humans in motion.ACM TOG, 2023
Mustafa Isik, Martin R ¨unz, Markos Georgopoulos, Taras Khakhulin, Jonathan Starck, Lourdes Agapito, and Matthias Nießner. Humanrf: High-fidelity neural radiance fields for humans in motion.ACM TOG, 2023. 3
2023
-
[38]
Learning to estimate hidden motions with global motion aggregation
Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley. Learning to estimate hidden motions with global motion aggregation. InICCV, 2021. 2
2021
-
[39]
Learning optical flow from a few matches
Shihao Jiang, Yao Lu, Hongdong Li, and Richard Hartley. Learning optical flow from a few matches. InCVPR, 2021. 2
2021
-
[40]
Co- tracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker: It is better to track together. InECCV, 2024. 2
2024
-
[41]
3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics (SIGGRAPH), 42(4):1–14, 2023
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics (SIGGRAPH), 42(4):1–14, 2023. 2, 3, 4, 1
2023
-
[42]
Michels, S¨oren Pirk, and Wojtek Palubicki
Andrzej Kokosza, Helge Wrede, Daniel Gonzalez Esparza, Milosz Makowski, Daoming Liu, Dominik L. Michels, S¨oren Pirk, and Wojtek Palubicki. Scintilla: Simulating combustible vegetation for wildfires.ACM Trans. Graph., 43(4), 2024. 2
2024
-
[43]
Robust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang. Robust consistent video depth estimation. InCVPR, 2021. 3
2021
-
[44]
Tapvid-3d: A benchmark for tracking any point in 3d
Skanda Koppula, Ignacio Rocco, Yi Yang, Joe Heyward, Joao Carreira, Andrew Zisserman, Gabriel Brostow, and Carl Doersch. Tapvid-3d: A benchmark for tracking any point in 3d. InNeurIPS, 2024. 2, 3
2024
-
[45]
DynMF: Neural motion factorization for real-time dynamic view synthesis with 3d gaussian splatting
Agelos Kratimenos, Jiahui Lei, and Kostas Daniilidis. DynMF: Neural motion factorization for real-time dynamic view synthesis with 3d gaussian splatting. InECCV, 2024. 3
2024
-
[46]
Monoc- ular dense 3d reconstruction of a complex dynamic scene from two perspective frames
Suryansh Kumar, Yuchao Dai, and Hongdong Li. Monoc- ular dense 3d reconstruction of a complex dynamic scene from two perspective frames. InICCV, 2017. 3
2017
-
[47]
Harley, Leonidas Guibas, and Kostas Daniilidis
Jiahui Lei, Yijia Weng, Adam W. Harley, Leonidas Guibas, and Kostas Daniilidis. Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds. InCVPR, 2025. 3
2025
-
[48]
Grounding image matching in 3d with mast3r, 2024
Vincent Leroy, Yohann Cabon, and Jerome Revaud. Grounding image matching in 3d with mast3r, 2024. 2, 3
2024
-
[49]
Tava: Template-free animatable volumetric actors
Ruilong Li, Julian Tanke, Minh V o, Michael Zollh ¨ofer, J¨urgen Gall, Angjoo Kanazawa, and Christoph Lassner. Tava: Template-free animatable volumetric actors. In ECCV, 2022. 3
2022
-
[50]
Neural 3d video synthesis from multi-view video
Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. In CVPR, 2022. 3
2022
-
[51]
Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker, Noah Snavely, Ce Liu, and William T. Freeman. Learning the depths of moving people by watching frozen people. In CVPR, 2019. 3
2019
-
[52]
Neural scene flow fields for space-time view syn- thesis of dynamic scenes
Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view syn- thesis of dynamic scenes. InCVPR, 2021. 3
2021
-
[53]
Dynibar: Neural dynamic image-based rendering
Zhengqi Li, Qianqian Wang, Forrester Cole, Richard Tucker, and Noah Snavely. Dynibar: Neural dynamic image-based rendering. InCVPR, 2023. 3
2023
-
[54]
Spacetime gaussian feature splatting for real-time dynamic view syn- thesis
Zhan Li, Zhang Chen, Zhong Li, and Yi Xu. Spacetime gaussian feature splatting for real-time dynamic view syn- thesis. InCVPR, pages 8508–8520, 2024. 3
2024
-
[55]
SIFT flow: Dense correspondence across scenes and its applications
Ce Liu, Jenny Yuen, and Antonio Torralba. SIFT flow: Dense correspondence across scenes and its applications. IEEE TPAMI, 2011. 2
2011
-
[56]
Modgs: Dy- namic gaussian splatting from casually-captured monocular videos
Qingming Liu, Yuan Liu, Jiepeng Wang, Xianqiang Lv, Peng Wang, Wenping Wang, and Junhui Hou. Modgs: Dy- namic gaussian splatting from casually-captured monocular videos. InICLR, 2025. 3
2025
-
[57]
David G. Lowe. Distinctive image features from scale- invariant keypoints.IJCV, 2004. 2
2004
-
[58]
Align3r: Aligned monocular depth estimation for dynamic videos
Jiahao Lu, Tianyu Huang, Peng Li, Zhiyang Dou, Cheng Lin, Zhiming Cui, Zhen Dong, Sai-Kit Yeung, Wenping Wang, and Yuan Liu. Align3r: Aligned monocular depth estimation for dynamic videos. InCVPR, 2025. 3
2025
-
[59]
Lucas and Takeo Kanade
Bruce D. Lucas and Takeo Kanade. An iterative image registration technique with an application to stereo vision. InInternational Joint Conference on Artificial Intelligence (IJCAI), 1981. 2
1981
-
[60]
Dynamic 3d gaussians: Tracking by per- sistent dynamic view synthesis
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by per- sistent dynamic view synthesis. In3DV, pages 800–809,
-
[61]
Consistent video depth estimation
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. Consistent video depth estimation. ACM TOG, 39(4), 2020. 3
2020
-
[62]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. InECCV, pages 405–421. Springer, 2020. 2, 3
2020
-
[63]
MFT: Long- term tracking of every pixel
Michal Neoral, Jonas Serych, and Jiri Matas. MFT: Long- term tracking of every pixel. InWACV, 2024. 2
2024
-
[64]
Newcombe, Dieter Fox, and Steven M
Richard A. Newcombe, Dieter Fox, and Steven M. Seitz. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. InCVPR, 2015. 3
2015
-
[65]
Delta: Dense efficient long-range 3d tracking for any video
Tuan Duc Ngo, Peiye Zhuang, Chuang Gan, Evange- los Kalogerakis, Sergey Tulyakov, Hsin-Ying Lee, and Chaoyang Wang. Delta: Dense efficient long-range 3d tracking for any video. InICLR, 2025. 2, 3
2025
-
[66]
C3DPO: Canonical 3d pose networks for non-rigid structure from motion
David Novotny, Nikhila Ravi, Benjamin Graham, Natalia Neverova, and Andrea Vedaldi. C3DPO: Canonical 3d pose networks for non-rigid structure from motion. InICCV,
-
[67]
Barron, Sofien Bouaziz, Dan B
Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B. Goldman, Steven M. Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. InICCV, 2021. 3
2021
-
[68]
Barron, Sofien Bouaziz, Dan B
Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B. Goldman, Ricardo Martin- Brualla, and Steven M. Seitz. Hypernerf: A higher- dimensional representation for topologically varying neural radiance fields.ACM TOG, 2021
2021
-
[69]
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. InCVPR, pages 10318–10327, 2021. 2, 3
2021
-
[70]
Dense monocular depth estimation in complex dy- namic scenes
Rene Ranftl, Vibhav Vineet, Qifeng Chen, and Vladlen Koltun. Dense monocular depth estimation in complex dy- namic scenes. InCVPR, 2016. 3
2016
-
[71]
Sudderth, and Jan Kautz
Zhile Ren, Orazio Gallo, Deqing Sun, Ming-Hsuan Yang, Erik B. Sudderth, and Jan Kautz. A fusion approach for multi-frame optical flow estimation. InWACV, 2019. 2
2019
-
[72]
Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary R. Bradski. ORB: An efficient alternative to SIFT or SURF. InICCV, 2011. 2
2011
-
[73]
Video pop-up: Monocular 3d reconstruction of dynamic scenes
Chris Russell, Rui Yu, and Lourdes Agapito. Video pop-up: Monocular 3d reconstruction of dynamic scenes. InECCV,
-
[74]
R3d3: Dense 3d reconstruction of dynamic scenes from multiple cameras
Aron Schmied, Tobias Fischer, Martin Danelljan, Marc Pollefeys, and Fisher Yu. R3d3: Dense 3d reconstruction of dynamic scenes from multiple cameras. InICCV, pages 3216–3226, 2023. 1
2023
-
[75]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 2
2016
-
[76]
A vote-and-verify strategy for fast spatial verification in image retrieval
Johannes Lutz Sch ¨onberger, True Price, Torsten Sattler, Jan-Michael Frahm, and Marc Pollefeys. A vote-and-verify strategy for fast spatial verification in image retrieval. In Asian Conference on Computer Vision (ACCV), 2016. 4
2016
-
[77]
Pixelwise view selection for unstructured multi-view stereo
Johannes Lutz Sch ¨onberger, Enliang Zheng, Marc Polle- feys, and Jan-Michael Frahm. Pixelwise view selection for unstructured multi-view stereo. InEuropean Conference on Computer Vision (ECCV), 2016. 2
2016
-
[78]
Videoflow: Exploiting temporal cues for multi-frame optical flow estimation
Xiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Videoflow: Exploiting temporal cues for multi-frame optical flow estimation. In ICCV, 2023. 2
2023
-
[79]
Nerf- player: A streamable dynamic scene representation with decomposed neural radiance fields.IEEE TVCG, 29(5),
Liangchen Song, Anpei Chen, Zhong Li, Zhang Chen, Lele Chen, Junsong Yuan, Yi Xu, and Andreas Geiger. Nerf- player: A streamable dynamic scene representation with decomposed neural radiance fields.IEEE TVCG, 29(5),
-
[80]
Track everything everywhere fast and robustly
Yunzhou Song, Jiahui Lei, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. Track everything everywhere fast and robustly. InECCV, 2024. 2
2024
-
[81]
Marigold-dc: Zero-shot monocular depth com- pletion with guided diffusion
Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, and Anton Obukhov. Marigold-dc: Zero-shot monocular depth com- pletion with guided diffusion. InICCV, 2025. 2
2025
-
[82]
Neural trajectory fields for dynamic novel view syn- thesis
Chaoyang Wang, Ben Eckart, Simon Lucey, and Orazio Gallo. Neural trajectory fields for dynamic novel view syn- thesis. InarXiv preprint arXiv:2105.05994, 2021. 3
2021 arXiv
-
[83]
Ode-gs: Latent odes for dynamic scene extrapolation with 3d gaussian splatting
Daniel Wang, Patrick Rim, Tian Tian, Dong Lao, Alex Wong, and Ganesh Sundaramoorthi. Ode-gs: Latent odes for dynamic scene extrapolation with 3d gaussian splatting. arXiv preprint arXiv:2506.05480, 2025. 3
2025 arXiv
-
[84]
Fourier plenoctrees for dynamic radiance field rendering in real-time
Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. Fourier plenoctrees for dynamic radiance field rendering in real-time. InCVPR, 2022. 3
2022
-
[85]
Tracking everything everywhere all at once
Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely. Tracking everything everywhere all at once. In ICCV, 2023. 2
2023
-
[86]
Efros, and Angjoo Kanazawa
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, and Angjoo Kanazawa. Continuous 3d perception model with persistent state. InCVPR, 2025. 3
2025
-
[87]
Dust3r: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. InCVPR, 2024. 2, 3
2024
-
[88]
Freetimegs: Free gaussian primitives at any- time anywhere for dynamic scene reconstruction
Yifan Wang, Peishan Yang, Zhen Xu, Jiaming Sun, Zhan- hua Zhang, Chen Yong, Hujun Bao, Sida Peng, and Xi- aowei Zhou. Freetimegs: Free gaussian primitives at any- time anywhere for dynamic scene reconstruction. InCVPR,
-
[89]
Srinivasan, Jonathan T
Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron, and Ira Kemelmacher-Shlizerman. Hu- mannerf: Free-viewpoint rendering of moving people from monocular video. InCVPR, 2022. 3
2022
-
[90]
Wagner, Sarker Miraz Mahfuz, Wo- jtek Pałubicki, Dominik L
Helge Wrede, Anton R. Wagner, Sarker Miraz Mahfuz, Wo- jtek Pałubicki, Dominik L. Michels, and S ¨oren Pirk. Fire- X: Extinguishing fire with stoichiometric heat release.ACM TOG, 44(6):1–17, 2025. 2
2025
-
[91]
4d gaussian splatting for real-time dynamic 11 scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xi- aopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xing- gang Wang. 4d gaussian splatting for real-time dynamic 11 scene rendering. InProceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pag...
2024
-
[92]
4d gaussian splatting for real-time dynamic scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xi- aopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xing- gang Wang. 4d gaussian splatting for real-time dynamic scene rendering. InCVPR, pages 20310–20320, 2024. 2, 3, 6, 7, 8
2024
-
[93]
3dgut: Enabling dis- torted cameras and secondary rays in gaussian splatting
Qi Wu, Janick Martinez Esturo, Ashkan Mirzaei, Nicolas Moenne-Loccoz, and Zan Gojcic. 3dgut: Enabling dis- torted cameras and secondary rays in gaussian splatting. CVPR, 2025. 3
2025
-
[94]
Space-time neural irradiance fields for free-viewpoint video
Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. InCVPR, 2021. 3
2021
-
[95]
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. InCVPR, 2024. 2, 3
2024
-
[96]
Spatialtrackerv2: 3d point tracking made easy
Yuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev, Yuri Makarov, Bingyi Kang, Xing Zhu, Hujun Bao, Yujun Shen, and Xiaowei Zhou. Spatialtrackerv2: 3d point tracking made easy. InICCV, 2025
2025
-
[97]
4dgt: Learning a 4d gaus- sian transformer using real-world monocular videos
Zhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou, Richard Newcombe, and Zhaoyang Lv. 4dgt: Learning a 4d gaus- sian transformer using real-world monocular videos. In NeurIPS, 2025. 2, 3
2025
-
[98]
Freeman, and Ce Liu
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic, Forrester Cole, Huiwen Chang, Deva Ramanan, William T. Freeman, and Ce Liu. Lasr: Learning artic- ulated shape reconstruction from a monocular video. In CVPR, 2021. 3
2021
-
[99]
Viser: Video-specific surface embeddings for articulated 3d shape reconstruction
Gengshan Yang, Deqing Sun, Varun Jampani, Daniel Vla- sic, Forrester Cole, Ce Liu, and Deva Ramanan. Viser: Video-specific surface embeddings for articulated 3d shape reconstruction. InNeurIPS, 2021
2021
-
[100]
Banmo: Build- ing animatable 3d neural models from many casual videos
Gengshan Yang, Minh V o, Natalia Neverova, Deva Ra- manan, Andrea Vedaldi, and Hanbyul Joo. Banmo: Build- ing animatable 3d neural models from many casual videos. InCVPR, 2022. 3
2022
-
[101]
Depth anything: Un- leashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Ji- ashi Feng, and Hengshuang Zhao. Depth anything: Un- leashing the power of large-scale unlabeled data. InCVPR,
-
[102]
Depth any- thing v2
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiao- gang Xu, Jiashi Feng, and Hengshuang Zhao. Depth any- thing v2. InNeurIPS, Red Hook, NY , USA, 2024. Curran Associates Inc. 2, 4, 7
2024
-
[103]
Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction
Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction. In CVPR, 2024. 3
2024
-
[104]
Real- time photorealistic dynamic scene representation and ren- dering with 4d gaussian splatting
Zeyu Yang, Hongye Yang, Zijie Pan, and Li Zhang. Real- time photorealistic dynamic scene representation and ren- dering with 4d gaussian splatting. InICLR, 2024. 3, 6, 7, 8, 1, 4
2024
-
[105]
Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera
Jae Shin Yoon, Kihwan Kim, Orazio Gallo, Hyun Soo Park, and Jan Kautz. Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera. In CVPR, 2020. 3
2020
-
[106]
Freeman, and Tali Dekel
Zhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman, and Tali Dekel. Consistent depth of moving objects in video.ACM TOG, 40(4), 2021. 3
2021
-
[107]
Zhoutong Zhang, Forrester Cole, Zhengqi Li, Michael Ru- binstein, Noah Snavely, and William T. Freeman. Structure and motion from casual videos. InECCV, 2022. 3
2022
-
[108]
Harley, Bokui Shen, Gordon Wet- zstein, and Leonidas J
Yang Zheng, Adam W. Harley, Bokui Shen, Gordon Wet- zstein, and Leonidas J. Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. InICCV,
-
[109]
Loop, Christian Theobalt, and Marc Stamminger
Michael Zollh ¨ofer, Matthias Nießner, Shahram Izadi, Christoph Rhemann, Christopher Zach, Matthew Fisher, Chenglei Wu, Andrew William Fitzgibbon, Charles T. Loop, Christian Theobalt, and Marc Stamminger. Real-time non-rigid reconstruction using an rgb-d camera.ACM TOG, 33, 20...
2014
-
[110]
In the video, we show the rendered results for different kinds of trajecto- ries, leaving the camera poses used for training
Novel View Trajectories To provide a qualitative demonstration of the capability of our model to generalize to novel views, we provide the videoresults and pipeline.mp4. In the video, we show the rendered results for different kinds of trajecto- ries, leaving the camera poses ...
-
[111]
For one scene, this video shows the relevant preprocessing and op- timization steps
Pipeline Visualization In the second half of the same video file we present a visu- alization of our proposed reconstruction pipeline. For one scene, this video shows the relevant preprocessing and op- timization steps. For each stage, we show the correspond- ing view for each...
-
[112]
To account for the undercon- strained scene geometry, we increased the depth loss by a factor of100
Implementation Details For the static scene optimization, we applied the original 3DGS implementation [41]. To account for the undercon- strained scene geometry, we increased the depth loss by a factor of100. To prevent the introduction of floating ar- tifacts in the backgroun...
-
[113]
Due to size limitations of the supplementary ma- terial, we provide only a single short example scene in the datasubdirectory
Code We include the Python source code for preprocessing, as well as for the static and dynamic optimization, in thecode directory. Due to size limitations of the supplementary ma- terial, we provide only a single short example scene in the datasubdirectory. The source code co...
-
[114]
With this setup, training a single scene takes approxi- mately30 minat a GPU memory usage of8 GBand a sys- tem memory usage of2 GB
Compute Hardware For training, we used a system with a 16-core AMD Ryzen Threadripper PRO 3955, an NVIDIA RTX A6000 with 48 GBof GPU memory, as well as128 GBof system mem- ory. With this setup, training a single scene takes approxi- mately30 minat a GPU memory usage of8 GBand ...
-
[115]
We provide a more detailed overview of all captured scene in Tab
Dataset Description For the validation of our proposed method, we captured a diverse dataset of different fuels and backgrounds. We provide a more detailed overview of all captured scene in Tab. 3. Each video has a duration of15 s, a resolution of 1280×720, and a frame rate of400 Hz
-
[116]
5 only reports the metrics averaged over all scenes, we provide the more detailed per-scene results of our pro- posed method in Tab
Per-Scene Results As Sec. 5 only reports the metrics averaged over all scenes, we provide the more detailed per-scene results of our pro- posed method in Tab. 4. The per-scene results for the two baselines we compared our method against are shown in Tabs. 5 and 6, whereas the ...
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.