REVIEW 3 major objections 5 minor 72 references
MGStream: Motion-aware 3D Gaussian for Streamable Dynamic Scene Reconstruction
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read MGStream claims that streaming dynamic scenes can be compressed and stabilized by updating only the Gaussians that move.
desk verdict MGStream is a credible incremental step for streamable dynamic NVS: motion-mask-guided deformation cuts storage 3-4x with equal or better quality, but the motion mask's recall is the main limitation to probe in review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the motion-related Gaussian set $G_m$, selected through a four-step chain: a motion mask from optical flow and temporal difference (Eq. (3)); Gaussian ID maps that record which Gaussian dominates each pixel's $\alpha$ blending (Eq. (5)); density-based clustering followed by Delaunay-based convex hulls to pull in the Gaussians inside the moving volume (Eqs. (6)-(7)); and, for emerging objects, an attention mask computed from large per-pixel errors after deformation (Eq. (9)) that selects the subset $G_{\text{new}}$ for spherical harmonic color updates (Eqs. (10)-(11)). This chain does the work of confining all per-frame deformation and optimization to a small, scene-relevant fraction of the representation, which is the source of both the storage savings and the temporal consistency.
What would settle it
Render a dynamic sequence containing a thin, fast-moving object (such as a swinging cord) or a slow color change that falls below the optical-flow and frame-difference thresholds, and inspect the frames for stale background or ghosting exactly at the masked-out pixels; if such artifacts appear, the central claim that motion-related Gaussians fully cover the dynamics is false. More directly, compare the motion mask against hand-labeled moving-object masks on the evaluation views and measure its recall.
Extended reading notes
Core claim
MGStream is built on the idea that a dynamic scene can be split into a frozen static Gaussian model and a small per-frame set of motion-related Gaussians. To locate that set, the paper computes a motion mask $\hat{M}$ (Eq. (3)) by thresholding optical flow and frame difference, then applying morphological dilation and erosion; Gaussian ID maps (Eq. (5)) back-project the masked pixels to the surface Gaussians $G_o$; and density-based clustering plus Delaunay-based convex hulls (Eqs. (6)-(7)) add the interior Gaussians $G_i$, yielding $G_m$. The dynamic is then modeled by rigid translation and rotation offsets from a hash-grid deformation field applied only to $G_m$ (Eq. (8)), and emerging objects are handled by an attention map (Eq. (9)) that selects a subset $G_{\text{new}}$ whose spherical harmonic coefficients are updated (Eqs. (10)-(11)). Because the static Gaussians are never touched, the rendered background stays consistent across frames, which is the mechanism behind the reported flicker reduction and storage savings.
Load-bearing premise
The method assumes that the motion mask produced by fixed thresholds on optical flow and frame difference, after morphological dilation and erosion, marks every pixel belonging to a moving or newly appearing object, so that no pixel requiring an update is left to the static Gaussians.
Editorial extensions
If this is right
- Per-frame storage scales with the amount of motion in the scene rather than with scene size, since only the selected motion-related Gaussians receive new offsets and color updates.
- Static regions of the rendered video remain unchanged across frames, which is what lowers flicker metrics such as Ewarp.
- Emerging objects can be rendered without adding or densifying Gaussians, by updating the spherical harmonic colors of already-selected motion Gaussians under an attention mask.
- The pipeline remains causal and online: each frame is processed from the previous frame's state without access to future frames.
- Training time stays close to the fastest streaming baselines because the deformation and optimization stages run for only 100 iterations each on small Gaussian subsets.
Reading between the lines
- A natural follow-up is to make the motion mask adaptive: learned or per-scene thresholds would likely improve recall for slow and thin motions, which fixed thresholds can miss.
- The same 'update only what moves' separation could be transferred to other primitive-based streaming representations by replacing Gaussian ID maps with per-primitive visibility or hit counts from the rasterizer.
- A direct stress test would measure motion-mask recall against labeled object masks on the evaluation datasets; dynamic-region PSNR should track that recall.
- The reported storage reductions suggest that quantizing or entropy-coding the per-frame deltas could lower the streaming bitrate further.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MGStream, a per-frame streaming method for dynamic novel view synthesis with 3D Gaussians. It computes a 2D motion mask by intersecting a thresholded optical-flow map with a morphologically smoothed frame-difference map (Eq. 3), back-projects the mask to Gaussian IDs via alpha-blending ID maps, expands the selected set with a clustering and convex-hull step (Eqs. 4-7), and then deforms only the motion-related Gaussians with a per-frame rigid transformation (Eq. 8). An attention map based on the rendered L1 error selects a subset of those Gaussians whose spherical-harmonic colors are further optimized for emerging content (Eqs. 9-11). Static Gaussians are frozen. Experiments on N3DV and MeetRoom report improved PSNR, lower storage, shorter training time, and lower warping error relative to StreamRF, Dynamic-3DGS, and 3DGStream, with ablations in Tables 3-5.
Significance. If the mechanism is robust, the paper's separation of motion-related and static Gaussians is an elegant and practical idea: it freezes the static backbone, which plausibly explains the reported flicker reduction and storage savings, and it keeps the per-frame training cost low. The paper ships an open-source implementation, includes component ablations, and reports a wide range of metrics. The core risk is that the entire pipeline is gated by one hand-thresholded motion mask whose recall is never measured, and the evaluation is limited to two benign datasets; the central claim of general superiority over streaming 3DGS methods therefore needs additional evidence.
major comments (3)
- [§4.1, Eq. (3)] The motion mask is the sole gate for all updates to the dynamic content, and the static Gaussians are never revisited by any loss. The mask is formed by thresholding optical flow at 1 pixel and temporal difference at 10, then applying morphological DILATE and ERODE with a kernel size of 20 and intersecting the two cues. The paper does not report the recall or stability of this mask. Any moving or emerging region that falls below these thresholds, or is eroded by the morphological operations, is neither deformed per Eq. (8) nor color-optimized per Eqs. (9)-(11), so the previous frame's content persists and produces ghosts or stale appearance. Because the two datasets contain relatively slow, textured, centrally placed motion, the favorable comparisons in Table 1 may be dataset-specific. Please provide a recall/precision analysis of the motion mask against a reference, a sensitivity study over the thresholds and kernel size, and either a fallback mechanism that revisits high-error static regions or evidence that such a mechanism is unnecessary.
- [§5.2, Tables 1 and 3-5] The headline comparisons lack error bars, and the dynamic-region PSNR is stated to be defined in the supplementary material, which is not available in the submitted manuscript. The differences between methods are small (for example, 31.84 versus 32.02 PSNR on N3DV), and the ablations in Tables 3-5 do not state which dataset or scene they use, making it impossible to judge how the conclusions transfer across scenes. Please report results over multiple runs with variance, move the dynamic-region evaluation protocol into the main text, and specify the dataset and scene used for each ablation.
- [§4.2, Eqs. (9)-(11)] The optimization stage updates only the spherical-harmonic colors of the Gaussians selected by the attention mask; it does not update position, scale, opacity, or density, and no densification mechanism is mentioned. For an object that emerges in a region not already covered by motion-related Gaussians from the previous frame, color-only updates can change the appearance but cannot create new geometry. The ablation in Table 4 shows that adding color optimization improves PSNR from 33.51 to 34.31, but this does not establish that the method handles genuinely new geometry. Please clarify whether emergence is always confined to regions already represented by Gm, or provide experiments with an object entering the field of view from outside.
minor comments (5)
- [§2.2] The term 'wrap-based methods' appears twice in Section 2.2 and should be 'warp-based methods' to match the terminology used in the Introduction.
- [Fig. 4 caption] The caption contains a typo: 'motio-related 3DGs' should be 'motion-related 3DGs'.
- [Table 1 caption] The phrase 'The detail abouts PSNR calculation' should be 'The details of the PSNR calculation'.
- [Eq. (3)] The symbols D and E for DILATE and ERODE are used without defining the structuring element or kernel shape; please state these explicitly.
- [Tables 3-5] The ablation tables would be easier to interpret if they stated the dataset, scene, and the number of frames averaged, and if the units of Ewarp were given.
Circularity Check
No circularity found: the motion mask, Gaussian selection, deformation, and optimization are defined from input signals and losses, not from the claims they are used to support.
full rationale
I walked the paper's derivation chain. The motion mask in Eq. (3) is computed from a pre-trained optical flow model and the temporal difference between adjacent frames; it is not defined in terms of the model's output or in terms of the reported metrics. The motion-related Gaussians are then selected through GIM back-projection and the clustering-based convex hull algorithm in Eqs. (4)-(7), using only the mask and existing Gaussian positions. The deformation in Eq. (8) and the optimization in Eqs. (9)-(11) are trained against the standard reconstruction loss L_color, and the attention map in Eq. (9) uses the rendered L1 error as an internal training signal to choose which Gaussians to update; this is a self-referential optimization strategy, not a validation quantity that is later reported as an independent prediction. The Ewarp metric uses optical flow, but the method is not fitted to Ewarp, so this is not a circular evaluation. No fitted parameter is renamed as a prediction, no load-bearing result is imported solely from the authors' own prior work, and no uniqueness or external-support claim is used to force the architecture. The central comparisons in Tables 1 and 2 are therefore independent of the method's construction. I find no circular step.
Assumptions & free parameters
free parameters (7)
- optical flow threshold tau =
1
- temporal difference threshold =
10
- morphological kernel size =
20
- DBSCAN eps =
2
- DBSCAN min_samples =
10
- attention threshold percentile =
99th percentile of L1 error
- training iterations =
100 deformation, 100 optimization
assumptions (4)
- domain assumption Static scene regions have constant appearance across frames; lighting and shadows do not change.
- domain assumption The motion mask from optical flow and temporal difference completely covers all moving and emerging object pixels.
- domain assumption GIM back-projection plus clustering-based convex hull maps the 2D motion mask to exactly the 3DGs belonging to moving objects.
- domain assumption The pre-trained optical flow model (GMflow) provides reliable motion estimates for the input videos.
Cite this review
Pith. "Pith review of MGStream: Motion-aware 3D Gaussian for Streamable Dynamic Scene Reconstruction." pith.science (2026). https://pith.science/paper/HH6VREF7
@misc{pith2026250513839,
author = {Pith},
title = {Pith review of: MGStream: Motion-aware 3D Gaussian for Streamable Dynamic Scene Reconstruction},
year = {2026},
howpublished = {\url{https://pith.science/paper/HH6VREF7}},
note = {Machine review of arXiv:2505.13839}
}
read the original abstract
3D Gaussian Splatting (3DGS) has gained significant attention in streamable dynamic novel view synthesis (DNVS) for its photorealistic rendering capability and computational efficiency. Despite much progress in improving rendering quality and optimization strategies, 3DGS-based streamable dynamic scene reconstruction still suffers from flickering artifacts and storage inefficiency, and struggles to model the emerging objects. To tackle this, we introduce MGStream which employs the motion-related 3D Gaussians (3DGs) to reconstruct the dynamic and the vanilla 3DGs for the static. The motion-related 3DGs are implemented according to the motion mask and the clustering-based convex hull algorithm. The rigid deformation is applied to the motion-related 3DGs for modeling the dynamic, and the attention-based optimization on the motion-related 3DGs enables the reconstruction of the emerging objects. As the deformation and optimization are only conducted on the motion-related 3DGs, MGStream avoids flickering artifacts and improves the storage efficiency. Extensive experiments on real-world datasets N3DV and MeetRoom demonstrate that MGStream surpasses existing streaming 3DGS-based approaches in terms of rendering quality, training/storage efficiency and temporal consistency. Our code is available at: https://github.com/pcl3dv/MGStream.
Figures
Reference graph
Works this paper leans on
-
[1]
Hyperreel: High-fidelity 6-dof video with ray- conditioned sampling
Benjamin Attal, Jia-Bin Huang, Christian Richardt, Michael Zollhoefer, Johannes Kopf, Matthew O’Toole, and Changil Kim. Hyperreel: High-fidelity 6-dof video with ray- conditioned sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16610–16620, 2023. 1, 7
work page 2023
-
[2]
Per-gaussian embedding- based deformation for deformable 3d gaussian splatting
Jeongmin Bae, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, and Youngjung Uh. Per-gaussian embedding- based deformation for deformable 3d gaussian splatting. In European Conference on Computer Vision, pages 321–335. Springer, 2024. 3
work page 2024
-
[3]
Loopsparsegs: Loop based sparse-view friendly gaussian splatting
Zhenyu Bao, Guibiao Liao, Kaichen Zhou, Kanglin Liu, Qing Li, and Guoping Qiu. Loopsparsegs: Loop based sparse-view friendly gaussian splatting. arXiv preprint arXiv:2408.00254, 2024. 2
arXiv 2024
-
[4]
Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields
Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 5855–5864,
-
[5]
Mip-nerf 360: Unbounded anti-aliased neural radiance fields
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5470–5479, 2022. 2
2022
-
[6]
Zip-nerf: Anti-aliased grid-based neural radiance fields
Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 19697–19705, 2023. 2
2023
-
[7]
Gary Bradski, Adrian Kaehler, et al. Opencv. Dr. Dobb’s journal of software tools, 3(2), 2000. 5
work page 2000
-
[8]
Immersive light field video with a layered mesh representation
Michael Broxton, John Flynn, Ryan Overbeck, Daniel Erick- son, Peter Hedman, Matthew Duvall, Jason Dourgarian, Jay Busch, Matt Whalen, and Paul Debevec. Immersive light field video with a layered mesh representation. ACM Trans- actions on Graphics (TOG), 39(4):86–1, 2020. 1
work page 2020
Show all 72 references
-
[9]
Unstructured lumigraph ren- dering
Chris Buehler, Michael Bosse, Leonard McMillan, Steven Gortler, and Michael Cohen. Unstructured lumigraph ren- dering. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques , pages 425– 432, 2001. 2
2001
-
[10]
Hexplane: A fast representa- tion for dynamic scenes
Ang Cao and Justin Johnson. Hexplane: A fast representa- tion for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 130–141, 2023. 7
2023
-
[11]
Plenoptic sampling
Jin-Xiang Chai, Xin Tong, Shing-Chow Chan, and Heung- Yeung Shum. Plenoptic sampling. In Proceedings of the 27th annual conference on Computer graphics and interac- tive techniques, pages 307–318, 2000. 2
2000
-
[12]
Gaussianeditor: Swift and control- lable 3d editing with gaussian splatting
Yiwen Chen, Zilong Chen, Chi Zhang, Feng Wang, Xi- aofeng Yang, Yikai Wang, Zhongang Cai, Lei Yang, Huaping Liu, and Guosheng Lin. Gaussianeditor: Swift and control- lable 3d editing with gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patt...
2024
-
[13]
High-quality streamable free-viewpoint video
Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Den- nis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics (ToG) , 34(4):1–13,
-
[14]
Unstructured light fields
Abe Davis, Marc Levoy, and Fredo Durand. Unstructured light fields. In Computer Graphics Forum, pages 305–314. Wiley Online Library, 2012. 2
2012
-
[15]
Motion2fusion: Real-time volumetric performance capture
Mingsong Dou, Philip Davidson, Sean Ryan Fanello, Sameh Khamis, Adarsh Kowdle, Christoph Rhemann, Vladimir Tankovich, and Shahram Izadi. Motion2fusion: Real-time volumetric performance capture. ACM Transactions on Graphics (ToG), 36(6):1–16, 2017. 1
2017
-
[16]
4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes
Yuanxing Duan, Fangyin Wei, Qiyu Dai, Yuhang He, Wen- zheng Chen, and Baoquan Chen. 4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes. In ACM SIGGRAPH 2024 Conference Papers , pages 1–11,
2024
-
[17]
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, J ¨org Sander, Xiaowei Xu, et al. A density-based algorithm for discovering clusters in large spatial databases with noise. In kdd, pages 226–231,
-
[18]
Plenoxels: Radiance fields without neural networks
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5501–5510, 2022. 2
2022
-
[19]
K-planes: Explicit radiance fields in space, time, and appearance
Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 12479–12488, 2023. 3, 7
2023
-
[20]
Hicom: Hierarchical coherent motion for streamable dynamic scene with 3d gaussian splatting
Qiankun Gao, Jiarui Meng, Chengxiang Wen, Jie Chen, and Jian Zhang. Hicom: Hierarchical coherent motion for streamable dynamic scene with 3d gaussian splatting. arXiv preprint arXiv:2411.07541, 2024. 3
2024 arXiv
-
[21]
Gaussianflow: Splatting gaussian dynamics for 4d content creation
Quankai Gao, Qiangeng Xu, Zhe Cao, Ben Mildenhall, Wen- chao Ma, Le Chen, Danhang Tang, and Ulrich Neumann. Gaussianflow: Splatting gaussian dynamics for 4d content creation. arXiv preprint arXiv:2403.12365, 2024. 4
2024 arXiv
-
[22]
Queen: Quantized efficient encoding of dynamic gaussians for streaming free-viewpoint videos
Sharath Girish, Tianye Li, Amrita Mazumdar, Abhinav Shri- vastava, Shalini De Mello, et al. Queen: Quantized efficient encoding of dynamic gaussians for streaming free-viewpoint videos. Advances in Neural Information Processing Systems, 37:43435–43467, 2024. 3
2024
-
[23]
Semantic gaussians: Open-vocabulary scene understanding with 3d gaussian splatting.arXiv preprint arXiv:2403.15624,
Jun Guo, Xiaojian Ma, Yue Fan, Huaping Liu, and Qing Li. Semantic gaussians: Open-vocabulary scene understanding with 3d gaussian splatting.arXiv preprint arXiv:2403.15624,
-
[24]
Motion-aware 3d gaussian splatting for efficient dynamic scene reconstruction
Zhiyang Guo, Wengang Zhou, Li Li, Min Wang, and Houqiang Li. Motion-aware 3d gaussian splatting for efficient dynamic scene reconstruction. arXiv preprint arXiv:2403.11447, 2024. 4
2024 arXiv
-
[25]
S4d: Streaming 4d real-world reconstruction with gaussians 9 and 3d control points
Bing He, Yunuo Chen, Guo Lu, Li Song, and Wenjun Zhang. S4d: Streaming 4d real-world reconstruction with gaussians 9 and 3d control points. arXiv preprint arXiv:2408.13036 ,
-
[26]
Point’n move: Interactive scene object manipu- lation on gaussian splatting radiance fields
Jiajun Huang, Hongchuan Yu, Jianjun Zhang, and Hammadi Nait-Charif. Point’n move: Interactive scene object manipu- lation on gaussian splatting radiance fields. IET Image Pro- cessing, 2024. 2
2024
-
[27]
Beyond face rotation: Global and local perception gan for photoreal- istic and identity preserving frontal view synthesis
Rui Huang, Shu Zhang, Tianyu Li, and Ran He. Beyond face rotation: Global and local perception gan for photoreal- istic and identity preserving frontal view synthesis. In Pro- ceedings of the IEEE international conference on computer vision, pages 2439–2448, 2017. 2
2017
-
[28]
Virtualized reality: Constructing virtual worlds from real scenes
Takeo Kanade, Peter Rander, and PJ Narayanan. Virtualized reality: Constructing virtual worlds from real scenes. IEEE multimedia, 4(1):34–47, 1997. 1
1997
-
[29]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[30]
Blind video temporal consistency via deep video prior.Advances in Neu- ral Information Processing Systems, 33:1083–1093, 2020
Chenyang Lei, Yazhou Xing, and Qifeng Chen. Blind video temporal consistency via deep video prior.Advances in Neu- ral Information Processing Systems, 33:1083–1093, 2020. 7
2020
-
[31]
Blind video deflickering by neural filtering with a flawed atlas
Chenyang Lei, Xuanchi Ren, Zhaoxiang Zhang, and Qifeng Chen. Blind video deflickering by neural filtering with a flawed atlas. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10439– 10448, 2023. 7
2023
-
[32]
Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normaliza- tion
Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normaliza- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 2...
-
[33]
Streaming radiance fields for 3d video synthe- sis
Lingzhi Li, Zhen Shen, Zhongshu Wang, Li Shen, and Ping Tan. Streaming radiance fields for 3d video synthe- sis. Advances in Neural Information Processing Systems, 35: 13485–13498, 2022. 2, 3, 6, 7
2022
-
[34]
Neural 3d video synthesis from multi-view video
Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. In Proceedings of the IEEE/CVF Conference on Computer Vi- si...
2022
-
[35]
Spacetime gaus- sian feature splatting for real-time dynamic view synthesis
Zhan Li, Zhang Chen, Zhong Li, and Yi Xu. Spacetime gaus- sian feature splatting for real-time dynamic view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8508–8520, 2024. 2, 3, 7
2024
-
[36]
Clip-gs: Clip-informed gaussian splatting for real-time and view- consistent 3d semantic understanding
Guibiao Liao, Jiankun Li, Zhenyu Bao, Xiaoqing Ye, Jingdong Wang, Qing Li, and Kanglin Liu. Clip-gs: Clip-informed gaussian splatting for real-time and view- consistent 3d semantic understanding. arXiv preprint arXiv:2404.14249, 2024. 2
2024 arXiv
-
[37]
Swings: Sliding window gaussian splatting for volumetric video streaming with arbi- trary length
Bangya Liu and Suman Banerjee. Swings: Sliding window gaussian splatting for volumetric video streaming with arbi- trary length. arXiv preprint arXiv:2409.07759, 2024. 3
2024 arXiv
-
[38]
Posegan: A pose- to-image translation framework for camera localization
Kanglin Liu, Qing Li, and Guoping Qiu. Posegan: A pose- to-image translation framework for camera localization. IS- PRS Journal of Photogrammetry and Remote Sensing , 166: 308–315, 2020. 2
2020
-
[39]
Dynamics-aware gaussian splat- ting streaming towards fast on-the-fly training for 4d recon- struction
Zhening Liu, Yingdong Hu, Xinjie Zhang, Jiawei Shao, Ze- hong Lin, and Jun Zhang. Dynamics-aware gaussian splat- ting streaming towards fast on-the-fly training for 4d recon- struction. arXiv preprint arXiv:2411.14847, 2024. 3
2024 arXiv
-
[40]
Scaffold-gs: Structured 3d gaussians for view-adaptive rendering
Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. Scaffold-gs: Structured 3d gaussians for view-adaptive rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20654–20664, 2024. 2
2024
-
[41]
Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis
Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713, 2023. 2, 3, 5, 6, 7
2023 arXiv
-
[42]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In European Conference on Computer Vision, pages 405–421. Springer, 2020. 1, 2
2020
-
[43]
Instant neural graphics primitives with a mul- tiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM transactions on graphics (TOG), 41(4):1–15, 2022. 2
2022
-
[44]
Compact3d: Com- pressing gaussian splat radiance field models with vector quantization
KL Navaneet, Kossar Pourahmadi Meibodi, Soroush Abbasi Koohpayegani, and Hamed Pirsiavash. Compact3d: Com- pressing gaussian splat radiance field models with vector quantization. arXiv preprint arXiv:2311.18159, 2023. 2
2023 arXiv
-
[45]
Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs
Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition,...
2022
-
[46]
Nerfies: Deformable neural radiance fields
Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5865–5874, 2021. 1
2021
-
[47]
Hypernerf: A higher- dimensional representation for topologically varying neural radiance fields
Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M Seitz. Hypernerf: A higher- dimensional representation for topologically varying neural radiance fields. arXiv preprint arXiv:2106.13228, 2021
2021 arXiv
-
[48]
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 10318–10327, 2021. 1, 2
2021
-
[49]
Langsplat: 3d language gaussian splatting
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister. Langsplat: 3d language gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20051–20060, 2024. 2
2024
-
[50]
Blazebvd: 10 Make scale-time equalization great again for blind video de- flickering
Xinmin Qiu, Congying Han, Zicheng Zhang, Bonan Li, Tiande Guo, Pingyu Wang, and Xuecheng Nie. Blazebvd: 10 Make scale-time equalization great again for blind video de- flickering. arXiv preprint arXiv:2403.06243, 2024. 7
2024 arXiv
-
[51]
Motion detection techniques using optical flow
Amir Akramin Shafie, Fadhlan Hafiz, MH Ali, et al. Motion detection techniques using optical flow. World Academy of Science, Engineering and Technology, 56:559–561, 2009. 4
2009
-
[52]
Nerf- player: A streamable dynamic scene representation with de- composed neural radiance fields.IEEE Transactions on Visu- alization and Computer Graphics , 29(5):2732–2742, 2023
Liangchen Song, Anpei Chen, Zhong Li, Zhang Chen, Lele Chen, Junsong Yuan, Yi Xu, and Andreas Geiger. Nerf- player: A streamable dynamic scene representation with de- composed neural radiance fields.IEEE Transactions on Visu- alization and Computer Graphics , 29(5):2732–2742, 2023. 7
2023
-
[53]
Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction
Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5459– 5469, 2022. 2
2022
-
[54]
3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos
Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, and Wei Xing. 3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free- viewpoint videos. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, ...
2024
-
[55]
Optical flow guided feature: A fast and robust motion representation for video action recognition
Shuyang Sun, Zhanghui Kuang, Lu Sheng, Wanli Ouyang, and Wei Zhang. Optical flow guided feature: A fast and robust motion representation for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1390–1399, 2018. 4
2018
-
[56]
Mixed neural voxels for fast multi- view video synthesis
Feng Wang, Sinan Tan, Xinghang Li, Zeyue Tian, Yafei Song, and Huaping Liu. Mixed neural voxels for fast multi- view video synthesis. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 19706– 19716, 2023. 7
2023
-
[57]
Sparsenerf: Distilling depth ranking for few-shot novel view synthesis
Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Zi- wei Liu. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 9065–9076,
-
[58]
Neural residual radiance fields for streamably free-viewpoint videos
Liao Wang, Qiang Hu, Qihan He, Ziyu Wang, Jingyi Yu, Tinne Tuytelaars, Lan Xu, and Minye Wu. Neural residual radiance fields for streamably free-viewpoint videos. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 76–87, 2023. 3
2023
-
[59]
Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 2
2021 arXiv
-
[60]
Bad-nerf: Bundle adjusted deblur neural radiance fields
Peng Wang, Lingzhe Zhao, Ruijie Ma, and Peidong Liu. Bad-nerf: Bundle adjusted deblur neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 4170–4179, 2023. 2
2023
-
[61]
4d gaussian splatting for real-time dynamic scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20310–20320,...
2024
-
[62]
Unifying flow, stereo and depth estimation
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying flow, stereo and depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. 7
2023
-
[63]
View inde- pendent generative adversarial network for novel view syn- thesis
Xiaogang Xu, Ying-Cong Chen, and Jiaya Jia. View inde- pendent generative adversarial network for novel view syn- thesis. In Proceedings of the IEEE/CVF international con- ference on computer vision, pages 7791–7800, 2019. 2
2019
-
[64]
Representing long volumet- ric video with temporal gaussian hierarchy
Zhen Xu, Yinghao Xu, Zhiyuan Yu, Sida Peng, Jiaming Sun, Hujun Bao, and Xiaowei Zhou. Representing long volumet- ric video with temporal gaussian hierarchy. ACM Transac- tions on Graphics (TOG), 43(6):1–18, 2024. 2, 3
2024
-
[65]
Freenerf: Im- proving few-shot neural rendering with free frequency reg- ularization
Jiawei Yang, Marco Pavone, and Yue Wang. Freenerf: Im- proving few-shot neural rendering with free frequency reg- ularization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8254–8263,
-
[66]
Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting
Zeyu Yang, Hongye Yang, Zijie Pan, and Li Zhang. Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting. arXiv preprint arXiv:2310.10642, 2023. 2, 3, 7
2023 arXiv
-
[67]
Deformable 3d gaussians for high- fidelity monocular dynamic scene reconstruction
Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high- fidelity monocular dynamic scene reconstruction. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20331–20341, 2024. 1, 2
2024
-
[68]
V ol- ume rendering of neural implicit surfaces
Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. V ol- ume rendering of neural implicit surfaces. Advances in Neu- ral Information Processing Systems, 34:4805–4815, 2021. 2
2021
-
[69]
Monosdf: Exploring monocu- lar geometric cues for neural implicit surface reconstruc- tion
Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sat- tler, and Andreas Geiger. Monosdf: Exploring monocu- lar geometric cues for neural implicit surface reconstruc- tion. Advances in neural information processing systems , 35:25018–25032, 2022. 2
2022
-
[70]
Mip-splatting: Alias-free 3d gaussian splat- ting
Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-splatting: Alias-free 3d gaussian splat- ting. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 19447–19456,
-
[71]
Fsgs: Real-time few-shot view synthesis using gaussian splatting
Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang. Fsgs: Real-time few-shot view synthesis using gaussian splatting. In European Conference on Computer Vision , pages 145–163. Springer, 2025. 2
2025
-
[72]
High-quality video view interpolation using a layered representation
C Lawrence Zitnick, Sing Bing Kang, Matthew Uyttendaele, Simon Winder, and Richard Szeliski. High-quality video view interpolation using a layered representation. ACM transactions on graphics (TOG), 23(3):600–608, 2004. 1 11
2004
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.