Pith. sign in

REVIEW 3 major objections 4 minor 68 references

Scale-GS: Efficient Scalable Gaussian Splatting via Redundancy-filtering Training on Streaming Content

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Scale-GS claims that streaming dynamic 3D scenes can be trained frame-by-frame in about three seconds with higher quality than prior methods, by filtering Gaussian redundancy through a coarse-to-fine scale hierarchy.

desk verdict A solid, incremental improvement for streaming 3DGS with a clean multi-scale design; the efficiency claim is directionally supported but underevidenced on long-run scale drift. read the letter →

arxiv 2508.21444 v1 pith:DMVX4ILG submitted 2025-08-29 cs.CV

classification cs.CV
keywords streamingGaussiansplattingmulti-scalerepresentationdynamicscenerenderingnovelviewsynthesisdeformationandspawningredundancyfilteringtrainingefficiencyanchor-basedGaussians
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that dynamic 3D scenes can be streamed and updated frame by frame at a few seconds of training per frame, instead of tens of minutes to hours, while improving rendering quality. The central move is to treat Gaussian spheres as redundant across scale, space, and time: coarse Gaussians are trained first on low-resolution views, finer Gaussians are activated only where gradients show unresolved dynamics, and static regions and uninformative viewpoints are masked out. Between frames, existing Gaussians are deformed and new ones are spawned only where deformation cannot capture new or wide-range motion, then redundant Gaussians are pruned. On three real-world multi-view datasets the paper reports higher PSNR than prior streaming methods with shorter per-frame training time and real-time rendering. This matters because low-latency training is the bottleneck for VR, AR, immersive video, and telepresence.

What carries the argument

The load-bearing mechanism is the anchor-based multi-scale Gaussian hierarchy. Each Gaussian sphere is assigned to a scale level by a size range learned from the first frame, and a clamp keeps it in that range; training runs coarse-to-fine at downsampled resolutions. Activation is governed by an average-gradient test per anchor: if the mean gradient after deformation exceeds a threshold that shrinks by a factor of four per level, the next finer level is turned on and octree spawning adds new Gaussians in high-gradient subspaces. This threshold test is what converts 'redundancy filtering' from a slogan into a concrete compute-allocation rule.

What would settle it

Take a long multi-view sequence where fine-scale detail enters only after many frames (for example, a close-up object moves into view or a flame ignites), run Scale-GS with the hierarchy fixed from frame one, and compare PSNR and per-frame time against re-estimating scale ranges periodically. A drop in quality or a rise in per-frame time in the later segment would show the first-frame scale assumption is the limiting factor.

Watch

Extended reading notes

Core claim

The paper's central claim: for streaming dynamic scenes, filtering redundant Gaussians by scale makes training faster than prior Gaussian streaming methods while raising rendering quality. Gaussians are organized in an anchor-based multi-scale hierarchy: coarse Gaussians train on low-resolution views first, and finer Gaussians activate only when an anchor's mean gradient exceeds a level-dependent threshold. Inter-frame motion uses a hybrid strategy—deformation applies residual updates to carried-over Gaussians, spawning adds new Gaussians in octree subspaces where deformation cannot represent new or wide-range motion. A bidirectional masking step drops static anchors and ranks camera views b

Load-bearing premise

The entire scale hierarchy—which Gaussian sizes belong to which level, and the activation thresholds—is fixed from the first frame, on the premise that later frames contain similar Gaussian sizes; if late-appearing content has very different scales, the allocation can be wrong.

Editorial extensions

If this is right

  • Per-frame training of long dynamic sequences drops to a few seconds, making on-the-fly streaming of free-viewpoint video practical for VR/AR and telepresence.
  • Compute is spent where gradients demand it: static regions and non-informative viewpoints are masked out, so training cost scales with scene change rather than scene size.
  • Rendering stays real-time after training, because the final representation is still a compact set of 3D Gaussians.
  • The hybrid deformation-plus-spawning update gives a middle path between deformation-only methods, which miss new content, and spawning-only methods, which are slow.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The first-frame scale assumption is the fragile premise: a sequence that introduces objects at very different Gaussian sizes later could misassign levels. A test would be to shift the scale distribution partway through a sequence and watch PSNR or activation behavior.
  • The same gradient-threshold machinery could be tied to a rate-distortion budget, letting the number of spawned Gaussians adapt to a target per-frame time rather than to a fixed threshold.
  • Viewpoint selection by relevance score could be reused for sparse-camera streaming, where only a few cameras need to be decoded per frame.
  • The octree spawning rule could be extended to long unbounded scenes by allowing anchors to be created and retired over time, not just Gaussians within fixed anchors.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents Scale-GS, an anchor-based multi-scale Gaussian splatting framework for streaming dynamic scene reconstruction. The method organizes Gaussians into L scale levels, trains coarse levels first on downsampled views, and selectively activates finer levels via a gradient-threshold criterion. A hybrid deformation/spawning module models inter-frame motion, a mask-based redundancy removal keeps Gaussian counts bounded, and a bidirectional adaptive masking step focuses optimization on dynamic anchors and informative viewpoints. The authors report state-of-the-art PSNR and lowest per-frame training time on NV3D, MeetRoom, and Google Immersive, together with qualitative results and ablations over scale count, deformation/spawning, and view selection.

Significance. If the reported results are robust, Scale-GS is a practically valuable contribution to real-time streaming Gaussian splatting, with consistent PSNR improvements of 0.85–1.46 dB over the previous best streaming method (IGS) and per-frame training time reductions on all three benchmarks. The paper is clearly written, the framework is well motivated, and the ablations cover the main design choices. However, the current evidence does not fully support the strength of the central efficiency claim: the final comparison and the hyperparameter selection share test scenes, no statistical uncertainty is reported, and a key modeling assumption about temporal stability of Gaussian scale statistics is not stress-tested. These issues are fixable and do not invalidate the method, but they need to be addressed before the claimed generality can be accepted.

major comments (3)
  1. [§IV-D, Table I] The central claim of 'most efficient training' rests on Table I, but the table reports aggregate means with no error bars, standard deviations, per-scene breakdown, or significance tests. The training-time differences are small on NV3D and MeetRoom (3.2 vs 3.6 s and 3.0 vs 3.2 s for Scale-GS vs IGS), so without repeated runs or per-scene variance it is unclear whether the efficiency advantage is meaningful. Please report per-scene results, number of runs, and variance, or at least a significance test on the paired scenes.
  2. [§IV-B and §IV-E.1] Hyperparameters central to the method are selected using scenes that later appear in the final comparison. The number of scales L=3 is chosen via an ablation on the Face Print scene from Google Immersive (Table II), and the threshold schedule τ_add^(l)=0.01/4^(l-1) is stated in §IV-B without a sensitivity study. Since Table I includes the Google Immersive dataset, the reported gains may partly reflect configuration fitting on test scenes. A held-out validation split or a robustness analysis over L and τ_add is needed to support the generality of the claimed efficiency/quality trade-off.
  3. [§III-B, Eq. (5), Eq. (12)] The multi-scale hierarchy is fixed from the initial frame: 'the distribution learned from the initial frame can reasonably approximate that of all frames' (§III-B). This is load-bearing because Eq. (5) clamps Gaussian scales into time-invariant bins and Eq. (12) uses those bins to decide spawning and level activation. The paper does not test what happens when later frames introduce content with substantially different Gaussian scale statistics (e.g., an object entering the field of view, a flame appearing, or a close-up with finer detail than anything in frame 0). In that regime, spawned Gaussians may be forced into inappropriate scale levels, degrading either PSNR or the selective-training savings. Please measure scale-distribution drift over the 300-frame sequences, report the fraction of spawned Gaussians hitting the Eq. (5) clamp, and/or include an experiment that re-initializes the
minor comments (4)
  1. [§III-E / abstract] The phrase 'bidirectional bidirectional adaptive masking' appears in §III-E and in the abstract; remove the duplicate.
  2. [Eq. (16)] The notation is unclear: d_{c_k} is described as 'the normal direction of the viewpoint,' but it should be the viewing direction vector. Please clarify the definitions of n_v and d_{c_k}.
  3. [Fig. 1] The caption states that circle radius corresponds to 'average storage per frame,' but storage is not reported in Table I or elsewhere. Either add storage numbers or adjust the caption.
  4. [§III-A / Eq. (3)] Minor typos include 'winthin' (should be 'within') and 'the the scale' in §III-B. These should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the efficiency and quality claims are measured benchmarks, and the initial-frame scale hierarchy is an explicit inductive premise rather than a prediction derived from the method's own outputs.

full rationale

The paper's central claims are empirical: 'Extensive experiments demonstrate that Scale-GS achieves superior visual quality while significantly reducing training time compared to state-of-the-art methods' (Sec. V) is supported by Table I measurements against external baselines. No equation in the manuscript forces the reported PSNR/training-time numbers from a fitted parameter; the comparisons to IGS, 3DGStream, HiCoM, and others are direct measurements. The multi-scale hierarchy is built from an explicit assumption, 'the distribution learned from the initial frame can reasonably approximate that of all frames' (Sec. III-B), and Eq. 5 clamps each Gaussian into precomputed scale bins. This is a design premise about temporal scale-stationarity, not a conclusion made equivalent to its inputs by construction: the scale bins are estimated once from the first frame and then applied to later frames, which is a prediction subject to empirical verification rather than a tautology. The level-count L=3 and thresholds tau_add^(l)=0.01/4^(l-1) are selected via ablations on Google Immersive scenes that also appear in Table I; this is a hyperparameter-selection concern that could affect fairness/generality, but it does not make the benchmark claim a derived consequence of the chosen values because the reported numbers are measured rather than computed from the hyperparameters. Citations to anchor-based Gaussian structures and level-specific gradient thresholds (e.g., [28] Scaffold-GS) are external support, and no uniqueness theorem or central premise is imported solely from the authors' own prior work. The skeptical concern that later frames could introduce new scale classes outside the initial bins is a limitation of the method's assumption, not a circular step in the derivation chain. I find no self-definitional prediction, no fitted parameter renamed as a result, and no load-bearing self-citation; hence score 0.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the first-frame-derived scale hierarchy, the gradient-threshold activation policy, and the view-selection heuristics. The thresholds and number of levels are tuned on the evaluation datasets, while the number of Gaussians spawned and the top-k view selection are unspecified. No new physical entities are introduced; the multi-scale anchors and spawned Gaussians are part of the same trained representation.

free parameters (6)
  • L (number of scale levels) = 3
    Chosen via ablation on the Face Print scene of Google Immersive (Table II), then used for all datasets; the ablation also shows quality degrades for L=4 and L=5.
  • tau_add base gradient threshold (level 1) = 0.01
    Level-specific thresholds {0.01, 0.0025, 0.000625} are set in Sec. IV-B; the 0.01 scale factor is hand-set and directly controls when finer levels and spawning are triggered, hence the quality-efficiency trade-off.
  • lambda_SSIM = 0.2
    Loss weight in Eq. (10), set in Sec. IV-B without sensitivity analysis.
  • lambda_r (redundancy removal weight) = 0.001
    Sparsity weight in Eq. (14), set in Sec. IV-B without sensitivity analysis.
  • Gaussians spawned per octree subspace = not stated
    Sec. III-C says 'a fixed number of Gaussians are randomly assigned' but does not give the number; it affects both quality and training time.
  • Top-k views and tau_view = not stated
    Eq. (15) defines view relevance with threshold tau_view and the text says 'top-ranked views are selected', but no k or threshold value is reported.
assumptions (5)
  • standard math Differentiable volume rendering with anchor-based Gaussian parameterization (Scaffold-GS)
    Eqs. (1)-(4) are taken from 3DGS [1] and Scaffold-GS [28] without proof; the paper builds on them.
  • domain assumption The Gaussian scale distribution estimated from the first frame approximates all frames
    Sec. III-B: 'the distribution learned from the initial frame can reasonably approximate that of all frames'. This fixes the scale ranges [s_min, s_max] of every level for the whole sequence.
  • domain assumption Inter-frame pixel differences back-projected to 3D indicate dynamic anchors
    Sec. III-E: 'temporal variations between two consecutive frames enable a coarse estimation of spatial motion... back-projected into 3D space'. Assumes this signal identifies motion despite occlusions and view shifts.
  • ad hoc to paper Gradient magnitude relative to tau_add = Vol/4^{l-1} measures whether the current scale can represent the motion
    Eq. (12) defines the activation threshold; the form is asserted with no derivation and the 0.01 factor is tuned. It is the main policy deciding when to spawn and advance scales.
  • ad hoc to paper View relevance score S(c_k) = sum of IoU indicators weighted by |n_v^T d_c| selects the most informative viewpoints
    Eqs. (15)-(16) define the view selection; this heuristic is not derived or validated against alternatives beyond the BAM ablation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scale-GS: Efficient Scalable Gaussian Splatting via Redundancy-filtering Training on Streaming Content." pith.science (2026). https://pith.science/paper/DMVX4ILG

@misc{pith2026250821444,
  author       = {Pith},
  title        = {Pith review of: Scale-GS: Efficient Scalable Gaussian Splatting via Redundancy-filtering Training on Streaming Content},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DMVX4ILG}},
  note         = {Machine review of arXiv:2508.21444}
}
read the original abstract

3D Gaussian Splatting (3DGS) enables high-fidelity real-time rendering, a key requirement for immersive applications. However, the extension of 3DGS to dynamic scenes remains limitations on the substantial data volume of dense Gaussians and the prolonged training time required for each frame. This paper presents \M, a scalable Gaussian Splatting framework designed for efficient training in streaming tasks. Specifically, Gaussian spheres are hierarchically organized by scale within an anchor-based structure. Coarser-level Gaussians represent the low-resolution structure of the scene, while finer-level Gaussians, responsible for detailed high-fidelity rendering, are selectively activated by the coarser-level Gaussians. To further reduce computational overhead, we introduce a hybrid deformation and spawning strategy that models motion of inter-frame through Gaussian deformation and triggers Gaussian spawning to characterize wide-range motion. Additionally, a bidirectional adaptive masking mechanism enhances training efficiency by removing static regions and prioritizing informative viewpoints. Extensive experiments demonstrate that \M~ achieves superior visual quality while significantly reducing training time compared to state-of-the-art methods.

Figures

Figures reproduced from arXiv: 2508.21444 by the authors.

Figure 1
Figure 1. The proposed Scale-GS under dynamic scene achieves best rendering quality with the shortest training time. The left figures show results of [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The framework of Scale-GS. (a) The multi-scale decomposition across different resolution levels, where finer scales capture increasingly detailed [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The detail of hybrid deformation-spawning strategy across multi [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Qualitative comparison results on the NV3D datasets(scene flame-steak). The frame index is 2, 151, and 183 from up to down. For each frame index, [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Qualitative comparison results on the Google Immersive datasets(scene Dog). The frame index is 1, 18 and 33 from up to down. For each frame [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 7
Figure 7. Figure 7: The visual quality(from left to right: full hybrid, w/o deformation, [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 67 canonical work pages

  1. [65]

    Reduction of forgetting by contextual variation during encoding using 360-degree video-based immersive virtual environments,

    T. Mizuho, T. Narumi, and H. Kuzuoka, “Reduction of forgetting by contextual variation during encoding using 360-degree video-based immersive virtual environments,” IEEE Trans. Vis. Comput. Graph , 2024

  2. [66]

    Wavelet-based fast decoding of 360 videos,

    C. Groth, S. Fricke, S. Castillo, and M. Magnor, “Wavelet-based fast decoding of 360 videos,” IEEE Trans. Vis. Comput. Graph , vol. 29, no. 5, pp. 2508–2516, 2023

  3. [1]

    3D Gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian splatting for real-time radiance field rendering.” ACM Trans. Graph. , vol. 42, no. 4, pp. 139–1, 2023

  4. [2]

    Sc- gs: Sparse-controlled gaussian splatting for editable dynamic scenes,

    Y .-H. Huang, Y .-T. Sun, Z. Yang, X. Lyu, Y .-P. Cao, and X. Qi, “Sc- gs: Sparse-controlled gaussian splatting for editable dynamic scenes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 4220– 4230

  5. [3]

    As-rigid-as-possible deformation of gaussian radiance fields,

    X. Tong, T. Shao, Y . Weng, Y . Yang, and K. Zhou, “As-rigid-as-possible deformation of gaussian radiance fields,” IEEE Trans. Vis. Comput. Graph., 2025

  6. [4]

    RGAvatar: Relightable 4D Gaussian avatar from monocular videos,

    Z. Fan, S.-S. Huang, Y . Zhang, D. Shang, J. Zhang, Y . Guo, and H. Huang, “RGAvatar: Relightable 4D Gaussian avatar from monocular videos,” IEEE Trans. Vis. Comput. Graph , 2025

  7. [5]

    Fov-GS: Foveated 3D Gaussian splatting for dynamic scenes,

    R. Fan, J. Wu, X. Shi, L. Zhao, Q. Ma, and L. Wang, “Fov-GS: Foveated 3D Gaussian splatting for dynamic scenes,” IEEE Trans. Vis. Comput. Graph, 2025

  8. [6]

    Deformable 3D Gaussians for high-fidelity monocular dynamic scene reconstruc- tion,

    Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3D Gaussians for high-fidelity monocular dynamic scene reconstruc- tion,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 20 331–20 341

Show all 68 references
  1. [7]

    Spacetime Gaussian feature splat- ting for real-time dynamic view synthesis,

    Z. Li, Z. Chen, Z. Li, and Y . Xu, “Spacetime Gaussian feature splat- ting for real-time dynamic view synthesis,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 8508–8520

  2. [8]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 20 310–20 320

  3. [9]

    4D Gaussian splatting with scale-aware residual field and adaptive optimization for real-time rendering of temporally complex dynamic scenes,

    J. Yan, R. Peng, L. Tang, and R. Wang, “4D Gaussian splatting with scale-aware residual field and adaptive optimization for real-time rendering of temporally complex dynamic scenes,” in Proc. 32nd ACM Int. Conf. Multimedia , 2024, pp. 7871–7880

  4. [10]

    3DGStream: On-the-fly training of 3D Gaussians for efficient streaming of photo- realistic free-viewpoint videos,

    J. Sun, H. Jiao, G. Li, Z. Zhang, L. Zhao, and W. Xing, “3DGStream: On-the-fly training of 3D Gaussians for efficient streaming of photo- realistic free-viewpoint videos,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2024, pp. 20 675–20 685

  5. [11]

    Tensor4d: Efficient neural 4D decomposition for high-fidelity dynamic reconstruc- tion and rendering,

    R. Shao, Z. Zheng, H. Tu, B. Liu, H. Zhang, and Y . Liu, “Tensor4d: Efficient neural 4D decomposition for high-fidelity dynamic reconstruc- tion and rendering,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2023, pp. 16 632–16 642

  6. [12]

    The lumigraph,

    S. J. Gortler, R. Grzeszczuk, R. Szeliski, and M. F. Cohen, “The lumigraph,” in Proc. 23rd Annu. Conf. Comput. Graph. Interactive Techn., 2023, pp. 453–464

  7. [13]

    3D-kernel foveated rendering for light fields,

    X. Meng, R. Du, J. F. JaJa, and A. Varshney, “3D-kernel foveated rendering for light fields,” IEEE Trans. Vis. Comput. Graph , vol. 27, no. 8, pp. 3350–3360, 2020

  8. [14]

    Light field rendering,

    M. Levoy and P. Hanrahan, “Light field rendering,” in Proc. 23rd Annu. Conf. Comput. Graph. Interactive Techn. , 2023, pp. 441–452

  9. [15]

    NeRF: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” Commun. ACM, vol. 65, no. 1, pp. 99–106, 2021

  10. [16]

    Tensorf: Tensorial radiance fields,

    A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su, “Tensorf: Tensorial radiance fields,” in Proc. Eur . Conf. Comput. Vis. Springer, 2022, pp. 333–350

  11. [17]

    Culling-based real-time rendering with accurate ray sampling for high-resolution light field 3D display,

    X.-S. Hu, X.-Y . Lin, Y .-J. Liu, M.-H. Xiang, Y .-Q. Guo, Y . Xing, and Q.- H. Wang, “Culling-based real-time rendering with accurate ray sampling for high-resolution light field 3D display,” IEEE Trans. Vis. Comput. Graph, 2024

  12. [18]

    Instant neural graphics primitives with a multiresolution hash encoding,

    T. M ¨uller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM Trans. Graph. , vol. 41, no. 4, pp. 1–15, 2022

  13. [19]

    Plenoxels: Radiance fields without neural networks,

    S. Fridovich-Keil, A. Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa, “Plenoxels: Radiance fields without neural networks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 5501– 5510

  14. [20]

    MobileNeRF: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures,

    Z. Chen, T. Funkhouser, P. Hedman, and A. Tagliasacchi, “MobileNeRF: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 16 569–16 578

  15. [21]

    A real-time method for inserting virtual objects into neural radiance fields,

    K. Ye, H. Wu, X. Tong, and K. Zhou, “A real-time method for inserting virtual objects into neural radiance fields,” IEEE Trans. Vis. Comput. Graph, 2024

  16. [22]

    Plenoctrees for real-time rendering of neural radiance fields,

    A. Yu, R. Li, M. Tancik, H. Li, R. Ng, and A. Kanazawa, “Plenoctrees for real-time rendering of neural radiance fields,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 5752–5761

  17. [23]

    Mip-NeRF: A multiscale representation for anti- aliasing neural radiance fields,

    J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan, “Mip-NeRF: A multiscale representation for anti- aliasing neural radiance fields,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2021, pp. 5855–5864

  18. [24]

    Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,

    J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, “Mip-NeRF 360: Unbounded anti-aliased neural radiance fields,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 5470– 5479

  19. [25]

    Relightable detailed human reconstruction from sparse flashlight images,

    J. Lu, T. Shao, H. Wang, Y .-L. Yang, Y . Yang, and K. Zhou, “Relightable detailed human reconstruction from sparse flashlight images,” IEEE Trans. Vis. Comput. Graph , 2024

  20. [26]

    Behind the scenes: Density fields for single view reconstruction,

    F. Wimbauer, N. Yang, C. Rupprecht, and D. Cremers, “Behind the scenes: Density fields for single view reconstruction,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 9076–9086. IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, VOL. XX, NO. XX, AUGU...

  21. [27]

    pixelNeRF: Neural radiance fields from one or few images,

    A. Yu, V . Ye, M. Tancik, and A. Kanazawa, “pixelNeRF: Neural radiance fields from one or few images,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2021, pp. 4578–4587

  22. [28]

    Scaffold- GS: Structured 3D Gaussians for view-adaptive rendering,

    T. Lu, M. Yu, L. Xu, Y . Xiangli, L. Wang, D. Lin, and B. Dai, “Scaffold- GS: Structured 3D Gaussians for view-adaptive rendering,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 20 654– 20 664

  23. [29]

    Octree- GS: Towards consistent real-time rendering with lod-structured 3D Gaussians,

    K. Ren, L. Jiang, T. Lu, M. Yu, L. Xu, Z. Ni, and B. Dai, “Octree- GS: Towards consistent real-time rendering with lod-structured 3D Gaussians,” arXiv:2403.17898, 2024

  24. [30]

    PGSR: Planar-based Gaussian splatting for efficient and high-fidelity surface reconstruction,

    D. Chen, H. Li, W. Ye, Y . Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang, “PGSR: Planar-based Gaussian splatting for efficient and high-fidelity surface reconstruction,” IEEE Trans. Vis. Comput. Graph , 2024

  25. [31]

    Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes,

    Z. Yu, T. Sattler, and A. Geiger, “Gaussian opacity fields: Efficient adaptive surface reconstruction in unbounded scenes,” ACM Trans. Graph., vol. 43, no. 6, pp. 1–13, 2024

  26. [32]

    HybridGS: Decoupling transients and statics with 2D and 3D Gaussian splatting,

    J. Lin, J. Gu, L. Fan, B. Wu, Y . Lou, R. Chen, L. Liu, and J. Ye, “HybridGS: Decoupling transients and statics with 2D and 3D Gaussian splatting,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2025, pp. 788–797

  27. [33]

    Hac: Hash-grid assisted context for 3D Gaussian splatting compression,

    Y . Chen, Q. Wu, W. Lin, M. Harandi, and J. Cai, “Hac: Hash-grid assisted context for 3D Gaussian splatting compression,” in Proc. Eur . Conf. Comput. Vis. Springer, 2024, pp. 422–438

  28. [34]

    Compact 3D Gaussian representation for radiance field,

    J. C. Lee, D. Rho, X. Sun, J. H. Ko, and E. Park, “Compact 3D Gaussian representation for radiance field,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2024, pp. 21 719–21 728

  29. [35]

    Compressed 3D Gaussian splatting for accelerated novel view synthesis,

    S. Niedermayr, J. Stumpfegger, and R. Westermann, “Compressed 3D Gaussian splatting for accelerated novel view synthesis,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 10 349– 10 358

  30. [36]

    MPGS: Multi-plane Gaussian splatting for compact scenes rendering,

    D. Li, S.-S. Huang, and H. Huang, “MPGS: Multi-plane Gaussian splatting for compact scenes rendering,” IEEE Trans. Vis. Comput. Graph, 2025

  31. [37]

    Colmap- free 3D Gaussian splatting,

    Y . Fu, S. Liu, A. Kulkarni, J. Kautz, A. A. Efros, and X. Wang, “Colmap- free 3D Gaussian splatting,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2024, pp. 20 796–20 805

  32. [38]

    LGM: Large multi-view Gaussian model for high-resolution 3D content creation,

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “LGM: Large multi-view Gaussian model for high-resolution 3D content creation,” in Proc. Eur . Conf. Comput. Vis. Springer, 2024, pp. 1–18

  33. [39]

    iVR-GS: Inverse volume rendering for explorable visualization via editable 3D Gaussian splatting,

    K. Tang, S. Yao, and C. Wang, “iVR-GS: Inverse volume rendering for explorable visualization via editable 3D Gaussian splatting,” IEEE Trans. Vis. Comput. Graph , 2025

  34. [40]

    Triplane meets Gaussian splatting: Fast and generalizable single-view 3D reconstruction with transformers,

    Z.-X. Zou, Z. Yu, Y .-C. Guo, Y . Li, D. Liang, Y .-P. Cao, and S.-H. Zhang, “Triplane meets Gaussian splatting: Fast and generalizable single-view 3D reconstruction with transformers,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 10 324–10 335

  35. [41]

    MVSNeRF: Fast generalizable radiance field reconstruction from multi- view stereo,

    A. Chen, Z. Xu, F. Zhao, X. Zhang, F. Xiang, J. Yu, and H. Su, “MVSNeRF: Fast generalizable radiance field reconstruction from multi- view stereo,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 14 124–14 133

  36. [42]

    GeoNeRF: Generalizing NeRF with geometry priors,

    M. M. Johari, Y . Lepoittevin, and F. Fleuret, “GeoNeRF: Generalizing NeRF with geometry priors,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2022, pp. 18 365–18 375

  37. [43]

    Splatter image: Ultra- fast single-view 3D reconstruction,

    S. Szymanowicz, C. Rupprecht, and A. Vedaldi, “Splatter image: Ultra- fast single-view 3D reconstruction,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 10 208–10 217

  38. [44]

    Transplat: Generalizable 3D Gaussian splatting from sparse multi-view images with transform- ers,

    C. Zhang, Y . Zou, Z. Li, M. Yi, and H. Wang, “Transplat: Generalizable 3D Gaussian splatting from sparse multi-view images with transform- ers,” in Proc. AAAI Conf. Artif. Intell. , vol. 39, no. 9, 2025, pp. 9869– 9877

  39. [45]

    GPS- Gaussian: Generalizable pixel-wise 3D Gaussian splatting for real-time human novel view synthesis,

    S. Zheng, B. Zhou, R. Shao, B. Liu, S. Zhang, L. Nie, and Y . Liu, “GPS- Gaussian: Generalizable pixel-wise 3D Gaussian splatting for real-time human novel view synthesis,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2024, pp. 19 680–19 690

  40. [46]

    pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D re- construction,

    D. Charatan, S. L. Li, A. Tagliasacchi, and V . Sitzmann, “pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D re- construction,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2024, pp. 19 457–19 467

  41. [47]

    DepthSplat: Connecting gaussian splatting and depth,

    H. Xu, S. Peng, F. Wang, H. Blum, D. Barath, A. Geiger, and M. Pollefeys, “DepthSplat: Connecting gaussian splatting and depth,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2025, pp. 16 453–16 463

  42. [48]

    MVSplat: Efficient 3D Gaussian splatting from sparse multi-view images,

    Y . Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.- J. Cham, and J. Cai, “MVSplat: Efficient 3D Gaussian splatting from sparse multi-view images,” in Proc. Eur . Conf. Comput. Vis. Springer, 2024, pp. 370–386

  43. [49]

    MVSGaussian: Fast generalizable gaussian splatting reconstruc- tion from multi-view stereo,

    T. Liu, G. Wang, S. Hu, L. Shen, X. Ye, Y . Zang, Z. Cao, W. Li, and Z. Liu, “MVSGaussian: Fast generalizable gaussian splatting reconstruc- tion from multi-view stereo,” in Proc. Eur . Conf. Comput. Vis. Springer, 2024, pp. 37–53

  44. [50]

    GS-LRM: Large reconstruction model for 3D Gaussian splatting,

    K. Zhang, S. Bi, H. Tan, Y . Xiangli, N. Zhao, K. Sunkavalli, and Z. Xu, “GS-LRM: Large reconstruction model for 3D Gaussian splatting,” in Proc. Eur . Conf. Comput. Vis. Springer, 2024, pp. 1–19

  45. [51]

    MVSNet: Depth inference for unstructured multi-view stereo,

    Y . Yao, Z. Luo, S. Li, T. Fang, and L. Quan, “MVSNet: Depth inference for unstructured multi-view stereo,” in Proc. Eur . Conf. Comput. Vis. , 2018, pp. 767–783

  46. [52]

    HexPlane: A fast representation for dynamic scenes,

    A. Cao and J. Johnson, “HexPlane: A fast representation for dynamic scenes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 130–141

  47. [53]

    K-planes: Explicit radiance fields in space, time, and appearance,

    S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa, “K-planes: Explicit radiance fields in space, time, and appearance,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 12 479–12 488

  48. [54]

    Forward flow for novel view synthesis of dynamic scenes,

    X. Guo, J. Sun, Y . Dai, G. Chen, X. Ye, X. Tan, E. Ding, Y . Zhang, and J. Wang, “Forward flow for novel view synthesis of dynamic scenes,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2023, pp. 16 022–16 033

  49. [55]

    High- fidelity and real-time novel view synthesis for dynamic scenes,

    H. Lin, S. Peng, Z. Xu, T. Xie, X. He, H. Bao, and X. Zhou, “High- fidelity and real-time novel view synthesis for dynamic scenes,” in SIGGRAPH Asia 2023 Conference Papers , 2023, pp. 1–9

  50. [56]

    DeVRF: Fast deformable voxel radiance fields for dynamic scenes,

    J.-W. Liu, Y .-P. Cao, W. Mao, W. Zhang, D. J. Zhang, J. Keppo, Y . Shan, X. Qie, and M. Z. Shou, “DeVRF: Fast deformable voxel radiance fields for dynamic scenes,” Adv. Neural Inf. Process. Syst., vol. 35, pp. 36 762– 36 775, 2022

  51. [57]

    D- NeRF: Neural radiance fields for dynamic scenes,

    A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer, “D- NeRF: Neural radiance fields for dynamic scenes,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 10 318–10 327

  52. [58]

    Neural residual radiance fields for streamably free-viewpoint videos,

    L. Wang, Q. Hu, Q. He, Z. Wang, J. Yu, T. Tuytelaars, L. Xu, and M. Wu, “Neural residual radiance fields for streamably free-viewpoint videos,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2023, pp. 76–87

  53. [59]

    NeRFPlayer: A streamable dynamic scene representation with decomposed neural radiance fields,

    L. Song, A. Chen, Z. Li, Z. Chen, L. Chen, J. Yuan, Y . Xu, and A. Geiger, “NeRFPlayer: A streamable dynamic scene representation with decomposed neural radiance fields,” IEEE Trans. Vis. Comput. Graph, vol. 29, no. 5, pp. 2732–2742, 2023

  54. [60]

    Streaming radiance fields for 3D video synthesis,

    L. Li, Z. Shen, Z. Wang, L. Shen, and P. Tan, “Streaming radiance fields for 3D video synthesis,” Adv. Neural Inf. Process. Syst. , vol. 35, pp. 13 485–13 498, 2022

  55. [61]

    HiCoM: Hierarchical coherent motion for dynamic streamable scenes with 3D Gaussian splatting,

    Q. Gao, J. Meng, C. Wen, J. Chen, and J. Zhang, “HiCoM: Hierarchical coherent motion for dynamic streamable scenes with 3D Gaussian splatting,” Adv. Neural Inf. Process. Syst. , vol. 37, pp. 80 609–80 633, 2024

  56. [62]

    Instant Gaussian stream: Fast and generalizable streaming of dynamic scene reconstruction via gaussian splatting,

    J. Yan, R. Peng, Z. Wang, L. Tang, J. Yang, J. Liang, J. Wu, and R. Wang, “Instant Gaussian stream: Fast and generalizable streaming of dynamic scene reconstruction via gaussian splatting,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2025, pp. 16 520–16 531

  57. [63]

    Motion-aware 3D Gaussian splatting for efficient dynamic scene reconstruction,

    Z. Guo, W. Zhou, L. Li, M. Wang, and H. Li, “Motion-aware 3D Gaussian splatting for efficient dynamic scene reconstruction,” IEEE Trans. Circuits Syst. Video Technol. , 2024

  58. [64]

    Dynamic 3D Gaussians: Tracking by persistent dynamic view synthesis,

    J. Luiten, G. Kopanas, B. Leibe, and D. Ramanan, “Dynamic 3D Gaussians: Tracking by persistent dynamic view synthesis,” in Proc. IEEE Int. Conf. 3D Vis. IEEE, 2024, pp. 800–809

  59. [67]

    Neural 3D video synthesis from multi-view video,

    T. Li, M. Slavcheva, M. Zollhoefer, S. Green, C. Lassner, C. Kim, T. Schmidt, S. Lovegrove, M. Goesele, R. Newcombe et al. , “Neural 3D video synthesis from multi-view video,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 5521–5531

  60. [68]

    Immersive light field video with a layered mesh representation,

    M. Broxton, J. Flynn, R. Overbeck, D. Erickson, P. Hedman, M. Duvall, J. Dourgarian, J. Busch, M. Whalen, and P. Debevec, “Immersive light field video with a layered mesh representation,” ACM Trans. Graph. , vol. 39, no. 4, pp. 86–1, 2020

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.