Pith. sign in

REVIEW 3 major objections 5 minor 35 references

LiDAR reflectance and structure-aware Salient Gaussians reconstruct self-driving scenes more accurately under complex lighting, with fewer primitives and less training time.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

LiDAR-reflectance-guided Salient Gaussians improve self-driving scene reconstruction under high ego-motion and complex lighting, beating OmniRe by 1.18 dB PSNR on Waymo Complex Lighting.

T0 review reviewed 2026-07-14 challenge →

load-bearing objection Solid systems package for multi-modal driving 3DGS; Complex Lighting claim is real in the table but the reflectance causal story is under-isolated. the 3 major comments →

arxiv 2603.12647 v3 pith:HVQJEMZ2 submitted 2026-03-13 cs.CV cs.AI

LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction

classification cs.CV cs.AI
keywords 3D Gaussian Splattingself-driving scene reconstructionLiDAR reflectanceSalient Gaussiansnovel view synthesisWaymo Open Datasetcross-modal consistency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Self-driving scene reconstruction from cameras alone fails under high ego-motion and complex lighting because RGB is lighting-sensitive and sparse views leave geometry under-constrained. This paper argues that raw LiDAR carries two underused signals—geometric structure and intensity that can be calibrated into lighting-invariant reflectance—and that both should be built into the Gaussian primitives themselves, not used only for depth or initialization. It introduces Salient Gaussians that are elongated along edges or flattened on planes, seeded from LiDAR geometric and reflectance feature points, then refined by a transform that upgrades or degrades Gaussians according to their linearity and planarity. Reflectance is attached as an extra material channel and its gradients are forced to agree with grayscale RGB gradients so material boundaries stay sharp. On Waymo sequences spanning dense traffic, high speed, complex lighting, and static scenes, the method reports higher PSNR/SSIM and lower LPIPS than strong baselines while using fewer Gaussians and finishing training faster; the largest gain is 1.18 dB PSNR under complex lighting. The practical payoff is editable, high-fidelity digital twins that can supply safer testing and richer training data for end-to-end driving models.

Core claim

Calibrating LiDAR intensity into a lighting-invariant reflectance channel, attaching it to each Gaussian, initializing structure-aware Salient Gaussians from LiDAR geometric and reflectance feature points, and jointly aligning reflectance–RGB gradients yields higher-fidelity self-driving scene reconstruction—especially under complex lighting—while reducing the number of Gaussians and training time relative to prior 3DGS methods that use LiDAR only for initialization or depth.

What carries the argument

Salient Gaussians: edge or planar ellipsoids that share a single non-dominant scale, seeded from LiDAR geometric/reflectance feature points, maintained by a linearity/planarity transform and anisotropic split, and supervised by a joint loss that matches gradient direction and normalized magnitude between rendered reflectance and grayscale RGB.

Load-bearing premise

That intensity corrected only by range and local incidence angle is accurate and lighting-invariant enough, after sparse projection, to serve as a reliable material channel whose gradients can be aligned with RGB under real sensor noise and calibration error.

What would settle it

On the same Waymo Complex Lighting sequences, re-run with deliberately corrupted incidence angles or uncorrected raw intensity; if the 1.18 dB PSNR gain over OmniRe disappears or reverses while geometry metrics stay similar, the reflectance prior is not carrying the claimed benefit.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes LR-SGS, a multi-modal 3D Gaussian Splatting pipeline for self-driving scene reconstruction. It introduces structure-aware Salient Gaussians (edge/planar primitives with a shared non-dominant scale, Eq. 4), initialized from LiDAR geometric edge/planar points (smoothness, Eq. 5) and reflectance edge points (Eq. 6), refined by a salient transform on linearity/planarity and type-aware density control (Fig. 3). LiDAR intensity is calibrated to reflectance via distance and incidence-angle correction (Eqs. 1–2) and stored as a Gaussian attribute; rendering produces color, depth, and reflectance (Eq. 7). Optimization uses RGB, LiDAR (depth + reflectance + reflectance-gradient), and a Joint Loss aligning gradient direction and normalized magnitude between reflectance and grayscale RGB (Eqs. 9–14). On 24 Waymo sequences (Dense Traffic, High-Speed, Complex Lighting, Static), LR-SGS reports better PSNR/SSIM/LPIPS than OmniRe, StreetGS, PVG and others, with fewer Gaussians and shorter training; the headline claim is +1.18 dB PSNR on Complex Lighting (30.51 vs OmniRe 29.33).

Significance. If the reported gains hold under broader evaluation, LR-SGS is a useful systems contribution to multi-modal 3DGS for autonomous driving. Salient Gaussians plus LiDAR feature initialization give a clear structural prior that improves fidelity and efficiency (Tables II–III, V; Fig. 5), and the reflectance channel with Joint Loss is a concrete way to exploit intensity beyond depth. Editable reconstructions (Fig. 1) matter for simulation and data synthesis. The work is incremental relative to OmniRe/StreetGS rather than a new representation paradigm, but the design choices are specific, ablated, and practically motivated. Strengths include consistent multi-category gains, component ablations with matching qualitative figures, and efficiency metrics (Gaussian count, training time, FPS).

major comments (3)
  1. Abstract and §I attribute the Complex Lighting gain (+1.18 dB PSNR vs OmniRe in Table I: 30.51 vs 29.33) primarily to calibrated reflectance as a lighting-invariant material channel plus Joint Loss. The only reflectance ablation (Table II, w/o Reflectance: 28.87 vs Ours 29.22, Δ≈0.35 dB) is averaged over all 24 sequences; there is no per-category breakdown isolating Complex Lighting. Salient Gaussians + LF-Init already improve structure and efficiency (Tables II–III, Fig. 5). Without a Complex-Lighting-only ablation that removes reflectance attribute and Joint Loss while keeping Salient Gaussians, the causal claim that reflectance solves complex lighting is not supported by the reported tables and may overstate the lighting-invariant prior relative to the structural prior.
  2. §IV.A evaluates 24 hand-selected Waymo sequences (6 per category) with every fourth frame held out, no multi-seed variance or error bars, and object masks taken from InvRGB+L [30]. Table I margins (e.g., Dense Traffic PSNR 28.89 vs OmniRe 28.44; Static 28.73 vs 28.23) are modest; without sequence IDs, selection criteria, or variance, it is hard to judge robustness or selection bias. A load-bearing claim of superior reconstruction across challenging self-driving scenes needs either a larger/public split, reported variance, or at least full sequence identifiers and a sensitivity check on mask quality.
  3. §III.A (Eqs. 1–2) recovers reflectance ρ from intensity after distance and incidence-angle correction using local normals from neighboring points, then projects sparse F_gt for supervision. The Complex Lighting story assumes this channel is sufficiently lighting-invariant and well-aligned under Waymo noise, calibration error, and sparse returns. The manuscript does not validate calibration quality (e.g., residual intensity–reflectance correlation under day/night, normal estimation failure cases, or sensitivity of Joint Loss Eq. 14 to projection/sparsity). If calibration is systematically biased, the material-channel prior is weaker than claimed; a short quantitative check or failure-case analysis would ground the assumption.
minor comments (5)
  1. Loss weights and τ_max/τ_min (§IV.A.3) are fixed without sensitivity analysis; a brief sweep or statement that results are stable in a neighborhood of these values would strengthen reproducibility.
  2. Fig. 4 qualitative comparisons would benefit from consistent zoom insets and identical exposure across methods, especially for Complex Lighting (c–d), so artifact differences are not confounded by display choices.
  3. Notation: F_gt / F'_gt for reflectance and gradient images vs F_G for rendered reflectance is easy to confuse with feature maps; consider ρ or R for reflectance throughout.
  4. Related work could more explicitly contrast InvRGB+L [30] and TCLC-GS [29] on how reflectance/intensity is used (material channel + joint gradient loss vs mesh/octree initialization), to clarify novelty of the Joint Loss design.
  5. Table V efficiency is on four sequences only; state whether the same four are used for all methods and whether FPS is measured under identical resolution/hardware.

Circularity Check

0 steps flagged

No circular derivation: empirical 3DGS engineering evaluated on held-out Waymo frames against external baselines; losses and hyperparameters do not force the reported PSNR by construction.

full rationale

LR-SGS is a systems/method paper: Salient Gaussian parameterization (Eq. 4), LOAM-style geometric feature extraction (Eq. 5), reflectance-edge extraction (Eq. 6), intensity-to-reflectance calibration (Eqs. 1–2 from standard LiDAR models), and composite losses (Eqs. 9–14) are design choices optimized by gradient descent. Novel-view metrics (Table I) are measured on every-fourth-frame held-out Waymo images against independent baselines (OmniRe, StreetGS, PVG, etc.). Loss weights and τ thresholds are fixed hyperparameters, not fitted to the headline +1.18 dB Complex Lighting number. Ablations (Tables II–IV) remove components and re-evaluate; none of the reported gains reduce algebraically to the training objective or to a self-citation uniqueness claim. Overlapping-author citations ([6], [19]) appear only as related work and do not underwrite the central result. No self-definitional loop, fitted-input-as-prediction, or ansatz-smuggling chain is present.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The method rests on standard 3DGS rendering, a physics-motivated but approximate intensity-to-reflectance model, hand-chosen thresholds for the salient transform, and the assumption that Waymo multi-camera/LiDAR calibration and object masks are accurate enough for joint optimization. No new physical entities are postulated; free parameters are the usual loss weights and two linearity/planarity thresholds.

free parameters (4)
  • τ_max (salient upgrade threshold) = 0.5
    Hand-set to 0.5; controls when a Non-Salient Gaussian is promoted based on linearity/planarity. Directly affects which primitives become Salient.
  • τ_min (salient degrade threshold) = 0.1
    Hand-set to 0.1; controls demotion of Salient Gaussians. Affects final Gaussian count and structure coverage.
  • loss weights λ_c, λ_depth, λ_fle, λ'_fle, λ_dir, λ_val = λ_c=λ_val=0.2; λ_depth=λ_fle=λ_dir=0.1; λ'_fle=0.05
    Hand-chosen (0.2 / 0.1 / 0.1 / 0.05 / 0.1 / 0.2). Balance photometric, depth, reflectance, and joint gradient terms; not derived from first principles.
  • Gaussian smoothing σ for Joint Loss = 1.2 px
    Fixed at 1.2 pixels before Scharr gradients; affects edge alignment strength.
axioms (4)
  • domain assumption LiDAR intensity I = η_all · ρ · cosα / R² can be inverted for reflectance ρ after estimating local normals from neighboring points (Eq. 1–2).
    Standard radiometric model used in Intensity-SLAM / RI-LIO; accuracy depends on normal quality and ignores multi-return / atmospheric effects.
  • domain assumption 3DGS α-blending (Eq. 7) plus sky compositing correctly models outdoor appearance when Gaussians are optimized with the stated losses.
    Inherited from Kerbl et al. 3DGS and OmniRe-style scene graphs; not re-derived.
  • ad hoc to paper Linearity L=(s1−s2)/s1 and planarity P=(s2−s3)/s1 of ordered scales reliably indicate edge vs planar structure for the salient transform.
    Inspired by MG-SLAM / Manhattan-world heuristics; thresholds and two-successive-evaluation rule are paper-specific.
  • domain assumption Object masks from InvRGB+L and Waymo poses/calibrations are accurate enough that dynamic/static decomposition does not dominate error.
    Implementation detail §IV.A.3; no sensitivity analysis provided.
invented entities (2)
  • Salient Gaussian (edge/planar reduced-parameter primitive with dominant direction) no independent evidence
    purpose: Capture edges and planes with fewer free scales while seeding from LiDAR feature points and allowing bidirectional transform.
    New representation relative to isotropic or fully anisotropic 3DGS; independent evidence is only the ablation tables inside this paper.
  • Reflectance attribute channel + Joint Loss (gradient direction & normalized magnitude consistency with RGB grayscale) no independent evidence
    purpose: Provide lighting-invariant material supervision and sharpen material boundaries via cross-modal edge alignment.
    Reflectance itself is classical; the specific Joint Loss formulation and attachment as a Gaussian attribute optimized jointly with RGB is paper-specific.

reviewed 2026-07-14 · how reviews work

0 comments
Cite this review

Pith. "Pith review of LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction." pith.science (2026). https://pith.science/paper/HVQJEMZ2

@misc{pith2026260312647,
  author       = {Pith},
  title        = {Pith review of: LR-SGS: Robust LiDAR-Reflectance-Guided Salient Gaussian Splatting for Self-Driving Scene Reconstruction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HVQJEMZ2}},
  note         = {Machine review of arXiv:2603.12647}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent 3D Gaussian Splatting (3DGS) methods have demonstrated the feasibility of self-driving scene reconstruction and novel view synthesis. However, most existing methods either rely solely on cameras or use LiDAR only for Gaussian initialization or depth supervision, while the rich scene information contained in point clouds, such as reflectance, and the complementarity between LiDAR and RGB have not been fully exploited, leading to degradation in challenging self-driving scenes, such as those with high ego-motion and complex lighting. To address these issues, we propose a robust and efficient LiDAR-reflectance-guided Salient Gaussian Splatting method (LR-SGS) for self-driving scenes, which introduces a structure-aware Salient Gaussian representation, initialized from geometric and reflectance feature points extracted from LiDAR and refined through a salient transform and improved density control to capture edge and planar structures. Furthermore, we calibrate LiDAR intensity into reflectance and attach it to each Gaussian as a lighting-invariant material channel, jointly aligned with RGB to enforce boundary consistency. Extensive experiments on the Waymo Open Dataset demonstrate that LR-SGS achieves superior reconstruction performance with fewer Gaussians and shorter training time. In particular, on Complex Lighting scenes, our method surpasses OmniRe by 1.18 dB PSNR.

Figures

Figures reproduced from arXiv: 2603.12647 by CM Jiang, DY Kong, F Zhu, H Zhu, XK Kuang, YJ Zhang, ZY Chen.

Figure 1
Figure 1. Figure 1: Overview of LR-SGS. Given RGB and LiDAR sequences as input, the method produces high-fidelity geometry, reflectance, and RGB renderings (left). Accurate modeling of background and objects enables realistic scene editing, including replacement and deletion (right). lighting conditions and significant ego-motion, which leads to texture inconsistencies and unstable optimization. Notably, the multi-modal data … view at source ↗
Figure 2
Figure 2. Figure 2: Method Overview. The initial scene Gaussians comprise Salient Gaussians from LiDAR feature points and Non-Salient Gaussians from SfM points. The scene is represented as a 3DGS scene graph with background, dynamic objects, and sky nodes. After obtaining the rendered Color, Depth, and Reflectance (Refle.) images, we optimize the scene parameters by minimizing a weighted sum of the Color, LiDAR, and Joint los… view at source ↗
Figure 2
Figure 2. Figure 2: Reward Evaluation in GRPO Training. (a) Standard ow-based GRPO methods evaluate generated samples under the single original condition, resulting in sparse reward mapping and insu cient inter-sample relationship exploration. (b) Our MV-GRPO leverages an augmented set of conditions to facilitate a dense multi-view mapping, fostering a comprehensive exploration of relationship among samples. the condition spa… view at source ↗
Figure 3
Figure 3. Figure 3: Our Transform and Split. The dashed line represents the Gaussian shape before executing split. ↑ and ↓ denote that the high and low threshold conditions are satisfied, respectively. In addition to geometric features, we leverage LiDAR reflectance to extract additional edge points. For point pi , we take its left and right neighboring point sets PM and PN along the same ring, and compute the reflectance gra… view at source ↗
Figure 3
Figure 3. Figure 3: Reward Ranking Varies with Conditions. Reward rankings of SDE sam￾ples across multiple semantically similar yet dierent conditions exhibit large variations, indicating that relying on a single condition for advantage estimation is inadequate. 3.2 Observation and Analysis As shown in [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative Comparison of Novel View Synthesis. (a) shows the Dense Traffic scene. (b) shows the High-Speed scene. (c) and (d) show the scene with Complex Lighting conditions. (e) shows the Static scene. Our method not only achieves high-quality reconstruction of dynamic objects, but also recovers very fine details of the background environment, maintaining consistent and stable reconstructions even under … view at source ↗
Figure 4
Figure 4. Figure 4: Overview of MV-GRPO. MV-GRPO leverages a exible Condition En￾hancer module (a pretrained VLM or LLM) to generate diverse augmented conditions for dense multi-view reward signals, facilitating comprehensive advantage estimation. Online VLM Enhancer. To dynamically capture the visual semantics of gen￾erated samples, a pretrained Vision-Language Model (VLM) is employed as an online Condition EnhancerVLME . Du… view at source ↗
Figure 5
Figure 5. Figure 5: Ablation study on Salient Gaussians. Salient Gaussians enable finer reconstruction of structural features in the environment. StreetGS and OmniRe exhibit blurred artifacts. Additionally, under high-speed that lowers co-visibility across frames, our approach is more sensitive to geometric and textural boundaries, as shown in [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 5
Figure 5. Figure 5: Distribution of Probability Drift at Dierent SDE Steps. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Ablation study on LiDAR Reflectance. We show the rendered image and depth of our method with and without LiDAR Reflectance. For clearer visual comparison, we increased the brightness in the zoomed-in regions [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 6
Figure 6. Figure 6: Reward Curves during Training. Our MV-GRPO outperforms baselines in both convergence speed and performance ceiling under various training settings [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Ablation of the Joint Loss. (a) the rendered image. (b) the rendered reflectance. Salient Gaussians improves rendering quality, accelerates convergence, and reduces training time. This improvement arises because Salient Gaussians better match edges and pla￾nar structures in the environment and require fewer param￾eters [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 7
Figure 7. Figure 7: Qualitative Comparisons with Baselines on HPS-v3. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Qualitative Comparisons with Baselines on Uni edReward-v2. [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

35 extracted references · 2 linked inside Pith

  1. [1]

    EmerneRF: Emergent spatial-temporal scene decomposition via self-supervision,

    J. Yang, B. Ivanovic, O. Litany, X. Weng, S. W. Kim, B. Li, T. Che, D. Xu, S. Fidler, M. Pavone, and Y . Wang, “EmerneRF: Emergent spatial-temporal scene decomposition via self-supervision,” inThe Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=ycv2z8TYur

  2. [2]

    3d gaussian splatting for real-time radiance field rendering,

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,”ACM Transactions on Graphics, vol. 42, no. 4, July 2023. [Online]. Available: https://repo-sam.inria.fr/fungraph/3d-gaussian-splatting/

  3. [3]

    Rgb-only gaussian splatting slam for unbounded outdoor scenes,

    S. Yu, C. Cheng, Y . Zhou, X. Yang, and H. Wang, “Rgb-only gaussian splatting slam for unbounded outdoor scenes,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 11 068–11 074

  4. [4]

    Hugs: Holistic urban 3d scene understanding via gaussian splatting,

    H. Zhou, J. Shao, L. Xu, D. Bai, W. Qiu, B. Liu, Y . Wang, A. Geiger, and Y . Liao, “Hugs: Holistic urban 3d scene understanding via gaussian splatting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 21 336–21 345

  5. [5]

    Large- scale gaussian splatting slam,

    Z. Xin, C. Wu, P. Huang, Y . Zhang, Y . Mao, and G. Huang, “Large- scale gaussian splatting slam,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 8478–8485

  6. [6]

    Fgo-slam: Enhancing gaussian slam with globally consistent opacity radiance field,

    F. Zhu, Y . Zhao, Z. Chen, B. Yu, and H. Zhu, “Fgo-slam: Enhancing gaussian slam with globally consistent opacity radiance field,” in2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 11 075–11 081

  7. [7]

    Intensity-slam: Intensity assisted localization and mapping for large scale environment,

    H. Wang, C. Wang, and L. Xie, “Intensity-slam: Intensity assisted localization and mapping for large scale environment,”IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 1715–1721, 2021

  8. [8]

    Ri- lio: Reflectivity image assisted tightly-coupled lidar-inertial odometry,

    Y . Zhang, Y . Tian, W. Wang, G. Yang, Z. Li, F. Jing, and M. Tan, “Ri- lio: Reflectivity image assisted tightly-coupled lidar-inertial odometry,” IEEE Robotics and Automation Letters, vol. 8, no. 3, pp. 1802–1809, 2023

  9. [9]

    Street gaussians: Modeling dynamic urban scenes with gaussian splatting,

    Y . Yan, H. Lin, C. Zhou, W. Wang, H. Sun, K. Zhan, X. Lang, X. Zhou, and S. Peng, “Street gaussians: Modeling dynamic urban scenes with gaussian splatting,” inECCV, 2024

  10. [10]

    Omnire: Omni urban scene reconstruction,

    Z. Chen, J. Yang, J. Huang, R. de Lutio, J. M. Esturo, B. Ivanovic, O. Litany, Z. Gojcic, S. Fidler, M. Pavone, L. Song, and Y . Wang, “Omnire: Omni urban scene reconstruction,” inThe Thirteenth Inter- national Conference on Learning Representations, 2025

  11. [11]

    Periodic vibration gaussian: Dynamic urban scene reconstruction and real-time rendering,

    Y . Chen, C. Gu, J. Jiang, X. Zhu, and L. Zhang, “Periodic vibration gaussian: Dynamic urban scene reconstruction and real-time rendering,”CoRR, vol. abs/2311.18561, 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2311.18561

  12. [12]

    Drivinggaussian: Composite gaussian splatting for surrounding dy- namic autonomous driving scenes,

    X. Zhou, Z. Lin, X. Shan, Y . Wang, D. Sun, and M.-H. Yang, “Drivinggaussian: Composite gaussian splatting for surrounding dy- namic autonomous driving scenes,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 21 634–21 643

  13. [13]

    Neural point light fields,

    J. Ost, I. Laradji, A. Newell, Y . Bahat, and F. Heide, “Neural point light fields,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18 419–18 429

  14. [14]

    Urban radiance fields,

    K. Rematas, A. Liu, P. P. Srinivasan, J. T. Barron, A. Tagliasacchi, T. Funkhouser, and V . Ferrari, “Urban radiance fields,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12 932–12 942

  15. [15]

    Real-time neural rasterization for large scenes,

    J. Y . Liu, Y . Chen, Z. Yang, J. Wang, S. Manivasagam, and R. Urtasun, “Real-time neural rasterization for large scenes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 8416–8427

  16. [16]

    Shine-mapping: Large-scale 3d mapping using sparse hierarchical implicit neural representations,

    X. Zhong, Y . Pan, J. Behley, and C. Stachniss, “Shine-mapping: Large-scale 3d mapping using sparse hierarchical implicit neural representations,” in2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 8371–8377

  17. [17]

    Urban radi- ance field representation with deformable neural mesh primitives,

    F. Lu, Y . Xu, G. Chen, H. Li, K.-Y . Lin, and C. Jiang, “Urban radi- ance field representation with deformable neural mesh primitives,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 465–476

  18. [18]

    Nerflets: Local radiance fields for efficient structure-aware 3d scene representation from 2d supervision,

    X. Zhang, A. Kundu, T. Funkhouser, L. Guibas, H. Su, and K. Genova, “Nerflets: Local radiance fields for efficient structure-aware 3d scene representation from 2d supervision,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8274–8284

  19. [19]

    Large-scale neural scene disentanglement approach for self- driving simulation,

    C. Jiang, R. Niu, Z. Chen, C. Hua, X. Kuang, X. Fu, B. Yu, and H. Zhu, “Large-scale neural scene disentanglement approach for self- driving simulation,”IEEE Transactions on Intelligent Vehicles, pp. 1– 12, 2024

  20. [20]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoor- thi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” inECCV, 2020

  21. [21]

    Mars: An instance-aware, modular and realistic simulator for autonomous driving,

    Z. Wu, T. Liu, L. Luo, Z. Zhong, J. Chen, H. Xiao, C. Hou, H. Lou, Y . Chen, R. Yanget al., “Mars: An instance-aware, modular and realistic simulator for autonomous driving,” inCAAI International Conference on Artificial Intelligence. Springer, 2023, pp. 3–15

  22. [22]

    Unisim: A neural closed-loop sensor simulator,

    Z. Yang, Y . Chen, J. Wang, S. Manivasagam, W.-C. Ma, A. J. Yang, and R. Urtasun, “Unisim: A neural closed-loop sensor simulator,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 1389–1399

  23. [23]

    Suds: Scalable urban dynamic scenes,

    H. Turki, J. Y . Zhang, F. Ferroni, and D. Ramanan, “Suds: Scalable urban dynamic scenes,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 375–12 385

  24. [24]

    Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video,

    T. Wu, F. Zhong, A. Tagliasacchi, F. Cole, and C. Oztireli, “Dˆ 2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video,”Advances in neural information processing systems, vol. 35, pp. 32 653–32 666, 2022

  25. [25]

    Real-time photorealistic dy- namic scene representation and rendering with 4d gaussian splatting,

    Z. Yang, H. Yang, Z. Pan, and L. Zhang, “Real-time photorealistic dy- namic scene representation and rendering with 4d gaussian splatting,” inICLR, 2024

  26. [26]

    4d gaussian splatting for real-time dynamic scene rendering,

    G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang, “4d gaussian splatting for real-time dynamic scene rendering,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 20 310–20 320

  27. [27]

    Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruc- tion,

    Z. Yang, X. Gao, W. Zhou, S. Jiao, Y . Zhang, and X. Jin, “Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruc- tion,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 20 331–20 341

  28. [28]

    Lidar-enhanced 3d gaussian splatting mapping,

    J. Shen, H. Yu, J. Wu, W. Yang, and G.-S. Xia, “Lidar-enhanced 3d gaussian splatting mapping,” in2025 IEEE International Conference on Robotics and Automation (ICRA), 2025, pp. 2048–2054

  29. [29]

    Tclc-gs: Tightly coupled lidar-camera gaussian splatting for autonomous driving: Supplementary materials,

    C. Zhao, S. Sun, R. Wang, Y . Guo, J.-J. Wan, Z. Huang, X. Huang, Y . V . Chen, and L. Ren, “Tclc-gs: Tightly coupled lidar-camera gaussian splatting for autonomous driving: Supplementary materials,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 91–106

  30. [30]

    Invrgb+ l: Inverse rendering of complex scenes with unified color and lidar reflectance modeling,

    X. Chen, B. Chandaka, C.-H. Lin, Y .-Q. Zhang, D. Forsyth, H. Zhao, and S. Wang, “Invrgb+ l: Inverse rendering of complex scenes with unified color and lidar reflectance modeling,”arXiv preprint arXiv:2507.17613, 2025

  31. [31]

    Object-aware bundle adjust- ment for correcting monocular scale drift,

    D. P. Frost, O. K ¨ahler, and D. W. Murray, “Object-aware bundle adjust- ment for correcting monocular scale drift,” in2016 IEEE International Conference on Robotics and Automation (ICRA), 2016, pp. 4770– 4776

  32. [32]

    Mg-slam: Structure gaussian splatting slam with manhattan world hy- pothesis,

    S. Liu, T. Deng, H. Zhou, L. Li, H. Wang, D. Wang, and M. Li, “Mg-slam: Structure gaussian splatting slam with manhattan world hy- pothesis,”IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 17 034–17 049, 2025

  33. [33]

    Loam: Lidar odometry and mapping in real-time,

    J. Zhang and S. Singh, “Loam: Lidar odometry and mapping in real-time,” inRobotics: Science and Systems, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:18612391

  34. [34]

    Scalability in perception for autonomous driving: Waymo open dataset,

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine, V . Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y . Zhang, J. Shlens, Z. Chen, and D. Anguelov, “Scalability in perception for autonomous driving: Waymo open dataset,” in2020 IEEE/CVF Conference o...

  35. [35]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595

This paper was first reviewed by grok-4.5 on July 14, 2026.