REVIEW 4 major objections 6 minor 109 references
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read GS-Occ3D claims a camera-only pipeline can reconstruct 3D occupancy from driving video, matching LiDAR-based labels on Waymo and beating them for zero-shot transfer to nuScenes.
desk verdict A real vision-only occupancy label curation pipeline with strong geometry results, but the headline claim of scalability is undercut by the deliberate exclusion of ego-static scenes that the authors admit the method cannot handle. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Octree-based Gaussian Surfel: a hierarchical octree of voxels, each holding up to m Gaussian surfels that locally approximate surface geometry, with level-of-detail that captures coarse structures like walls and roads alongside fine details like vegetation and object edges. Ground is handled by dedicated Gaussian surfels initialized by projecting camera poses onto the xy-plane with a fixed height offset, then regularized to stay planar; dynamic vehicles are initialized from a vision-based 3D tracker, with learnable pose corrections refining each box. The label-curation stage turns the reconstructed point cloud into training labels by slicing per-frame ranges, aggregating dynamic-object points in box coordinates, and ray-casting so only first-hit voxels are marked as observed, which separates free space from unobserved space.
What would settle it
Measure the IoU of CVT-Occ trained on GS-Occ3D labels versus Occ3D labels on the full Occ3D-Waymo validation set, including the ego-static scenes that the paper excludes; if the vision-only labels do worse than LiDAR labels when those scenes are kept, the claimed matching generalization is broken.
Extended reading notes
Core claim
The paper's central claim is that a vision-only pipeline can replace LiDAR-based annotation for downstream occupancy models, and that the bottleneck is not the sensor but the scene representation. GS-Occ3D optimizes an octree-based Gaussian Surfel field over each long driving sequence, with explicit ground surfels and separately tracked dynamic vehicles, producing dense point clouds that are converted into occupancy labels by frame-wise division, multi-frame aggregation, and ray-casting. Experiments on Waymo report state-of-the-art geometry reconstruction (CD 0.56) and, when the resulting labels train CVT-Occ, generalization on Occ3D-Waymo that is described as reasonable and overall comparable to LiDAR labels, with superior zero-shot IoU on nuScenes. The authors state this is the first vision-only reconstruction of the full Waymo dataset.
Load-bearing premise
The evaluation assumes that removing ego-static scenes is a fair way to compare labels, even though those are exactly the scenes where vision-only reconstruction is known to fail.
Editorial extensions
If this is right
- If the method scales as claimed, occupancy label curation no longer requires LiDAR-equipped fleets; ordinary camera-equipped vehicles can supply auto-labeling data.
- Downstream occupancy models trained on vision-only labels inherit wider coverage than LiDAR labels, including regions beyond LiDAR range and tall structures such as high-rise buildings.
- Zero-shot generalization on Occ3D-nuScenes indicates that vision-only labels are less tied to a particular sensor configuration or scene distribution than LiDAR-based labels.
- Reconstructing the full Waymo dataset provides a large camera-only label base for future occupancy and scene-completion research.
- The richer 66-class semantics available from RGB segmentation compared with Occ3D's 16 classes could make vision-only labels useful beyond geometry, for objects like motorcycles, lane markings, and crosswalks.
Reading between the lines
- If vision-only reconstruction proves robust to different camera rigs and weather, the same octree-surfel decomposition could be adapted to other offboard auto-labeling tasks, such as lane geometry extraction or object-level mesh generation.
- The paper's admitted failure on ego-static scenes suggests a practical hybrid: keep vision-only reconstruction for moving-vehicle stretches and add a lightweight motion estimator or multi-view prior only for stationary segments.
- A direct test of the scalability claim would be to reconstruct labels from a different camera rig, such as nuScenes images, and evaluate on Occ3D-Waymo, checking whether the zero-shot advantage is symmetric across sensor setups.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GS-Occ3D, a vision-only framework for reconstructing 3D occupancy from multi-view driving sequences. The method initializes an octree-based Gaussian surfel representation from SfM point clouds, explicitly models ground planes and dynamic vehicles, and trains with photometric and geometric losses. After reconstructing the Waymo dataset, the authors curate binary occupancy labels via frame-wise division, multi-frame aggregation, and ray-casting, then use these labels to train the CVT-Occ downstream occupancy model. Experiments report state-of-the-art Chamfer Distance on the Waymo Static-32 geometry split (0.56 vs. 0.82 for the best explicit baseline), comparable downstream occupancy performance on a restricted Occ3D-Waymo validation set (44.7 vs. 57.4 IoU), superior zero-shot generalization on Occ3D-nuScenes (33.4 vs. 31.4 IoU), and favorable fitting results on the authors' own label sets.
Significance. If the central claim holds, the paper introduces a scalable vision-only alternative to LiDAR-based occupancy label curation, potentially enabling crowdsourced data to be used for auto-labeling at scale. The work is significant in scope: it reports reconstructing the full Waymo dataset, and the downstream experiment is a direct comparison of vision-only labels against an established LiDAR-based benchmark. The main strengths are the clear decomposition into static, ground, and dynamic components; the efficient octree-surfel representation (80 MB storage, 0.8 h training per scene); and the explicit downstream validation with a strong occupancy model. The geometry result is substantial, and the zero-shot nuScenes improvement, even if modest, is a concrete and falsifiable positive result. The main weakness is the benchmark restriction to non-ego-static scenes, which is acknowledged by the authors and must be resolved before the headline claims can be taken at face value.
major comments (4)
- [Section 4.3, Datasets; Table 2] The exclusion of ego-static scenes from both training and evaluation is the most load-bearing assumption in the paper. The text states that vision-only methods 'cannot be reliably handled' on these scenes and that they are excluded 'for fairness', leaving 637/798 training and 165/202 validation scenes. Since the Occ3D-baseline model is not given this advantage, the Table 2 comparison of 44.7 vs. 57.4 IoU on 'Occ3D-Val(Waymo)' is not a comparison on the full official benchmark, and the abstract's claim of effectiveness on Occ3D-Waymo is therefore overstated. Please report results on the full official Occ3D-Waymo validation split, or at least provide a per-split breakdown (ego-static vs. non-ego-static), so the reader can assess the impact of this selection. The current presentation makes it impossible to determine whether the vision-only labels are comparable on the full distribution or only on a favorable subset.
- [Section 4.2; Table 1] The Chamfer Distance metric is computed against LiDAR point clouds, and the paper itself notes that reconstruction beyond the LiDAR coverage is counted as error ('our reasonable reconstruction beyond the LiDAR coverage is also counted as regions with high CD'). This conflation of range with inaccuracy undermines the single-number 'SOTA geometry' claim in Table 1. Please report CD restricted to a common evaluation mask, e.g., voxels or points within the LiDAR field of view, or provide a range-limited precision/recall curve. Without such a restriction, the 0.56 CD value may reflect both genuine geometric accuracy and the advantage of not emitting points in regions where LiDAR has no ground truth.
- [Section 3.1; Equations (4)-(5)] The experimental section does not specify the values of the key hyperparameters needed to reproduce the reported results, including the number of Gaussian surfels per voxel m, the base voxel size epsilon, the loss weights lambda_geo, lambda_obj, lambda_road, lambda_sky, lambda_s, lambda_d, and lambda_n, the learning rates, and the number of optimization iterations. Since no code release is mentioned, these omitted values are load-bearing for the paper's quantitative claims. Please add a supplementary table with all training hyperparameters and the exact protocol used for the Waymo Static-32 and full-dataset reconstructions.
- [Section 4.5; Table 3] Table 3 states that all methods use 'our Ground Gaussians for fairness'. The ground Gaussian component is a core novel contribution of the proposed method, and giving it to PGSR, 2DGS, and GVKF changes the comparison from 'our method vs. the baselines as published' to 'our method vs. baselines augmented with one of our components'. This could be fair as a controlled study, but it also makes the SOTA claim sensitive to an unreported interaction: if the ground gaussians are more compatible with some representations than others, the ranking in Table 3 may not reflect the methods' intrinsic performance. Please provide an additional comparison where baselines use their own default initialization/ground handling, or justify in more detail why the ground gaussians constitute a neutral initialization for all methods.
minor comments (6)
- [Section 4.1] The 'Waymo Static-32' split is referenced to [88] but never defined in the paper; please state the number of scenes, the selection criterion, and the split composition so the geometry comparison is self-contained.
- [Table 1] The F2-NeRF Chamfer Distance of 886.77 is several orders of magnitude larger than all other values in the table; please verify whether this is a typo, a unit issue, or the result of a catastrophic failure, and clarify it in the text.
- [Section 4.3; Table 2] F1, Precision, and Recall are reported in Table 2 but are not defined in the metrics paragraph; please specify how these are computed for voxel occupancy (e.g., binary voxel-level scores over the visible region).
- [References and author list] Several reference titles contain incorrect spacing ('V oteflow', 'V oxformer'), and the author name 'Moonjun Goon' in the author list differs from 'Moonjun Gong' in reference [41]; please harmonize these.
- [Section 3.2, Multi-frame Aggregation] The description 'following a process similar to [71]' is vague for the coordinate transformation used to aggregate dynamic-object points; please specify the box coordinate system, the association of points to boxes, and how occluded or mis-associated points are filtered.
- [Tables 1 and 3] The 3Cam rows for PGSR, 2DGS, and GVKF in Table 3 appear to coincide with the values in Table 1, but Table 1 does not state which camera configuration was used; please clarify the relationship between the two tables.
Circularity Check
No significant circularity: the proposed vision-only labels are validated against independent LiDAR-based Occ3D benchmarks, and no fitted quantity is renamed as a prediction.
full rationale
The paper's derivation chain is: detector-free SfM and ground gaussians initialize an octree Gaussian-surfel reconstruction; the reconstructed point cloud is divided into frames, aggregated, and voxelized; CVT-Occ is trained on those voxel labels; and the trained model is evaluated on the external Occ3D-Waymo and Occ3D-nuScenes validation sets whose labels are LiDAR-derived. None of the equations (1)-(5) or the curation steps define the predicted occupancy in terms of Occ3D labels, and no parameter is fitted to the evaluation target. The rows in Table 2 that use Ours-Val/Ours-Train as evaluation labels are explicitly presented as the paper's own labels; the load-bearing comparisons are the rows against Occ3D-Val(Waymo) (44.7 vs 57.4 IoU) and Occ3D-Val(nuScenes) (33.4 vs 31.4 IoU), which use independent LiDAR-based ground truth. The self-citations (Occ3D [71], CVT-Occ [94]) are used as benchmark and baseline model, not as the justification for the reconstruction; the octree and dynamic-scene inspirations cite external groups ([62], [87]). The manuscript's Limitations point (3) and Section 4.3 admit that ego-static scenes are excluded because vision-only reconstruction cannot reliably handle them, and this is a real benchmark-fairness concern for the Waymo comparison (the full validation set is 202 scenes, of which 165 are used), but it is not circularity: the exclusion is a selection criterion, not a fitted parameter, and the nuScenes zero-shot test is independent of that selection. No circular step can be quoted, so the correct finding is no significant circularity.
Assumptions & free parameters
free parameters (5)
- Number of Gaussian surfels per voxel, m =
not specified
- Base voxel size epsilon =
not specified
- Ground height offset for ground surfels =
not specified
- Loss weights lambda_geo, lambda_obj, lambda_road, lambda_sky, lambda_s, lambda_d, lambda_n =
not specified
- Frame-wise perception range and target point count per sweep =
not specified
assumptions (6)
- domain assumption Road is approximately parallel to the camera plane, and ground elevation can be recovered from the nearest camera pose with a fixed height offset.
- domain assumption Dynamic objects are vehicles with 3D bounding boxes from a vision-based tracker, and all sampled points inside a box belong to that object.
- domain assumption The SfM sparse point cloud is a reliable scene skeleton and coverage proxy for occupancy.
- domain assumption LiDAR point clouds serve as ground truth for geometric accuracy, and geometry outside LiDAR coverage can still be scored by CD.
- ad hoc to paper Excluding ego-static scenes from training and evaluation is a fair way to compare vision-only labels with LiDAR labels.
- domain assumption Binary, non-semantic occupancy labels are sufficient to train and evaluate the downstream occupancy model.
Cite this review
Pith. "Pith review of GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting." pith.science (2026). https://pith.science/paper/SN6AQMZX
@misc{pith2026250719451,
author = {Pith},
title = {Pith review of: GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/SN6AQMZX}},
note = {Machine review of arXiv:2507.19451}
}
read the original abstract
Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents leveraging vast amounts of potential crowdsourced data for auto-labeling. To address this, we propose GS-Occ3D, a scalable vision-only framework that directly reconstructs occupancy. Vision-only occupancy reconstruction poses significant challenges due to sparse viewpoints, dynamic scene elements, severe occlusions, and long-horizon motion. Existing vision-based methods primarily rely on mesh representation, which suffer from incomplete geometry and additional post-processing, limiting scalability. To overcome these issues, GS-Occ3D optimizes an explicit occupancy representation using an Octree-based Gaussian Surfel formulation, ensuring efficiency and scalability. Additionally, we decompose scenes into static background, ground, and dynamic objects, enabling tailored modeling strategies: (1) Ground is explicitly reconstructed as a dominant structural element, significantly improving large-area consistency; (2) Dynamic vehicles are separately modeled to better capture motion-related occupancy patterns. Extensive experiments on the Waymo dataset demonstrate that GS-Occ3D achieves state-of-the-art geometry reconstruction results. By curating vision-only binary occupancy labels from diverse urban scenes, we show their effectiveness for downstream occupancy models on Occ3D-Waymo and superior zero-shot generalization on Occ3D-nuScenes. It highlights the potential of large-scale vision-based occupancy reconstruction as a new paradigm for scalable auto-labeling. Project Page: https://gs-occ3d.github.io/
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. In ICCV, 2021. 2
2021
-
[2]
Se- mantickitti: A dataset for semantic scene understanding of lidar sequences
Jens Behley, Martin Garbade, Andres Milioto, Jan Quen- zel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. Se- mantickitti: A dataset for semantic scene understanding of lidar sequences. In Proceedings of the IEEE/CVF inter- national conference on computer vision, pages 9297–9307,
-
[3]
Langocc: Self-supervised open vocabulary occupancy estimation via volume rendering
Simon Boeder, Fabian Gigengack, and Benjamin Risse. Langocc: Self-supervised open vocabulary occupancy estimation via volume rendering. arXiv preprint arXiv:2407.17310, 2024. 3
arXiv 2024
-
[4]
Simon Boeder, Fabian Gigengack, and Benjamin Risse. Gaussianflowocc: Sparse and weakly supervised occu- pancy estimation using gaussian splatting and temporal flow. arXiv preprint arXiv:2502.17288, 2025. 3
arXiv 2025
-
[5]
nuscenes: A mul- timodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A mul- timodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621–11631, 2020. 3
2020
-
[6]
Monoscene: Monocular 3d semantic scene completion
Anh-Quan Cao and Raoul De Charette. Monoscene: Monocular 3d semantic scene completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 3991–4001, 2022. 3
2022
-
[7]
Gaussrender: Learning 3d occupancy with gaussian rendering, 2025
Loïck Chambon, Eloi Zablocki, Alexandre Boulch, Mick- aël Chen, and Matthieu Cord. Gaussrender: Learning 3d occupancy with gaussian rendering, 2025. 3
2025
-
[8]
Argoverse: 3d tracking and forecasting with rich maps
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jag- jeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Pe- ter Carr, Simon Lucey, Deva Ramanan, et al. Argoverse: 3d tracking and forecasting with rich maps. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8748–8757, 2019. 3
2019
Show all 109 references
-
[9]
Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction
Danpeng Chen, Hai Li, Weicai Ye, Yifan Wang, Weijian Xie, Shangjin Zhai, Nan Wang, Haomin Liu, Hujun Bao, and Guofeng Zhang. Pgsr: Planar-based gaussian splatting for efficient and high-fidelity surface reconstruction. arXiv preprint arXiv:2406.06521, 2024. 2, 5, 6, 8
2024 arXiv
-
[10]
Dogaussian: Distributed- oriented gaussian splatting for large-scale 3d reconstruction via gaussian consensus
Yu Chen and Gim Hee Lee. Dogaussian: Distributed- oriented gaussian splatting for large-scale 3d reconstruction via gaussian consensus. arXiv preprint arXiv:2405.13943,
-
[11]
Omnire: Omni urban scene reconstruction
Ziyu Chen, Jiawei Yang, Jiahui Huang, Riccardo de Lutio, Janick Martinez Esturo, Boris Ivanovic, Or Litany, Zan Go- jcic, Sanja Fidler, Marco Pavone, Li Song, and Yue Wang. Omnire: Omni urban scene reconstruction. arXiv preprint arXiv:2408.16760, 2024. 2
2024 arXiv
-
[12]
Trackocc: Camera-based 4d panop- tic occupancy tracking
Zhuoguang Chen, Kenan Li, Xiuyu Yang, Tao Jiang, Yim- ing Li, and Hang Zhao. Trackocc: Camera-based 4d panop- tic occupancy tracking. arXiv preprint arXiv:2503.08471,
-
[13]
Long3r: Long sequence streaming 3d re- construction, 2025
Zhuoguang Chen, Minghui Qin, Tianyuan Yuan, Zhe Liu, and Hang Zhao. Long3r: Long sequence streaming 3d re- construction, 2025. 8
2025
-
[14]
Mask2former for video instance segmentation
Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexan- der Kirillov, Rohit Girdhar, and Alexander G Schwing. Mask2former for video instance segmentation. arXiv preprint arXiv:2112.10764, 2021. 3, 8
2021 arXiv
-
[15]
Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaus- sian splatting
Xiao Cui, Weicai Ye, Yifan Wang, Guofeng Zhang, Wen- gang Zhou, and Houqiang Li. Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaus- sian splatting. arXiv preprint arXiv:2410.04354, 2024. 2
2024 arXiv
-
[16]
High-quality surface re- construction using gaussian surfels
Pinxuan Dai, Jiamin Xu, Wenxiang Xie, Xinguo Liu, Huamin Wang, and Weiwei Xu. High-quality surface re- construction using gaussian surfels. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024. 2
2024
-
[17]
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow tran...
2024
-
[18]
Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds
Zhiwen Fan, Wenyan Cong, Kairun Wen, Kevin Wang, Jian Zhang, Xinghao Ding, Danfei Xu, Boris Ivanovic, Marco Pavone, Georgios Pavlakos, et al. Instantsplat: Un- bounded sparse-view pose-free gaussian splatting in 40 sec- onds. arXiv preprint arXiv:2403.20309, 2(3):4, 2024. 8
2024 arXiv
-
[19]
Rogs: Large scale road surface reconstruction with meshgrid gaussian
Zhiheng Feng, Wenhua Wu, Tianchen Deng, and Hesh- eng Wang. Rogs: Large scale road surface reconstruction with meshgrid gaussian. arXiv preprint arXiv:2405.14342,
-
[20]
Cc-3dt: Panoramic 3d object tracking via cross-camera fusion
Tobias Fischer, Yung-Hsu Yang, Suryansh Kumar, Min Sun, and Fisher Yu. Cc-3dt: Panoramic 3d object tracking via cross-camera fusion. arXiv preprint arXiv:2212.01247,
-
[21]
Sugar: Surface- aligned gaussian splatting for efficient 3d mesh recon- struction and high-quality mesh rendering
Antoine Guédon and Vincent Lepetit. Sugar: Surface- aligned gaussian splatting for efficient 3d mesh recon- struction and high-quality mesh rendering. arXiv preprint arXiv:2311.12775, 2023. 2
2023 arXiv
-
[22]
Streetsurf: Extending multi-view im- plicit surface reconstruction to street views
Jianfei Guo, Nianchen Deng, Xinyang Li, Yeqi Bai, Bo- tian Shi, Chiyu Wang, Chenjing Ding, Dongliang Wang, and Yikang Li. Streetsurf: Extending multi-view im- plicit surface reconstruction to street views. arXiv preprint arXiv:2306.04988, 2023. 2, 5, 6
2023 arXiv
-
[23]
Dist-4d: Disentangled spatiotemporal diffusion with metric depth for 4d driving scene generation
Jiazhe Guo, Yikang Ding, Xiwu Chen, Shuo Chen, Bohan Li, Yingshuang Zou, Xiaoyang Lyu, Feiyang Tan, Xiaojuan Qi, Zhiheng Li, et al. Dist-4d: Disentangled spatiotemporal diffusion with metric depth for 4d driving scene generation. arXiv preprint arXiv:2503.15208, 2025. 2
2025 arXiv
-
[24]
Detector-free struc- ture from motion, 2023
Xingyi He, Jiaming Sun, Yifan Wang, Sida Peng, Qixing Huang, Hujun Bao, and Xiaowei Zhou. Detector-free struc- ture from motion, 2023. 3
2023
-
[25]
Splatad: Real-time li- dar and camera rendering with 3d gaussian splatting for au- tonomous driving, 2024
Georg Hess, Carl Lindström, Maryam Fatemi, Christoffer Petersson, and Lennart Svensson. Splatad: Real-time li- dar and camera rendering with 3d gaussian splatting for au- tonomous driving, 2024. 2
2024
-
[26]
Denoising dif- fusion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models, 2020. 2
2020
-
[27]
2d gaussian splatting for geometrically ac- curate radiance fields
Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2d gaussian splatting for geometrically ac- curate radiance fields. arXiv preprint arXiv:2403.17888 ,
-
[28]
Tri-perspective view for vision- based 3d semantic occupancy prediction
Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou, and Jiwen Lu. Tri-perspective view for vision- based 3d semantic occupancy prediction. In Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition, pages 9223–9232, 2023. 3
2023
-
[29]
Selfocc: Self-supervised vision-based 3d oc- cupancy prediction
Yuanhui Huang, Wenzhao Zheng, Borui Zhang, Jie Zhou, and Jiwen Lu. Selfocc: Self-supervised vision-based 3d oc- cupancy prediction. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 19946–19956, 2024. 3
2024
-
[30]
Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy prediction
Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou, and Jiwen Lu. Gaussianformer: Scene as gaussians for vision-based 3d semantic occupancy prediction. In Eu- ropean Conference on Computer Vision , pages 376–393. Springer, 2024. 1, 3
2024
-
[31]
Gausstr: Foundation model-aligned gaussian transformer for self-supervised 3d spatial understanding
Haoyi Jiang, Liu Liu, Tianheng Cheng, Xinjie Wang, Tian- wei Lin, Zhizhong Su, Wenyu Liu, and Xinggang Wang. Gausstr: Foundation model-aligned gaussian transformer for self-supervised 3d spatial understanding. arXiv preprint arXiv:2412.13193, 2024. 3
2024 arXiv
-
[32]
Horizon-gs: Unified 3d gaussian splatting for large-scale aerial-to-ground scenes, 2024
Lihan Jiang, Kerui Ren, Mulin Yu, Linning Xu, Junt- ing Dong, Tao Lu, Feng Zhao, Dahua Lin, and Bo Dai. Horizon-gs: Unified 3d gaussian splatting for large-scale aerial-to-ground scenes, 2024. 2
2024
-
[33]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph., 42(4):139–1,
-
[34]
A hierarchical 3d gaussian representation for real- time rendering of very large datasets
Bernhard Kerbl, Andreas Meuleman, Georgios Kopanas, Michael Wimmer, Alexandre Lanvin, and George Dret- takis. A hierarchical 3d gaussian representation for real- time rendering of very large datasets. ACM Transactions on Graphics (TOG), 43(4):1–15, 2024. 2
2024
-
[35]
Grounding image matching in 3d with mast3r
Vincent Leroy, Yohann Cabon, and Jérôme Revaud. Grounding image matching in 3d with mast3r. InEuropean Conference on Computer Vision , pages 71–91. Springer,
-
[36]
Uniscene: Unified occupancy-centric driving scene generation
Bohan Li, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Chi Zhang, Tiancai Wang, et al. Uniscene: Unified occupancy-centric driving scene generation. arXiv preprint arXiv:2412.05435, 2024. 2
2024 arXiv
-
[37]
Occmamba: Semantic occu- pancy prediction with state space models
Heng Li, Yuenan Hou, Xiaohan Xing, Yuexin Ma, Xiao Sun, and Yanyong Zhang. Occmamba: Semantic occu- pancy prediction with state space models. InProceedings of the Computer Vision and Pattern Recognition Conference , pages 11949–11959, 2025. 3
2025
-
[38]
Mtgs: Multi- traversal gaussian splatting, 2025
Tianyu Li, Yihang Qiu, Zhenhua Wu, Carl Lindström, Peng Su, Matthias Nießner, and Hongyang Li. Mtgs: Multi- traversal gaussian splatting, 2025. 2
2025
-
[39]
Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo
Yinhao Li, Han Bao, Zheng Ge, Jinrong Yang, Jianjian Sun, and Zeming Li. Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo. InPro- ceedings of the AAAI Conference on Artificial Intelligence, pages 1486–1494, 2023. 3
2023
-
[40]
V oxformer: Sparse voxel transformer for camera-based 3d semantic scene completion
Yiming Li, Zhiding Yu, Christopher Choy, Chaowei Xiao, Jose M Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar. V oxformer: Sparse voxel transformer for camera-based 3d semantic scene completion. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern ...
2023
-
[41]
Sscbench: A large-scale 3d semantic scene completion benchmark for autonomous driving
Yiming Li, Sihang Li, Xinhao Liu, Moonjun Gong, Ke- nan Li, Nuo Chen, Zijun Wang, Zhiheng Li, Tao Jiang, Fisher Yu, et al. Sscbench: A large-scale 3d semantic scene completion benchmark for autonomous driving. In 2024 IEEE/RSJ International Conference on Intelligent Robots and...
2024
-
[42]
Fb-occ: 3d occu- pancy prediction based on forward-backward view trans- formation
Zhiqi Li, Zhiding Yu, David Austin, Mingsheng Fang, Shiyi Lan, Jan Kautz, and Jose M Alvarez. Fb-occ: 3d occu- pancy prediction based on forward-backward view trans- formation. arXiv preprint arXiv:2307.01492, 2023. 3
2023 arXiv
-
[43]
Bevformer: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers.IEEE Transactions on Pat- tern Analysis and Machine Intelligence, 2024
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chong- hao Sima, Tong Lu, Qiao Yu, and Jifeng Dai. Bevformer: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers.IEEE Transactions on Pat- tern Analysis and Machine Intelligence, 2024. 3
2024
-
[44]
Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d
Yiyi Liao, Jun Xie, and Andreas Geiger. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, 45(3):3292–3310, 2022. 3
2022
-
[45]
Vastgaussian: Vast 3d gaussians for large scene reconstruction
Jiaqi Lin, Zhihao Li, Xiao Tang, Jianzhuang Liu, Shiyong Liu, Jiayue Liu, Yangdi Lu, Xiaofei Wu, Songcen Xu, You- liang Yan, et al. Vastgaussian: Vast 3d gaussians for large scene reconstruction. arXiv preprint arXiv:2402.17427 ,
-
[46]
V oteflow: Enforcing local rigidity in self-supervised scene flow, 2025
Yancong Lin, Shiming Wang, Liangliang Nan, Julian Kooij, and Holger Caesar. V oteflow: Enforcing local rigidity in self-supervised scene flow, 2025. 2
2025
-
[47]
Lidar-based 4d occu- pancy completion and forecasting
Xinhao Liu, Moonjun Gong, Qi Fang, Haoyu Xie, Yim- ing Li, Hang Zhao, and Chen Feng. Lidar-based 4d occu- pancy completion and forecasting. In 2024 IEEE/RSJ In- ternational Conference on Intelligent Robots and Systems (IROS), pages 11102–11109. IEEE, 2024. 1, 3
2024
-
[48]
Citygaussian: Real-time high-quality large-scale scene rendering with gaussians
Yang Liu, Chuanchen Luo, Lue Fan, Naiyan Wang, Jun- ran Peng, and Zhaoxiang Zhang. Citygaussian: Real-time high-quality large-scale scene rendering with gaussians. In European Conference on Computer Vision, pages 265–282. Springer, 2024. 2
2024
-
[49]
Let occ flow: Self- supervised 3d occupancy flow prediction
Yili Liu, Linzhan Mou, Xuan Yu, Chenrui Han, Sitong Mao, Rong Xiong, and Yue Wang. Let occ flow: Self- supervised 3d occupancy flow prediction. arXiv preprint arXiv:2407.07587, 2024. 3
2024 arXiv
-
[50]
Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation
Zhijian Liu, Haotian Tang, Alexander Amini, Xinyu Yang, Huizi Mao, Daniela L Rus, and Song Han. Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation. In 2023 IEEE international conference on robotics and automation (ICRA), pages 2774–2781. IEEE,
2023
-
[51]
Urban radiance field represen- tation with deformable neural mesh primitives
Fan Lu, Yan Xu, Guang Chen, Hongsheng Li, Kwan-Yee Lin, and Changjun Jiang. Urban radiance field represen- tation with deformable neural mesh primitives. In ICCV, pages 465–476, 2023. 2
2023
-
[52]
Scaffold-gs: Structured 3d gaussians for view-adaptive rendering
Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. Scaffold-gs: Structured 3d gaussians for view-adaptive rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20654–20664, 2024. 3
2024
-
[53]
Infinicube: Un- bounded and controllable dynamic 3d driving scene gen- eration with world-guided video models, 2024
Yifan Lu, Xuanchi Ren, Jiawei Yang, Tianchang Shen, Zhangjie Wu, Jun Gao, Yue Wang, Siheng Chen, Mike Chen, Sanja Fidler, and Jiahui Huang. Infinicube: Un- bounded and controllable dynamic 3d driving scene gen- eration with world-guided video models, 2024. 2
2024
-
[54]
Zopp: A framework of zero-shot offboard panoptic perception for autonomous driving, 2024
Tao Ma, Hongbin Zhou, Qiusheng Huang, Xuemeng Yang, Jianfei Guo, Bo Zhang, Min Dou, Yu Qiao, Botian Shi, and Hongsheng Li. Zopp: A framework of zero-shot offboard panoptic perception for autonomous driving, 2024. 1
2024
-
[55]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. 65(1):99–106, 2021. 2
2021
-
[56]
Instant neural graphics primitives with a multiresolution hash encoding
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM TOG, 2022. 2
2022
-
[57]
Scaling diffusion models to real-world 3d lidar scene completion
Lucas Nunes, Rodrigo Marcuzzi, Benedikt Mersch, Jens Behley, and Cyrill Stachniss. Scaling diffusion models to real-world 3d lidar scene completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14770–14780, 2024. 2
2024
-
[58]
Neural scene graphs for dynamic scenes
Julian Ost, Fahim Mannan, Nils Thuerey, Julian Knodt, and Felix Heide. Neural scene graphs for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 2856–2865, 2021. 2
2021
-
[59]
Co-occ: Coupling explicit feature fusion with volume rendering regularization for multi-modal 3d semantic occupancy prediction
Jingyi Pan, Zipeng Wang, and Lin Wang. Co-occ: Coupling explicit feature fusion with volume rendering regularization for multi-modal 3d semantic occupancy prediction. IEEE Robotics and Automation Letters, 2024. 3
2024
-
[60]
Renderocc: Vision-centric 3d occupancy pre- diction with 2d rendering supervision
Mingjie Pan, Jiaming Liu, Renrui Zhang, Peixiang Huang, Xiaoqi Li, Hongwei Xie, Bing Wang, Li Liu, and Shang- hang Zhang. Renderocc: Vision-centric 3d occupancy pre- diction with 2d rendering supervision. In 2024 IEEE Inter- national Conference on Robotics and Automation (ICRA...
2024
-
[61]
Time will tell: New outlooks and a baseline for temporal multi- view 3d object detection
Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, Kris Kitani, Masayoshi Tomizuka, and Wei Zhan. Time will tell: New outlooks and a baseline for temporal multi- view 3d object detection. arXiv preprint arXiv:2210.02443,
-
[62]
Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians
Kerui Ren, Lihan Jiang, Tao Lu, Mulin Yu, Linning Xu, Zhangkai Ni, and Bo Dai. Octree-gs: Towards consistent real-time rendering with lod-structured 3d gaussians. arXiv preprint arXiv:2403.17898, 2024. 3
2024 arXiv
-
[63]
Lmscnet: Lightweight multiscale 3d semantic completion
Luis Roldao, Raoul De Charette, and Anne Verroust- Blondet. Lmscnet: Lightweight multiscale 3d semantic completion. In 2020 International Conference on 3D Vi- sion (3DV), pages 111–119. IEEE, 2020. 3
2020
-
[64]
Gvkf: Gaus- sian voxel kernel functions for highly efficient surface re- construction in open scenes
Gaochao Song, Chong Cheng, and Hao Wang. Gvkf: Gaus- sian voxel kernel functions for highly efficient surface re- construction in open scenes. Advances in Neural Informa- tion Processing Systems, 37:104792–104815, 2025. 2, 5, 6, 8
2025
-
[65]
Coda-4dgs: Dynamic gaussian splatting with context and deformation awareness for autonomous driving
Rui Song, Chenwei Liang, Yan Xia, Walter Zimmer, Hu Cao, Holger Caesar, Andreas Festag, and Alois Knoll. Coda-4dgs: Dynamic gaussian splatting with context and deformation awareness for autonomous driving. arXiv preprint arXiv:2503.06744, 2025. 2
2025 arXiv
-
[66]
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser. Semantic scene completion from a single depth image. In Proceedings of the IEEE conference on computer vision and pattern recog- nition, pages 1746–1754, 2017. 3
2017
-
[67]
Don’t shake the wheel: Momentum- aware planning in end-to-end autonomous driving
Ziying Song, Caiyan Jia, Lin Liu, Hongyu Pan, Yongchang Zhang, Junming Wang, Xingyu Zhang, Shaoqing Xu, Lei Yang, and Yadan Luo. Don’t shake the wheel: Momentum- aware planning in end-to-end autonomous driving. In Pro- ceedings of the Computer Vision and Pattern Recognition Co...
2025
-
[68]
Loftr: Detector-free local feature matching with transformers
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 8922–8931, 2021. 3
2021
-
[69]
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aure- lien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer ...
2020
-
[70]
Block-nerf: Scalable large scene neural view synthesis
Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In CVPR, pages 8248– 8258, 2022. 2
2022
-
[71]
Occ3d: A large-scale 3d occupancy prediction benchmark for autonomous driving
Xiaoyu Tian, Tao Jiang, Longfei Yun, Yucheng Mao, Huitong Yang, Yue Wang, Yilun Wang, and Hang Zhao. Occ3d: A large-scale 3d occupancy prediction benchmark for autonomous driving. Advances in Neural Information Processing Systems, 36:64318–64330, 2023. 1, 3, 5, 6
2023
-
[72]
Neurad: Neural rendering for autonomous driving
Adam Tonderski, Carl Lindström, Georg Hess, William Ljungbergh, Lennart Svensson, and Christoffer Petersson. Neurad: Neural rendering for autonomous driving. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14895–14904, 2024. 2
2024
-
[73]
Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs
Haithem Turki, Deva Ramanan, and Mahadev Satya- narayanan. Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs. In CVPR, pages 12922– 12931, 2022. 2
2022
-
[74]
Suds: Scalable urban dynamic scenes
Haithem Turki, Jason Y Zhang, Francesco Ferroni, and Deva Ramanan. Suds: Scalable urban dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 12375–12385, 2023. 2
2023
-
[75]
Dn-splatter: Depth and normal priors for gaussian splatting and meshing.arXiv preprint arXiv:2403.17822, 2024
Matias Turkulainen, Xuqian Ren, Iaroslav Melekhov, Otto Seiskari, Esa Rahtu, and Juho Kannala. Dn-splatter: Depth and normal priors for gaussian splatting and meshing.arXiv preprint arXiv:2403.17822, 2024. 2
2024 arXiv
-
[76]
Vggt: Visual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Visual geometry grounded transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference , pages 5294–5306, 2025. 8
2025
-
[77]
Unifying appearance codes and bilateral grids for driving scene gaussian splatting, 2025
Nan Wang, Yuantao Chen, Lixing Xiao, Weiqing Xiao, Bo- han Li, Zhaoxi Chen, Chongjie Ye, Shaocong Xu, Saining Zhang, Ziyang Yan, Pierre Merriaux, Lei Lei, Tianfan Xue, and Hao Zhao. Unifying appearance codes and bilateral grids for driving scene gaussian splatting, 2025. 2
2025
-
[78]
Neus: Learning neural implicit surfaces by volume rendering for multi-view re- construction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view re- construction. arXiv preprint arXiv:2106.10689 , 2021. 5, 6
2021 arXiv
-
[79]
F2-nerf: Fast neural radiance field training with free camera trajectories
Peng Wang, Yuan Liu, Zhaoxi Chen, Lingjie Liu, Ziwei Liu, Taku Komura, Christian Theobalt, and Wenping Wang. F2-nerf: Fast neural radiance field training with free camera trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages ...
2023
-
[80]
Dust3r: Geometric 3d vision made easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 20697–20709, 2024. 8
2024
-
[81]
Openoccupancy: A large scale benchmark for surrounding semantic occupancy perception
Xiaofeng Wang, Zheng Zhu, Wenbo Xu, Yunpeng Zhang, Yi Wei, Xu Chi, Yun Ye, Dalong Du, Jiwen Lu, and Xin- gang Wang. Openoccupancy: A large scale benchmark for surrounding semantic occupancy perception. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vis...
2023
-
[82]
Uniocc: A unified benchmark for occupancy forecasting and prediction in autonomous driving
Yuping Wang, Xiangyu Huang, Xiaokang Sun, Mingx- uan Yan, Shuo Xing, Zhengzhong Tu, and Jiachen Li. Uniocc: A unified benchmark for occupancy forecasting and prediction in autonomous driving. arXiv preprint arXiv:2503.24381, 2025. 3
2025 arXiv
-
[83]
Surroundocc: Multi-camera 3d occu- pancy prediction for autonomous driving
Yi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu, Jie Zhou, and Jiwen Lu. Surroundocc: Multi-camera 3d occu- pancy prediction for autonomous driving. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 21729–21740, 2023. 3
2023
-
[84]
Mars: An instance-aware, modular and realistic simulator for au- tonomous driving
Zirui Wu, Tianyu Liu, Liyi Luo, Zhide Zhong, Jianteng Chen, Hongmin Xiao, Chao Hou, Haozhe Lou, Yuan- tao Chen, Runyi Yang, Yuxin Huang, Xiaoyu Ye, Zike Yan, Yongliang Shi, Yiyi Liao, and Hao Zhao. Mars: An instance-aware, modular and realistic simulator for au- tonomous drivi...
2023
-
[85]
Cruise: Cooperative reconstruction and editing in v2x scenarios using gaussian splatting, 2025
Haoran Xu, Saining Zhang, Peishuo Li, Baijun Ye, Xiaoxue Chen, Huan ang Gao, Jv Zheng, Xiaowei Song, Ziqiao Peng, Run Miao, Jinrang Jia, Yifeng Shi, Guangqi Yi, Hang Zhao, Hao Tang, Hongyang Li, Kaicheng Yu, and Hao Zhao. Cruise: Cooperative reconstruction and editing in v2x s...
2025
-
[86]
Grid-guided neural radiance fields for large urban scenes
Linning Xu, Yuanbo Xiangli, Sida Peng, Xingang Pan, Nanxuan Zhao, Christian Theobalt, Bo Dai, and Dahua Lin. Grid-guided neural radiance fields for large urban scenes. In CVPR, pages 8296–8306, 2023. 2
2023
-
[87]
Street gaussians for modeling dynamic ur- ban scenes
Yunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang, Haiyang Sun, Kun Zhan, Xianpeng Lang, Xiaowei Zhou, and Sida Peng. Street gaussians for modeling dynamic ur- ban scenes. arXiv preprint arXiv:2401.01339, 2024. 2, 4
2024 arXiv
-
[88]
Emernerf: Emergent spatial-temporal scene decomposition via self-supervision
Jiawei Yang, Boris Ivanovic, Or Litany, Xinshuo Weng, Se- ung Wook Kim, Boyi Li, Tong Che, Danfei Xu, Sanja Fi- dler, Marco Pavone, and Yue Wang. Emernerf: Emergent spatial-temporal scene decomposition via self-supervision. In International Conference on Learning Representations ,
-
[89]
Spectrally pruned gaussian fields with neural com- pensation, 2024
Runyi Yang, Zhenxin Zhu, Zhou Jiang, Baijun Ye, Xiaoxue Chen, Yifei Zhang, Yuantao Chen, Jian Zhao, and Hao Zhao. Spectrally pruned gaussian fields with neural com- pensation, 2024. 2
2024
-
[90]
X-scene: Large-scale driving scene gen- eration with high fidelity and flexible controllability
Yu Yang, Alan Liang, Jianbiao Mei, Yukai Ma, Yong Liu, and Gim Hee Lee. X-scene: Large-scale driving scene gen- eration with high fidelity and flexible controllability. arXiv preprint arXiv:2506.13558, 2025. 2
2025
-
[91]
Unisim: A neural closed-loop sensor simulator
Ze Yang, Yun Chen, Jingkang Wang, Sivabalan Mani- vasagam, Wei-Chiu Ma, Anqi Joyce Yang, and Raquel Ur- tasun. Unisim: A neural closed-loop sensor simulator. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 1389–1399, 2023. 2
2023
-
[92]
Blending distributed nerfs with tri-stage robust pose optimization
Baijun Ye, Caiyun Liu, Xiaoyu Ye, Yuantao Chen, Yuhai Wang, Zike Yan, Yongliang Shi, Hao Zhao, and Guyue Zhou. Blending distributed nerfs with tri-stage robust pose optimization. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 7975–
2024
-
[93]
Bevdiffuser: Plug-and-play dif- fusion model for bev denoising with ground-truth guidance
Xin Ye, Burhaneddin Yaman, Sheng Cheng, Feng Tao, Ab- hirup Mallik, and Liu Ren. Bevdiffuser: Plug-and-play dif- fusion model for bev denoising with ground-truth guidance. In Proceedings of the Computer Vision and Pattern Recog- nition Conference, pages 1495–1504, 2025. 2
2025
-
[94]
Cvt-occ: Cost volume temporal fusion for 3d occupancy prediction
Zhangchen Ye, Tao Jiang, Chenfeng Xu, Yiming Li, and Hang Zhao. Cvt-occ: Cost volume temporal fusion for 3d occupancy prediction. In European Conference on Com- puter Vision, pages 381–397. Springer, 2024. 3, 6, 7
2024
-
[95]
Gsdf: 3dgs meets sdf for improved render- ing and reconstruction
Mulin Yu, Tao Lu, Linning Xu, Lihan Jiang, Yuanbo Xian- gli, and Bo Dai. Gsdf: 3dgs meets sdf for improved render- ing and reconstruction. arXiv preprint arXiv:2403.16964,
-
[96]
Gaussian opacity fields: Efficient and compact surface reconstruction in unbounded scenes
Zehao Yu, Torsten Sattler, and Andreas Geiger. Gaussian opacity fields: Efficient and compact surface reconstruction in unbounded scenes. arXiv preprint arXiv:2404.10772 ,
-
[97]
Presight: Enhancing au- tonomous vehicle perception with city-scale nerf priors
Tianyuan Yuan, Yucheng Mao, Jiawei Yang, Yicheng Liu, Yue Wang, and Hang Zhao. Presight: Enhancing au- tonomous vehicle perception with city-scale nerf priors. In European Conference on Computer Vision, pages 323–339. Springer, 2024. 2
2024
-
[98]
Futuresight- drive: Thinking visually with spatio-temporal cot for au- tonomous driving
Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, and Xing Wei. Futuresight- drive: Thinking visually with spatio-temporal cot for au- tonomous driving. arXiv preprint arXiv:2505.17685, 2025. 2
2025 arXiv
-
[99]
Rade-gs: Ras- terizing depth in gaussian splatting
Baowen Zhang, Chuan Fang, Rakesh Shrestha, Yixun Liang, Xiaoxiao Long, and Ping Tan. Rade-gs: Ras- terizing depth in gaussian splatting. arXiv preprint arXiv:2406.01467, 2024. 2
2024 arXiv
-
[100]
Occnerf: Advancing 3d occupancy prediction in lidar-free environ- ments
Chubin Zhang, Juncheng Yan, Yi Wei, Jiaxin Li, Li Liu, Yansong Tang, Yueqi Duan, and Jiwen Lu. Occnerf: Advancing 3d occupancy prediction in lidar-free environ- ments. arXiv preprint arXiv:2312.09243, 2023. 3
2023 arXiv
-
[101]
Drone-assisted road gaussian splatting with cross- view uncertainty
Saining Zhang, Baijun Ye, Xiaoxue Chen, Yuantao Chen, Zongzheng Zhang, Cheng Peng, Yongliang Shi, and Hao Zhao. Drone-assisted road gaussian splatting with cross- view uncertainty. arXiv preprint arXiv:2408.15242, 2024. 2
2024 arXiv
-
[102]
Occformer: Dual-path transformer for vision-based 3d semantic occu- pancy prediction
Yunpeng Zhang, Zheng Zhu, and Dalong Du. Occformer: Dual-path transformer for vision-based 3d semantic occu- pancy prediction. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision , pages 9433–9443,
-
[103]
Veon: V ocabulary-enhanced occupancy prediction
Jilai Zheng, Pin Tang, Zhongdao Wang, Guoqing Wang, Xiangxuan Ren, Bailan Feng, and Chao Ma. Veon: V ocabulary-enhanced occupancy prediction. In European Conference on Computer Vision , pages 92–108. Springer,
-
[104]
Gaussianad: Gaussian- centric end-to-end autonomous driving
Wenzhao Zheng, Junjie Wu, Yao Zheng, Sicheng Zuo, Zixun Xie, Longchao Yang, Yong Pan, Zhihui Hao, Peng Jia, Xianpeng Lang, et al. Gaussianad: Gaussian- centric end-to-end autonomous driving. arXiv preprint arXiv:2412.10371, 2024. 1, 3
2024 arXiv
-
[105]
World4drive: End-to-end autonomous driving via intention-aware physical latent world model, 2025
Yupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang, Yuhang Zheng, Yinfeng Gao, Pengfei Li, Teng Zhang, Zhongpu Xia, Peng Jia, and Dongbin Zhao. World4drive: End-to-end autonomous driving via intention-aware physical latent world model, 2025. 2
2025
-
[106]
Hugsim: A real-time, photo-realistic and closed-loop simulator for autonomous driving, 2024
Hongyu Zhou, Longzhong Lin, Jiabao Wang, Yichong Lu, Dongfeng Bai, Bingbing Liu, Yue Wang, Andreas Geiger, and Yiyi Liao. Hugsim: A real-time, photo-realistic and closed-loop simulator for autonomous driving, 2024. 2
2024
-
[107]
Hugs: Holistic urban 3d scene understanding via gaussian splatting
Hongyu Zhou, Jiahao Shao, Lu Xu, Dongfeng Bai, We- ichao Qiu, Bingbing Liu, Yue Wang, Andreas Geiger, and Yiyi Liao. Hugs: Holistic urban 3d scene understanding via gaussian splatting. arXiv preprint arXiv:2403.12722, 2024. 2
2024 arXiv
-
[108]
Driving- gaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes
Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. Driving- gaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes. arXiv preprint arXiv:2312.07920, 2023. 2
2023 arXiv
-
[109]
Occgs: Zero-shot 3d oc- cupancy reconstruction with semantic and geometric-aware gaussian splatting
Xiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei, Nan Dong, and Ming-Hsuan Yang. Occgs: Zero-shot 3d oc- cupancy reconstruction with semantic and geometric-aware gaussian splatting. arXiv preprint arXiv:2502.04981, 2025. 3
2025 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.