REVIEW 4 major objections 6 minor 1 cited by
Moving Forward: A Review of Autonomous Driving Software and Hardware Systems
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper argues that no single hardware type can efficiently run the full autonomous-driving stack, so future accelerators will be heterogeneous: CPUs and GPUs plus task-specific cores and programmable FPGAs or CGRAs.
desk verdict A competent, useful survey whose original CPU/GPU benchmark is too under-specified to carry the heterogeneity thesis; worth refereeing with requests for methodology and reframing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by workload-diversity characterization. The paper defines arithmetic intensity (FLOPs per byte, equivalently FLOPs per memory operation) for the three end-to-end benchmarks and plots layer-level intensity for InterFuser (Fig. 9), showing some layers are compute-bound and others memory-bound. This intensity spread is the mechanism that motivates heterogeneous accelerators and processing-in-memory, since memory-bound layers benefit from computation placed near or inside DRAM while compute-bound layers benefit from parallel SIMD-style engines.
What would settle it
A head-to-head study that runs the full modular stack (perception, prediction, planning) alongside the three end-to-end models on the same four CPUs and four GPUs, with a sourced real-time threshold, and finds a single device class meeting all latency and energy budgets would directly weaken the claim that heterogeneous hardware is required.
Extended reading notes
Core claim
The central claim is that the future of self-driving accelerators lies in heterogeneous architectures rather than a single dominant device class. The paper reaches this through workload diversity: the evaluated end-to-end models have markedly different parameter counts, FLOPs, and arithmetic intensities, and within a single model such as InterFuser, individual layers range over several orders of magnitude in FLOPs-per-memory-operation. It reports that on the four CPU systems almost none of the end-to-end models can sustain the required 10 FPS, while GPUs both meet and exceed it with better energy efficiency, and it argues that no single hardware type can be optimal across such a spread. The conclusion is therefore a multi-core SoC that mixes general-purpose CPUs/GPUs, task-specific accelerators including processing-in-memory cores, and programmable components such as FPGAs or CGRAs to preserve long-term adaptability.
Load-bearing premise
The argument rests on assuming the three end-to-end models and the 10 FPS threshold fairly represent what autonomous-driving hardware must run, so the Figure 8 measurements can stand in for the full software stack.
Editorial extensions
If this is right
- If the heterogeneity claim is right, next-generation automotive SoCs will need to co-design general-purpose CPU/GPU cores with task-specific accelerators and programmable fabric rather than relying on a single accelerator type.
- Memory-bound layers identified by low arithmetic intensity become the natural targets for processing-in-memory cores, while compute-bound layers can stay on GPU or neural-network accelerators.
- Because autonomous vehicles have 10- to 15-year lifespans, the winning hardware platforms will include programmable elements (FPGAs or CGRAs) so they can absorb software updates and new models.
- The CPU-only path to level-4/5 autonomy becomes untenable if the 10 FPS threshold and the measured CPU results hold, reinforcing GPUs as the baseline and specialized accelerators as the next step.
Reading between the lines
- Editorial inference: the paper's proof-of-concept measurements would be much stronger if repeated on a standard benchmark suite covering full modular stacks (perception, prediction, planning) and end-to-end policies, since the chosen three models sample only part of the software space.
- Editorial inference: the arithmetic-intensity evidence suggests the first PIM deployments in an autonomous vehicle would target LiDAR point-cloud preprocessing, sensor-fusion attention layers, and other low-intensity layers, but the paper does not commit to a specific placement.
- Editorial inference: the 10 FPS threshold is treated as given, yet it is load-bearing; an independently sourced real-time requirement could shift the CPU-vs-GPU conclusion and deserves explicit validation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of autonomous driving (AD) software and hardware systems. It reviews sensor inputs, datasets, simulators, the modular and end-to-end software architectures, and commercial hardware platforms including GPUs, FPGAs, and SoCs such as Tesla FSD, NVIDIA DRIVE Orin, and Mobileye EyeQ6. It contributes a small benchmark study (Fig. 8) measuring FPS and FPS/W of three end-to-end models—TransFuser, InterFuser, and MILE—on four CPU/GPU systems, plus a layer-level arithmetic-intensity analysis of InterFuser (Fig. 9). On this basis, Sections V.B–V.D argue that single-type homogeneous hardware is suboptimal and that future AD accelerators will be increasingly heterogeneous, combining CPUs, GPUs, task-specific accelerators including PIM, and programmable logic such as FPGAs or CGRAs.
Significance. If its central thesis is accepted, the paper provides a useful organizing perspective for AD hardware design, aligning with industry trends in heterogeneous SoCs. Its main strength is the broad synthesis of current software and hardware stacks with specific commercial examples, and the attempt to ground a hardware argument in a small set of measurements rather than speculation. The benchmark data, however, are not currently sufficient to carry the heterogeneity conclusion: the measurements lack methodological detail and error bars, the '10 FPS' threshold is unsourced, and the test models do not span the diversity of the software stack described earlier. The qualitative conclusion is defensible and consistent with the literature, but the empirical support needs substantial strengthening. The paper is likely to be useful to practitioners and researchers entering the area, but in its present form it does not meet the standard for a fully supported experimental claim.
major comments (4)
- [§V.A, Fig. 8, Table VII] The quantitative claim in Section V.A that 'none of the CPUs solely—except for one with borderline results—can meet the minimum required performance of 10 FPS' is not supported by the information provided. There is no description of the measurement methodology: batch size, input resolution, precision, inference framework and version, CPU/GPU power states, number of runs, or the method used to compute FPS/W are all absent. The '10 FPS' threshold is asserted without any citation or derivation, and the single 'borderline' CPU data point cannot be interpreted without confidence intervals or run-to-run variance. Because this result is later cited in Section V.B as the 'brief demonstration' that single-type hardware is suboptimal, the missing methodology is load-bearing for the paper's hardware argument.
- [§V.A, Table VII] The comparison labeled CPU-only versus GPU-only is confounded by the system configurations. System 1 is a Jetson AGX Orin, which is a heterogeneous SoC containing both CPU and GPU on the same die, and System 2 is a laptop-class system with both a Ryzen 9 CPU and an RTX 3060 GPU. The paper does not state how 'CPU-only' execution was isolated (e.g., whether the GPU was disabled), how the GPU measurements were taken on the same systems, or how power and thermal sharing between the CPU and GPU affected the measurements. Without this information, the direct CPU-versus-GPU comparison in Fig. 8 is not reproducible and its validity is unclear.
- [§V.B, §V.D] The logical link from Fig. 8 to the heterogeneity thesis is incomplete. The benchmark compares different homogeneous CPU systems against different homogeneous GPU systems; it does not compare a homogeneous GPU-only configuration with a heterogeneous CPU+GPU+FPGA/CGRA/PIM configuration on the same workloads. Showing that these CPUs are slower than these GPUs on three end-to-end models does not demonstrate that 'managing these diverse models with a single type of hardware leads to suboptimal performance.' The paper should either add a direct homogeneous-versus-heterogeneous comparison or substantially soften the claim, explicitly stating that Fig. 8 supports only the narrower observation that the evaluated CPUs are less performant than the evaluated GPUs on these models.
- [§V.A, Table VI] The selection of TransFuser, InterFuser, and MILE is not justified as representative of the autonomous driving software stack surveyed in Section III. These three models are all transformer-based, camera+LiDAR, end-to-end driving models; they do not cover the diverse workloads described earlier, such as 2D/3D object detection CNNs (YOLO, VoxelNet, PointPillars), point-cloud networks, tracking, trajectory-prediction GNNs, or planning algorithms. The paper's conclusion that future accelerators must handle 'diverse computational and memory requirements' relies on this representativeness, but no argument or evidence is given that the three chosen models span that diversity. Without such justification, the empirical results in Fig. 8 and the layer analysis in Fig. 9 (which is only for InterFuser) cannot be generalized to the full AD stack.
minor comments (6)
- [§II.B] The KITTI dataset is cited as reference [17], which is actually the Contraction Hierarchies routing paper (Geisberger et al.); a proper citation for the KITTI vision benchmark suite is missing or mis-numbered.
- [§III.B.1] The text states that TransFuser uses 'ResNets [84] and RegNets [85]', but reference [85] is a model-predictive motion planner for the IARA car; the RegNet backbone should be cited to Radosavovic et al., which is already reference [90].
- [Fig. 8, Table VII] The device labels CPU_1 through GPU_4 in Fig. 8 are not mapped directly to the System 1–4 rows of Table VII, making the figure hard to interpret; adding a legend or using the system names directly would improve clarity.
- [§I (page 2)] There is a typo: 'To set he stage for this' should read 'To set the stage for this.'
- [§V.A] The sentence 'Worth noting is that these results indicate even a single GPU can exhibit varying performance and efficiency across different models' is grammatically awkward and should be rephrased for clarity.
- [§V.A] The statement that projections indicate autonomous vehicles 'will dominate 95% of the market by 2050' is an imprecise reading of reference [103], which is primarily about emissions from onboard computing; the claim should be reworded to match the source's actual projection (e.g., vehicle-miles traveled share) or removed.
Circularity Check
No significant circularity: the paper's benchmark and survey reasoning are externally grounded, and no prediction reduces to its inputs by construction.
full rationale
This paper is a survey with a speculative forward-looking conclusion, not a derivation chain in which outputs are fitted from inputs. Section V.A reports empirical FPS and FPS/W measurements of three published end-to-end models (TransFuser, InterFuser, MILE) on four CPU and four GPU systems. These measurements are not used to fit any parameter that is then renamed as a prediction; the numbers are external benchmark observations, and the conclusion that GPUs outperform CPUs on these workloads follows directly from the reported measurements. The central Section V.D claim, that future self-driving accelerators will be heterogeneous, is justified by the diverse computational and memory requirements of the surveyed software stack, by the architectural diversity of existing systems (Tesla FSD, DRIVE Orin, EyeQ6), and by qualitative arguments about sparsity and arithmetic intensity. Even though the CPU-vs-GPU comparison does not logically prove that heterogeneous architectures are necessary, that is an evidentiary gap or a weakness in inductive support, not circularity. There are no load-bearing self-citations: the references are to external prior work, and no uniqueness theorem or prior result by the same authors is invoked to force a conclusion. The unsourced 10 FPS threshold and missing measurement methodology are rigor concerns, but they do not make the argument circular. The paper is therefore self-contained against external benchmarks, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Minimum 10 FPS performance threshold =
10 FPS
assumptions (3)
- ad hoc to paper The three end-to-end models TransFuser, InterFuser, and MILE are representative of the computational and memory demands of autonomous driving software stacks.
- domain assumption TDP values from hardware specifications are an adequate proxy for power consumption when computing FPS/W efficiency.
- domain assumption The reported FPS measurements are accurate despite the absence of a measurement methodology.
Cite this review
Pith. "Pith review of Moving Forward: A Review of Autonomous Driving Software and Hardware Systems." pith.science (2026). https://pith.science/paper/SAAVA7D5
@misc{pith2026241110291,
author = {Pith},
title = {Pith review of: Moving Forward: A Review of Autonomous Driving Software and Hardware Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/SAAVA7D5}},
note = {Machine review of arXiv:2411.10291}
}
read the original abstract
With their potential to significantly reduce traffic accidents, enhance road safety, optimize traffic flow, and decrease congestion, autonomous driving systems are a major focus of research and development in recent years. Beyond these immediate benefits, they offer long-term advantages in promoting sustainable transportation by reducing emissions and fuel consumption. Achieving a high level of autonomy across diverse conditions requires a comprehensive understanding of the environment. This is accomplished by processing data from sensors such as cameras, radars, and LiDARs through a software stack that relies heavily on machine learning algorithms. These ML models demand significant computational resources and involve large-scale data movement, presenting challenges for hardware to execute them efficiently and at high speed. In this survey, we first outline and highlight the key components of self-driving systems, covering input sensors, commonly used datasets, simulation platforms, and the software architecture. We then explore the underlying hardware platforms that support the execution of these software systems. By presenting a comprehensive view of autonomous driving systems and their increasing demands, particularly for higher levels of autonomy, we analyze the performance and efficiency of scaled-up off-the-shelf GPU/CPU-based systems, emphasizing the challenges within the computational components. Through examples showcasing the diverse computational and memory requirements in the software stack, we demonstrate how more specialized hardware and processing closer to memory can enable more efficient execution with lower latency. Finally, based on current trends and future demands, we conclude by speculating what a future hardware platform for autonomous driving might look like.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Open-Source Autonomous Driving Software Platforms: Comparison of Autoware and Apollo
Autoware and Apollo differ in module design, and Apollo's shared-memory middleware is faster but more memory-hungry than Autoware's serialized DDS middleware.
Reference graph
Works this paper leans on
-
[1]
Evaluation of fuel consumption and emissions benefits of connected and automated vehicles in mixed traffic flow,
H. Li, H. Li, Y . Hu, T. Xia, Q. Miao, and J. Chu, “Evaluation of fuel consumption and emissions benefits of connected and automated vehicles in mixed traffic flow,” Frontiers in Energy Research, vol. 11, p. 1207449, 2023
2023
-
[2]
Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,
S. International, “Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles,” 2021
2021
-
[3]
Autonomous vehicle decision-making and control in complex and unconventional scenar- ios—a review,
F. Sana, N. L. Azad, and K. Raahemifar, “Autonomous vehicle decision-making and control in complex and unconventional scenar- ios—a review,” Machines, vol. 11, no. 7, p. 676, 2023
2023
-
[4]
A survey of autonomous driving: Common practices and emerging technologies,
E. Yurtsever, J. Lambert, A. Carballo, and K. Takeda, “A survey of autonomous driving: Common practices and emerging technologies,” IEEE Access, vol. 8, pp. 58443–58469, 2019
2019
-
[5]
Autonomous driving in urban environments: Boss and the urban challenge,
C. Urmson, J. Anhalt, J. A. Bagnell, C. R. Baker, R. Bittner, M. N. Clark, J. M. Dolan, D. Duggins, T. Galatali, C. Geyer, M. Gittle- man, S. Harbaugh, M. Hebert, T. M. Howard, S. Kolski, A. Kelly, M. Likhachev, M. McNaughton, N. Miller, K. M. Peterson, B. Pilnick, R. R. Rajkumar, P. E. Rybski, B. Salesky, Y .-W. Seo, S. Singh, J. M. Snider, A. Stentz, W....
2008
-
[6]
C. S. Badue, R. Guidolini, R. V . Carneiro, P. Azevedo, V . B. Cardoso, A. Forechi, L. F. R. Jesus, R. Berriel, T. M. Paix ˜ao, F. W. Mutz, T. Oliveira-Santos, and A. F. de Souza, “Self-driving cars: A survey,” ArXiv, vol. abs/1901.04407, 2019
arXiv 1901
-
[7]
Introducing the 5th-generation waymo driver: Informed by experience, designed for scale, engineered to tackle more environments
“Introducing the 5th-generation waymo driver: Informed by experience, designed for scale, engineered to tackle more environments.” https: //waymo.com/blog/2020/03/introducing-5th-generation-waymo-driver. html. Online
2020
-
[8]
Apollo: Open source autonomous driving
“Apollo: Open source autonomous driving.” https://github.com/ ApolloAuto/apollo. Online
Show all 114 references
-
[9]
Autopilot
I. Tesla, “Autopilot.” https://www.tesla.com/autopilot, 2023. Online
2023
-
[10]
Environmental-driven approach towards level 5 self-driving,
M. Hurair, J. Ju, and J. Han, “Environmental-driven approach towards level 5 self-driving,” Sensors, vol. 24, no. 2, p. 485, 2024
2024
-
[11]
A novel approach for detecting road based on two-stream fusion fully convolutional network,
X. Lv, Z. yi Liu, J. Xin, and N. Zheng, “A novel approach for detecting road based on two-stream fusion fully convolutional network,” 2018 IEEE Intelligent Vehicles Symposium (IV) , pp. 1464–1469, 2018
2018
-
[12]
The pascal visual object classes (voc) challenge,
M. Everingham, L. V . Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International Journal of Computer Vision , vol. 88, pp. 303–338, 2010
2010
-
[13]
Imagenet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” 2009 IEEE Conference on Computer Vision and Pattern Recognition , pp. 248–255, 2009
2009
-
[14]
Microsoft coco: Common objects in context,
T.-Y . Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European Conference on Computer Vision , 2014
2014
-
[15]
Learning multiple layers of features from tiny images,
A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009
2009
-
[16]
The cityscapes dataset for semantic urban scene understanding,
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 3213–3223, 2016
2016
-
[17]
Exact routing in large road networks using contraction hierarchies,
R. Geisberger, P. Sanders, D. Schultes, and C. Vetter, “Exact routing in large road networks using contraction hierarchies,”Transp. Sci., vol. 46, pp. 388–404, 2012
2012
-
[18]
V oxelnet: End-to-end learning for point cloud based 3d object detection,
Y . Zhou and O. Tuzel, “V oxelnet: End-to-end learning for point cloud based 3d object detection,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 4490–4499, 2017
2018
-
[19]
Pointpillars: Fast encoders for object detection from point clouds,
A. H. Lang, S. V ora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), pp. 12689–12697, 2018
2019
-
[20]
Multi-view 3d object detection network for autonomous driving,
X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object detection network for autonomous driving,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 6526–6534, 2016
2017
-
[21]
Joint 3d proposal generation and object detection from view aggregation,
J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. L. Waslander, “Joint 3d proposal generation and object detection from view aggregation,” 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 1–8, 2017
2018
-
[22]
Scalability in percep- tion for autonomous driving: Waymo open dataset,
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine, V . Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. M. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y . Zhang, J. Shlens, Z. Chen, and D. Anguelov, “Scalability in p...
2020
-
[23]
Large scale interactive motion forecasting for autonomous driving : The waymo open motion dataset,
S. M. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y . Chai, B. Sapp, C. Qi, Y . Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V . Vasudevan, A. McCauley, J. Shlens, and D. Anguelov, “Large scale interactive motion forecasting for autonomous driving : The waymo op...
2021
-
[24]
Womd-lidar: Raw sensor dataset benchmark for motion forecasting,
K. Chen, R. Ge, H. Qiu, R. Ai-Rfou, C. Qi, X. Zhou, Z. Yang, S. M. Ettinger, P. Sun, Z. Leng, M. A. Mustafa, I. Bogun, W. Wang, M. Tan, and D. Anguelov, “Womd-lidar: Raw sensor dataset benchmark for motion forecasting,” ArXiv, vol. abs/2304.03834, 2023
2023 arXiv
-
[25]
Autoware
A. Foundation, “Autoware.” https://github.com/autowarefoundation/ autoware, 2023. Online
2023
-
[26]
The apolloscape open dataset for autonomous driving and its application,
P. Wang, X. Huang, X. Cheng, D. Zhou, Q. Geng, and R. Yang, “The apolloscape open dataset for autonomous driving and its application,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 42, pp. 2702–2719, 2018
2018
-
[27]
Lgsvl simulator: A high fidelity simulator for autonomous driving,
G. Rong, B. H. Shin, H. Tabatabaee, Q. Lu, S. Lemke, M. Mo ˇzeiko, E. Boise, G. Uhm, M. Gerow, S. Mehta, et al. , “Lgsvl simulator: A high fidelity simulator for autonomous driving,” arXiv preprint arXiv:2005.03778, 2020
2005 arXiv
-
[28]
CARLA: An open urban driving simulator,
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun, “CARLA: An open urban driving simulator,” in Proceedings of the 1st Annual Conference on Robot Learning , pp. 1–16, 2017
2017
-
[29]
Faster r-cnn: Towards real- time object detection with region proposal networks,
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster r-cnn: Towards real- time object detection with region proposal networks,” IEEE Transac- tions on Pattern Analysis and Machine Intelligence , vol. 39, pp. 1137– 1149, 2015
2015
-
[30]
You only look once: Unified, real-time object detection,
J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 779–788, 2015
2016
-
[31]
Ssd: Single shot multibox detector,
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C.-Y . Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European Conference on Computer Vision , 2015
2015
-
[32]
Second: Sparsely embedded convolutional detection,
Y . Yan, Y . Mao, and B. Li, “Second: Sparsely embedded convolutional detection,” Sensors (Basel, Switzerland) , vol. 18, 2018
2018
-
[33]
Pointnet: Deep learning on point sets for 3d classification and segmentation,
C. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 77–85, 2016
2017
-
[34]
Pointnet++: Deep hierarchical feature learning on point sets in a metric space,
C. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in Neural Information Processing Systems, 2017
2017
-
[35]
Frustum pointnets for 3d object detection from rgb-d data,
C. Qi, W. Liu, C. Wu, H. Su, and L. J. Guibas, “Frustum pointnets for 3d object detection from rgb-d data,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 918–927, 2017
2018
-
[36]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” CoRR, vol. abs/1409.1556, 2014
2014 arXiv
-
[37]
Attention is all you need,
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems , 2017
2017
-
[38]
Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y . Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1080–1089, 2022
2022
-
[39]
Unifying voxel- based representation with transformer for 3d object detection,
Y . Li, Y . Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel- based representation with transformer for 3d object detection,” ArXiv, vol. abs/2206.00630, 2022
2022 arXiv
-
[40]
3d object detection for au- tonomous driving: A comprehensive survey,
J. Mao, S. Shi, X. Wang, and H. Li, “3d object detection for au- tonomous driving: A comprehensive survey,” International Journal of Computer Vision, vol. 131, pp. 1909 – 1963, 2022
1909
-
[41]
3d object tracking using rgb and lidar data,
A. Asvadi, P. Gir ˜ao, P. Peixoto, and U. J. C. Nunes, “3d object tracking using rgb and lidar data,” 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC) , pp. 1255–1260, 2016
2016
-
[42]
Fast multiple objects detection and tracking fusing color camera and 3d lidar for intelligent vehicles,
S. Hwang, N. Kim, Y . Choi, S. Lee, and I.-S. Kweon, “Fast multiple objects detection and tracking fusing color camera and 3d lidar for intelligent vehicles,” 2016 13th International Conference on Ubiquitous Robots and Ambient Intelligence (URAI) , pp. 234–239, 2016
2016
-
[43]
Fast r-cnn,
R. B. Girshick, “Fast r-cnn,” 2015
2015
-
[44]
Spin-images: A representation for 3-d surface match- ing,
A. E. Johnson, “Spin-images: A representation for 3-d surface match- ing,” 1997
1997
-
[45]
Using spin images for efficient object recognition in cluttered 3d scenes,
A. E. Johnson and M. Hebert, “Using spin images for efficient object recognition in cluttered 3d scenes,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 21, pp. 433–449, 1999
1999
-
[46]
Yolov3: An incremental improvement,
J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” ArXiv, vol. abs/1804.02767, 2018
2018 arXiv
-
[47]
Yolo9000: Better, faster, stronger,
J. Redmon and A. Farhadi, “Yolo9000: Better, faster, stronger,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6517–6525, 2016
2017
-
[48]
Yolov4: Optimal speed and accuracy of object detection,
A. Bochkovskiy, C.-Y . Wang, and H.-Y . M. Liao, “Yolov4: Optimal speed and accuracy of object detection,” ArXiv, vol. abs/2004.10934, 2020
2004 arXiv
-
[49]
Traffic light recognition using convolutional neural networks: A survey,
S. Pavlitska, N. Lambing, A. K. Bangaru, and J. M. Z ¨ollner, “Traffic light recognition using convolutional neural networks: A survey,” 2023 IEEE 26th International Conference on Intelligent Transportation Systems (ITSC), pp. 2790–2796, 2023
2023
-
[50]
Traffic light detection using tensorflow object detection framework,
T. V . Janahiraman and M. S. M. Subuhan, “Traffic light detection using tensorflow object detection framework,” 2019 IEEE 9th International Conference on System Engineering and Technology (ICSET) , pp. 108– 113, 2019
2019
-
[51]
Traffic lights detection and recognition method based on the improved yolov4 algorithm,
Q. Wang, Q. Zhang, X. Liang, Y . Wang, C. Zhou, and V . I. Mikulovich, “Traffic lights detection and recognition method based on the improved yolov4 algorithm,” Sensors (Basel, Switzerland) , vol. 22, 2021
2021
-
[52]
A two-stage framework for diverse traffic light recognition based on individual signal detection,
S.-Y . Lin and H.-Y . Lin, “A two-stage framework for diverse traffic light recognition based on individual signal detection,” in Mediterranean Conference on Pattern Recognition and Artificial Intelligence , 2021
2021
-
[53]
A sensor fusion approach for localization with cumulative error elimination,
F. Zhang, H. Stahle, G. Chen, C.-W. Chen, C. Simon, C. Buckl, and A. Knoll, “A sensor fusion approach for localization with cumulative error elimination,” 2012 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI) , pp. 1–6, 2012
2012
-
[54]
Road marking detection using lidar reflective intensity data and its application to vehicle localization,
A. Y . Hata and D. F. Wolf, “Road marking detection using lidar reflective intensity data and its application to vehicle localization,” 17th International IEEE Conference on Intelligent Transportation Systems (ITSC), pp. 584–589, 2014
2014
-
[55]
Sensor fusion-based low-cost vehicle localization system for complex urban environments,
J. K. Suhr, J. Jang, D. Min, and H. G. Jung, “Sensor fusion-based low-cost vehicle localization system for complex urban environments,” IEEE Transactions on Intelligent Transportation Systems , vol. 18, pp. 1078–1086, 2017
2017
-
[56]
A survey of the state-of-the-art localization techniques and their potentials for autonomous vehicle applications,
S. Kuutti, S. Fallah, K. V . Katsaros, M. Dianati, F. Mccullough, and A. Mouzakitis, “A survey of the state-of-the-art localization techniques and their potentials for autonomous vehicle applications,” IEEE Inter- net of Things Journal , vol. 5, pp. 829–846, 2018
2018
-
[57]
Multimodal trajectory pre- dictions for autonomous driving using deep convolutional networks,
H. Cui, V . Radosavljevic, F.-C. Chou, T.-H. Lin, T. Nguyen, T.-K. Huang, J. G. Schneider, and N. Djuric, “Multimodal trajectory pre- dictions for autonomous driving using deep convolutional networks,” 2019 International Conference on Robotics and Automation (ICRA) , pp. 2090–...
2019
-
[58]
Vectornet: Encoding hd maps and agent dynamics from vectorized representation,
J. Gao, C. Sun, H. Zhao, Y . Shen, D. Anguelov, C. Li, and C. Schmid, “Vectornet: Encoding hd maps and agent dynamics from vectorized representation,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11522–11530, 2020
2020
-
[59]
Bert: Pre- training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre- training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics, 2019
2019
-
[60]
Wayformer: Motion forecasting via simple & efficient attention networks,
N. Nayakanti, R. Al-Rfou, A. Zhou, K. Goel, K. S. Refaat, and B. Sapp, “Wayformer: Motion forecasting via simple & efficient attention networks,” 2023 IEEE International Conference on Robotics and Automation (ICRA) , pp. 2980–2987, 2022
2023
-
[61]
Risky action recognition in lane change video clips using deep spatiotemporal networks with segmentation mask transfer,
E. Yurtsever, Y . Liu, J. Lambert, C. Miyajima, E. Takeuchi, K. Takeda, and J. L. Hansen, “Risky action recognition in lane change video clips using deep spatiotemporal networks with segmentation mask transfer,” 2019 IEEE Intelligent Transportation Systems Conference (ITSC), p...
2019
-
[62]
Concrete problems for autonomous vehicle safety: Advantages of bayesian deep learning,
R. T. McAllister, Y . Gal, A. Kendall, M. van der Wilk, A. Shah, R. Cipolla, and A. Weller, “Concrete problems for autonomous vehicle safety: Advantages of bayesian deep learning,” in International Joint Conference on Artificial Intelligence , 2017
2017
-
[63]
Highway hierarchies hasten exact shortest path queries,
P. Sanders and D. Schultes, “Highway hierarchies hasten exact shortest path queries,” in Embedded Systems and Applications , 2005
2005
-
[64]
Reach for a*: Efficient point-to-point shortest path algorithms,
A. V . Goldberg, H. Kaplan, and R. F. Werneck, “Reach for a*: Efficient point-to-point shortest path algorithms,” in Workshop on Algorithm Engineering and Experimentation , 2006
2006
-
[65]
Efficient constrained path planning via search in state lattices,
M. Pivtoraiko and A. Kelly, “Efficient constrained path planning via search in state lattices,” 2005
2005
-
[66]
Time-bounded lattice for efficient planning in dynamic environments,
A. Kushleyev and M. Likhachev, “Time-bounded lattice for efficient planning in dynamic environments,” 2009 IEEE International Confer- ence on Robotics and Automation , pp. 1662–1668, 2009
2009
-
[67]
On the application of the d* search algorithm to time-based planning on lattice graphs,
M. Rufli and R. Y . Siegwart, “On the application of the d* search algorithm to time-based planning on lattice graphs,” in European Conference on Mobile Robots , 2009
2009
-
[68]
Center-based 3d object detection and tracking,
T. Yin, X. Zhou, and P. Kr ¨ahenb¨uhl, “Center-based 3d object detection and tracking,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11779–11788, 2020
2021
-
[69]
Multi-camera fusion in apollo software distribution,
G. Kulathunga, A. Buyval, and A. S. Klimchik, “Multi-camera fusion in apollo software distribution,” IFAC-PapersOnLine, 2019
2019
-
[70]
Spatial as deep: Spatial cnn for traffic scene understanding,
X. Pan, J. Shi, P. Luo, X. Wang, and X. Tang, “Spatial as deep: Spatial cnn for traffic scene understanding,” in AAAI Conference on Artificial Intelligence, 2017
2017
-
[71]
Mobilenetv2: Inverted residuals and linear bottlenecks,
M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 4510–4520, 2018
2018
-
[72]
Data driven prediction archi- tecture for autonomous driving and its application on apollo platform,
K. Xu, X. Xiao, J. Miao, and Q. Luo, “Data driven prediction archi- tecture for autonomous driving and its application on apollo platform,” 2020 IEEE Intelligent Vehicles Symposium (IV) , pp. 175–181, 2020
2020
-
[73]
End-to-end autonomous driving: Challenges and frontiers,
L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li, “End-to-end autonomous driving: Challenges and frontiers,” ArXiv, vol. abs/2306.16927, 2023
2023 arXiv
-
[74]
Learning to drive in a day,
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J. M. Allen, V .-D. Lam, A. Bewley, and A. Shah, “Learning to drive in a day,” 2019 International Conference on Robotics and Automation (ICRA) , pp. 8248–8254, 2018
2019
-
[75]
Cirl: Controllable imitative reinforcement learning for vision-based self-driving,
X. Liang, T. Wang, L. Yang, and E. P. Xing, “Cirl: Controllable imitative reinforcement learning for vision-based self-driving,” ArXiv, vol. abs/1807.03776, 2018
2018 arXiv
-
[76]
Trans- fuser: Imitation with transformer-based sensor fusion for autonomous driving,
K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger, “Trans- fuser: Imitation with transformer-based sensor fusion for autonomous driving,” IEEE Transactions on Pattern Analysis and Machine Intelli- gence, vol. 45, pp. 12878–12895, 2022
2022
-
[77]
Safety-enhanced autonomous driving using interpretable sensor fusion transformer,
H. Shao, L. Wang, R. Chen, H. Li, and Y . T. Liu, “Safety-enhanced autonomous driving using interpretable sensor fusion transformer,” in Conference on Robot Learning , 2022
2022
-
[78]
End to end learning for self-driving cars,
M. Bojarski, D. W. del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Zieba, “End to end learning for self-driving cars,” ArXiv, vol. abs/1604.07316, 2016
2016 arXiv
-
[79]
Urban driving with conditional imitation learning,
J. Hawke, R. Shen, C. Gurau, S. Sharma, D. Reda, N. Nikolov, P. Mazur, S. Micklethwaite, N. Griffiths, A. Shah, and A. Kendall, “Urban driving with conditional imitation learning,” 2020 IEEE Inter- national Conference on Robotics and Automation (ICRA), pp. 251–257, 2019
2020
-
[80]
Exploring the limitations of behavior cloning for autonomous driving,
F. Codevilla, E. Santana, A. M. L ´opez, and A. Gaidon, “Exploring the limitations of behavior cloning for autonomous driving,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 9328–9337, 2019
2019
-
[81]
Alvinn, an autonomous land vehicle in a neural network,
D. A. Pomerleau, “Alvinn, an autonomous land vehicle in a neural network,” 2015
2015
-
[82]
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,
D. Sun, X. Yang, M.-Y . Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 8934– 8943, 2017
2018
-
[83]
Self- attention generative adversarial networks,
H. Zhang, I. J. Goodfellow, D. N. Metaxas, and A. Odena, “Self- attention generative adversarial networks,” ArXiv, vol. abs/1805.08318, 2018
2018 arXiv
-
[84]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2015
2016
-
[85]
A model-predictive motion planner for the iara autonomous car,
V . B. Cardoso, J. Oliveira, T. Teixeira, C. S. Badue, F. W. Mutz, T. Oliveira-Santos, L. de Paula Veronese, and A. F. de Souza, “A model-predictive motion planner for the iara autonomous car,” 2017 IEEE International Conference on Robotics and Automation (ICRA) , pp. 225–230, 2016
2017
-
[86]
Mastering the game of go with deep neural networks and tree search,
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V . Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis,...
2016
-
[87]
Continuous control with deep reinforcement learning,
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. M. O. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” CoRR, vol. abs/1509.02971, 2015
2015 arXiv
-
[88]
Driving with llms: Fusing object- level vector modality for explainable autonomous driving,
L. Chen, O. Sinavski, J. H ¨unermann, A. Karnsund, A. J. Willmott, D. Birch, D. Maund, and J. Shotton, “Driving with llms: Fusing object- level vector modality for explainable autonomous driving,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) , pp. 14...
2024
-
[89]
Tesla ai day 2021
“Tesla ai day 2021.” https://www.youtube.com/watch?v= j0z4FweCy4M. Online
2021
-
[90]
Designing network design spaces,
I. Radosavovic, R. P. Kosaraju, R. B. Girshick, K. He, and P. Doll ´ar, “Designing network design spaces,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 10425–10433, 2020
2020
-
[91]
Efficientdet: Scalable and efficient object detection,
M. Tan, R. Pang, and Q. V . Le, “Efficientdet: Scalable and efficient object detection,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10778–10787, 2019
2020
-
[92]
Accelerating the pony.ai av sensor data pro- cessing pipeline
“Accelerating the pony.ai av sensor data pro- cessing pipeline.” https://developer.nvidia.com/blog/ accelerating-the-pony-av-sensor-data-processing-pipeline/. Accessed: 2024-05-31
2024
-
[93]
Compute solution for tesla’s full self-driving computer,
E. Talpes, A. Gorti, G. S. Sachdev, D. D. Sarma, G. Venkataramanan, P. J. Bannon, B. McGee, B. Floering, A. Jalote, C. Hsiong, and S. Arora, “Compute solution for tesla’s full self-driving computer,” IEEE Micro, vol. 40, pp. 25–35, 2020
2020
-
[94]
Technology - pony.ai
“Technology - pony.ai.” https://pony.ai/tech?lang=en. Accessed: 2024- 05-31
2024
-
[95]
The architectural implications of autonomous driv- ing: Constraints and acceleration,
S.-C. Lin, Y . Zhang, C.-H. Hsu, M. Skach, M. E. Haque, L. Tang, and J. Mars, “The architectural implications of autonomous driv- ing: Constraints and acceleration,” Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and...
2018
-
[96]
NUVO-6108GC Industrial-grade GPU Computing Platform
“NUVO-6108GC Industrial-grade GPU Computing Platform.” Online
-
[97]
Jetson Orin Family — NVIDIA
NVIDIA Corporation, “Jetson Orin Family — NVIDIA.” https: //www.nvidia.com/en-us/autonomous-machines/embedded-systems/ jetson-orin/, 2024. Accessed: 2024-05-21
2024
-
[98]
Nvidia drive in-vehicle computing for autonomous vehicles
“Nvidia drive in-vehicle computing for autonomous vehicles.” https: //www.nvidia.com/en-us/self-driving-cars/in-vehicle-computing/. On- line
-
[99]
Eyeq chip technology
“Eyeq chip technology.” https://www.mobileye.com/technology/ eyeq-chip/. Online
-
[100]
Hardware acceleration of deep neural networks for autonomous driving on fpga- based soc,
G. Sciangula, F. Restuccia, A. Biondi, and G. C. Buttazzo, “Hardware acceleration of deep neural networks for autonomous driving on fpga- based soc,” 2022 25th Euromicro Conference on Digital System Design (DSD), pp. 406–414, 2022
2022
-
[101]
Tesla vision update: Replacing ultrasonic sensors with tesla vision
“Tesla vision update: Replacing ultrasonic sensors with tesla vision.” https://www.tesla.com/en eu/support/transitioning-tesla-vision. On- line
-
[102]
Model-based imitation learning for urban driving,
A. Hu, G. Corrado, N. Griffiths, Z. Murez, C. Gurau, H. Yeo, A. Kendall, R. Cipolla, and J. Shotton, “Model-based imitation learning for urban driving,” ArXiv, vol. abs/2210.07729, 2022
2022 arXiv
-
[103]
Data centers on wheels: Emis- sions from computing onboard autonomous vehicles,
S. Sudhakar, V . Sze, and S. Karaman, “Data centers on wheels: Emis- sions from computing onboard autonomous vehicles,” IEEE Micro , vol. 43, no. 1, pp. 29–39, 2023
2023
-
[104]
Lazypim: An efficient cache coherence mechanism for processing-in-memory,
A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, K. Hsieh, K. T. Malladi, H. Zheng, and O. Mutlu, “Lazypim: An efficient cache coherence mechanism for processing-in-memory,” IEEE Computer Architecture Letters, vol. 16, pp. 46–50, 2017
2017
-
[105]
Cellular logic-in-memory arrays,
W. H. Kautz, “Cellular logic-in-memory arrays,” IEEE Transactions on Computers, vol. C-18, pp. 719–727, 1969
1969
-
[106]
A logic-in-memory computer,
H. S. Stone, “A logic-in-memory computer,” IEEE Transactions on Computers, vol. C-19, pp. 73–78, 1970
1970
-
[107]
New potential breakthrough memory: Hbm2 aquabolt
“New potential breakthrough memory: Hbm2 aquabolt.” https:// semiconductor.samsung.com/dram/hbm/hbm2-aquabolt/. Online
-
[108]
Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory,
D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A programmable digital neuromorphic architecture with high-density 3d memory,” in 2016 ACM/IEEE 43rd Annual Interna- tional Symposium on Computer Architecture (ISCA) , pp. 380–392, 2016
2016
-
[109]
Decomposing a scene into geomet- ric and semantically consistent regions,
S. Gould, R. Fulton, and D. Koller, “Decomposing a scene into geomet- ric and semantically consistent regions,” 2009 IEEE 12th International Conference on Computer Vision , pp. 1–8, 2009
2009
-
[110]
Tetris: Scalable and efficient neural network acceleration with 3d memory,
M. Gao, J. Pu, X. S. Yang, M. Horowitz, and C. E. Kozyrakis, “Tetris: Scalable and efficient neural network acceleration with 3d memory,” Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, 2017
2017
-
[111]
Eyeriss: An energy- efficient reconfigurable accelerator for deep convolutional neural net- works,
Y . hsin Chen, T. Krishna, J. S. Emer, and V . Sze, “Eyeriss: An energy- efficient reconfigurable accelerator for deep convolutional neural net- works,” IEEE Journal of Solid-State Circuits , vol. 52, pp. 127–138, 2016
2016
-
[112]
Ambit: In- memory accelerator for bulk bitwise operations using commodity dram technology,
V . Seshadri, D. Lee, T. Mullins, H. Hassan, A. Boroumand, J. S. Kim, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Ambit: In- memory accelerator for bulk bitwise operations using commodity dram technology,”2017 50th Annual IEEE/ACM International Symposium on Microa...
2017
-
[113]
Simdram: An end-to-end framework for bit-serial simd computing in dram,
N. Hajinazar, G. F. Oliveira, S. Gregorio, J. D. Ferreira, N. M. Ghiasi, M. Patel, M. Alser, S. Ghose, J. G. Luna, and O. Mutlu, “Simdram: An end-to-end framework for bit-serial simd computing in dram,” ArXiv, vol. abs/2105.12839, 2021
2021 arXiv
-
[114]
Dracc: a dram based accelerator for accurate cnn inference,
Q. Deng, L. Jiang, Y . Zhang, M. Zhang, and J. Yang, “Dracc: a dram based accelerator for accurate cnn inference,” 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC) , pp. 1–6, 2018
2018
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.