Pith. sign in

REVIEW 3 cited by

BEVGPT: Generative Pre-trained Large Model for Autonomous Driving Prediction, Decision-Making, and Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10357 v1 pith:NEZUCIMY submitted 2023-10-16 cs.RO

classification cs.RO
keywords drivingplanningautonomousdecision-makingframeworkmotionpredictionmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prediction, decision-making, and motion planning are essential for autonomous driving. In most contemporary works, they are considered as individual modules or combined into a multi-task learning paradigm with a shared backbone but separate task heads. However, we argue that they should be integrated into a comprehensive framework. Although several recent approaches follow this scheme, they suffer from complicated input representations and redundant framework designs. More importantly, they can not make long-term predictions about future driving scenarios. To address these issues, we rethink the necessity of each module in an autonomous driving task and incorporate only the required modules into a minimalist autonomous driving framework. We propose BEVGPT, a generative pre-trained large model that integrates driving scenario prediction, decision-making, and motion planning. The model takes the bird's-eye-view (BEV) images as the only input source and makes driving decisions based on surrounding traffic scenarios. To ensure driving trajectory feasibility and smoothness, we develop an optimization-based motion planning method. We instantiate BEVGPT on Lyft Level 5 Dataset and use Woven Planet L5Kit for realistic driving simulation. The effectiveness and robustness of the proposed framework are verified by the fact that it outperforms previous methods in 100% decision-making metrics and 66% motion planning metrics. Furthermore, the ability of our framework to accurately generate BEV images over the long term is demonstrated through the task of driving scenario prediction. To the best of our knowledge, this is the first generative pre-trained large model for autonomous driving prediction, decision-making, and motion planning with only BEV images as input.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

    cs.CV 2024-12 conditional novelty 6.0 of 10

    VLM-AD uses GPT-4o-generated reasoning and action annotations as auxiliary supervision to improve end-to-end autonomous driving planning without VLM inference.

  2. RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

    cs.CV 2024-12 conditional novelty 5.0 of 10

    One 8B multimodal model trained jointly on six driving datasets outperforms individual specialists on average and transfers zero-shot to three unseen driving benchmarks.

  3. PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving

    cs.RO 2024-12 conditional novelty 4.0 of 10

    PKRD-CoT structures multimodal LLM prompts into perception, knowledge, reasoning, and decision steps, and the authors report improved driving decision accuracy for GPT-4.0 and several other models.

Pith tools