Pith. sign in

REVIEW 8 cited by

Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03877 v3 pith:3DMQ23XP submitted 2024-06-06 cs.RO cs.CV

classification cs.ROcs.CV
keywords drivinge2e-adunderbench2driveautonomousclosed-loopmannermethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In an era marked by the rapid scaling of foundation models, autonomous driving technologies are approaching a transformative threshold where end-to-end autonomous driving (E2E-AD) emerges due to its potential of scaling up in the data-driven manner. However, existing E2E-AD methods are mostly evaluated under the open-loop log-replay manner with L2 errors and collision rate as metrics (e.g., in nuScenes), which could not fully reflect the driving performance of algorithms as recently acknowledged in the community. For those E2E-AD methods evaluated under the closed-loop protocol, they are tested in fixed routes (e.g., Town05Long and Longest6 in CARLA) with the driving score as metrics, which is known for high variance due to the unsmoothed metric function and large randomness in the long route. Besides, these methods usually collect their own data for training, which makes algorithm-level fair comparison infeasible. To fulfill the paramount need of comprehensive, realistic, and fair testing environments for Full Self-Driving (FSD), we present Bench2Drive, the first benchmark for evaluating E2E-AD systems' multiple abilities in a closed-loop manner. Bench2Drive's official training data consists of 2 million fully annotated frames, collected from 13638 short clips uniformly distributed under 44 interactive scenarios (cut-in, overtaking, detour, etc), 23 weathers (sunny, foggy, rainy, etc), and 12 towns (urban, village, university, etc) in CARLA v2. Its evaluation protocol requires E2E-AD models to pass 44 interactive scenarios under different locations and weathers which sums up to 220 routes and thus provides a comprehensive and disentangled assessment about their driving capability under different situations. We implement state-of-the-art E2E-AD models and evaluate them in Bench2Drive, providing insights regarding current status and future directions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A simulation-to-real pipeline (HASS synthetic hard cases, scenario-aware prompts, and an image-to-ego geometry encoder) improves an MLLM's open-loop planning on nuScenes, especially in hard scenarios.

  2. Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.

  3. From Failures to Fixes: LLM-Driven Scenario Repair for Self-Evolving Autonomous Driving

    cs.CV 2025-05 reject novelty 6.0 of 10

    SERA uses LLM-driven failure analysis and scenario retrieval to select training scenarios for few-shot fine-tuning, improving simulated autonomous driving scores.

  4. iPad: Iterative Proposal-centric End-to-End Autonomous Driving

    cs.CV 2025-05 conditional novelty 6.0 of 10

    iPad achieves top NAVSIM and Bench2Drive driving scores by iteratively refining sparse candidate trajectories with proposal-anchored attention over camera images.

  5. GEMINUS: Dual-aware Global and Scene-Adaptive Mixture-of-Experts for End-to-End Autonomous Driving

    cs.CV 2025-07 conditional novelty 5.0 of 10

    GEMINUS reports state-of-the-art closed-loop driving scores on Bench2Drive with a monocular camera by routing each situation to either a global expert or a scene-specialized expert based on scenario confidence.

  6. ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving

    cs.RO 2025-07 conditional novelty 5.0 of 10

    ReAL-AD combines VLM-generated strategy and tactical commands with a two-stage trajectory decoder, cutting open-loop L2 error and collision rate by about a third on nuScenes and Bench2Drive.

  7. CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CogAD reports state-of-the-art open-loop and closed-loop planning results by combining hierarchical scene-to-instance perception with intent-to-trajectory planning and dual-level uncertainty.

  8. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

Pith tools