Pith. sign in

REVIEW 13 cited by

DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.07656 v2 pith:L4IGSFN2 submitted 2025-03-07 cs.LG cs.CVcs.RO

classification cs.LGcs.CVcs.RO
keywords taskautonomousdrivetransformerdrivinge2e-adqueriessystembenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end autonomous driving (E2E-AD) has emerged as a trend in the field of autonomous driving, promising a data-driven, scalable approach to system design. However, existing E2E-AD methods usually adopt the sequential paradigm of perception-prediction-planning, which leads to cumulative errors and training instability. The manual ordering of tasks also limits the system`s ability to leverage synergies between tasks (for example, planning-aware perception and game-theoretic interactive prediction and planning). Moreover, the dense BEV representation adopted by existing methods brings computational challenges for long-range perception and long-term temporal fusion. To address these challenges, we present DriveTransformer, a simplified E2E-AD framework for the ease of scaling up, characterized by three key features: Task Parallelism (All agent, map, and planning queries direct interact with each other at each block), Sparse Representation (Task queries direct interact with raw sensor features), and Streaming Processing (Task queries are stored and passed as history information). As a result, the new framework is composed of three unified operations: task self-attention, sensor cross-attention, temporal cross-attention, which significantly reduces the complexity of system and leads to better training stability. DriveTransformer achieves state-of-the-art performance in both simulated closed-loop benchmark Bench2Drive and real world open-loop benchmark nuScenes with high FPS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decaying Turbulence and the Riemann Hypothesis: The number theory behind the infinite-time singularity

    hep-th 2026-04 unverdicted novelty 8.0 of 10

    Freely decaying incompressible turbulence possesses a universal Euler-ensemble attractor whose continuum Mellin spectrum is controlled by the non-trivial zeros of the Riemann zeta function, producing an infinite-time ...

  2. Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces

    cs.LG 2026-03 conditional novelty 7.0 of 10

    DeLL combines DPMM dual knowledge spaces with front-door causal adjustment and a non-autoregressive evolutionary decoder to reduce catastrophic forgetting and spurious correlations in lifelong end-to-end autonomous driving.

  3. WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A three-stage distillation converts an autoregressive driving VLA into a block-causal masked diffusion model, preserving planning accuracy while decoding 2.8x faster (15.1x with optimized kernels).

  4. Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    LCS steers a VLA driving model's latent features toward command-specific centroids to improve command following at roughly half the inference cost of two-pass CFG.

  5. MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Block-wise Modal Joint Attention over image, LiDAR, and diffusion action tokens yields 88.9 PDMS / 88.4 EPDMS on NAVSIM without anchors or auxiliary supervision.

  6. GeoWorldAD: Geometry World Action Model for Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Grounding an autonomous-driving action model in ego-aligned multi-scale 3D geometry and latent future-geometry tokens improves NAVSIM closed-loop PDMS/EPDMS over prior geometry- and world-model-based planners.

  7. BeyondSight: Object Permanence for End-to-End Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A sparse-query end-to-end driving model that maintains persistent actor hypotheses across full occlusions, plus a nuScenes extension that supervises and evaluates unobservable actors, cuts planning L2 error and raises...

  8. AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Trajectory-pattern anchors bridge VLA reasoning and continuous residual flow, yielding 77.28% success rate on Bench2Drive closed-loop driving.

  9. Driving Like Yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    Person2Drive is a new benchmark that generates personalized driving datasets via simulation, quantifies styles with MMD and KL metrics, and adapts E2E-AD models using a style reward framework.

  10. AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

    cs.RO 2026-01 unverdicted novelty 6.0 of 10

    Conditioning speed planning on the predicted path and relabeling synthetic cut-ins yields SOTA Bench2Drive scores (DS 89.07, SR 73.18%).

  11. CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

    cs.RO 2026-07 conditional novelty 5.5 of 10

    Residual waypoint RL around a frozen VLA prior, scaled via heterogeneous CARLA/H100 infrastructure, raises closed-loop driving score and success rate on longest6 v2 and Bench2Drive.

  12. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

  13. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

Pith tools