Pith. sign in

REVIEW 11 cited by

Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00959 v1 pith:3PWWGJR2 submitted 2024-07-01 cs.AI cs.RO

classification cs.AIcs.RO
keywords reasoningdrivinglong-tailplanningautonomouscapabilitiesend-to-endtoken
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The autonomous driving industry is increasingly adopting end-to-end learning from sensory inputs to minimize human biases in system design. Traditional end-to-end driving models, however, suffer from long-tail events due to rare or unseen inputs within their training distributions. To address this, we propose TOKEN, a novel Multi-Modal Large Language Model (MM-LLM) that tokenizes the world into object-level knowledge, enabling better utilization of LLM's reasoning capabilities to enhance autonomous vehicle planning in long-tail scenarios. TOKEN effectively alleviates data scarcity and inefficient tokenization by leveraging a traditional end-to-end driving model to produce condensed and semantically enriched representations of the scene, which are optimized for LLM planning compatibility through deliberate representation and reasoning alignment training stages. Our results demonstrate that TOKEN excels in grounding, reasoning, and planning capabilities, outperforming existing frameworks with a 27% reduction in trajectory L2 error and a 39% decrease in collision rates in long-tail scenarios. Additionally, our work highlights the importance of representation alignment and structured reasoning in sparking the common-sense reasoning capabilities of MM-LLMs for effective planning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OpenLongTail: Generative Scaling of Long-Tail Driving Data

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Pose-informed diffusion with Plücker rays, depth warps, and cross-view memory converts monocular long-tail videos into multi-view assets that improve closed-loop driving robustness nearly to ground-truth multi-view levels.

  2. DriveQA: Passing the Driving Knowledge Test

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DriveQA is a new multimodal driving-knowledge benchmark showing that LLMs and MLLMs struggle with right-of-way, numerical traffic rules, and sign variations, with modest transfer gains to nuScenes and BDD.

  3. RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding

    cs.RO 2025-07 conditional novelty 6.0 of 10

    RCG replaces handcrafted adversarial scenario scoring with a crash-grounded embedding and k-NN selection, yielding a 9.2% average relative improvement in ego success.

  4. SEAL: Vision-Language Model-Based Safe End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling

    cs.RO 2025-06 reject novelty 6.0 of 10

    SEAL extends a V2X vision-language driving model with GPT-4o-generated snow/fog data, gated scenario attention, and contrastive learning, reporting improved planning accuracy on synthetic long-tail tests.

  5. ZeroVO: Visual Odometry with Minimal Assumptions

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-frame visual odometry model using estimated depth, language priors, and semi-supervised pseudo-label filtering achieves zero-shot metric-scale pose estimation across multiple driving datasets.

  6. ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving

    cs.RO 2025-07 conditional novelty 5.0 of 10

    ReAL-AD combines VLM-generated strategy and tactical commands with a two-stage trajectory decoder, cutting open-loop L2 error and collision rate by about a third on nuScenes and Bench2Drive.

  7. World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    World4Drive couples multiple driving intentions with a latent world model to generate, score, and select trajectories, reporting state-of-the-art perception-free planning on nuScenes and NavSim.

  8. RoCA: Robust Cross-Domain End-to-End Autonomous Driving

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoCA, a Gaussian-process codebook over ego and agent tokens, improves cross-domain generalization and adaptation of end-to-end autonomous driving models without extra inference cost.

  9. CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CogAD reports state-of-the-art open-loop and closed-loop planning results by combining hierarchical scene-to-instance perception with intent-to-trajectory planning and dual-level uncertainty.

  10. Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A structured survey of LLM-based trajectory prediction methods, organized into trajectory-language mapping, multimodal fusion, and constraint-based reasoning, with benchmarks, metrics, and future directions.

  11. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

Pith tools