Pith. sign in

REVIEW 5 cited by

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14729 v3 pith:4GRUYRTK submitted 2025-01-24 cs.CV

classification cs.CV
keywords scenedrivinggenerationhermesunderstandingworldmodelunified
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting and reasoning about the driving environment. In this paper, we present a unified Driving World Model named HERMES. We seamlessly integrate 3D scene understanding and future scene evolution (generation) through a unified framework in driving scenarios. Specifically, HERMES leverages a Bird's-Eye View (BEV) representation to consolidate multi-view spatial information while preserving geometric relationships and interactions. We also introduce world queries, which incorporate world knowledge into BEV features via causal attention in the Large Language Model, enabling contextual enrichment for understanding and generation tasks. We conduct comprehensive studies on nuScenes and OmniDrive-nuScenes datasets to validate the effectiveness of our method. HERMES achieves state-of-the-art performance, reducing generation error by 32.4% and improving understanding metrics such as CIDEr by 8.0%. The model and code will be publicly released at https://github.com/LMD0311/HERMES.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A new benchmark, STSnu, uses 971 verified multiple-choice questions from NuScenes to test driving vision-language models' spatio-temporal reasoning, and shows they lag far behind text-only LLMs given perfect trajectories.

  2. CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Forecasting future LiDAR from historical camera–radar inputs pretrains a transferable CR BEV backbone that improves long-horizon geometry prediction and many nuScenes driving tasks.

  3. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  4. Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A joint video and LiDAR generation framework for driving scenes, conditioned on shared scene layouts, VLM captions, and BEV features, achieves SOTA generation and downstream perception gains on nuScenes.

  5. DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    DriveX predicts future latent BEV features from driving video and shows consistent, modest gains on occupancy, flow, and end-to-end driving, though no code is released.

Pith tools