Pith. sign in

REVIEW 5 cited by

DiVE: DiT-based Video Generation with Enhanced Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.01595 v1 pith:7SNGFKVH submitted 2024-09-03 cs.CV

classification cs.CV
keywords videosconsistentcontrolgenerationproposedcasescornerdit-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent video generation works are proposed to tackcle the mentioned problem, i.e. models built on top of Diffusion Transformers (DiT), works are still missing which are targeted on exploring the potential for multi-view videos generation scenarios. Noticeably, we propose the first DiT-based framework specifically designed for generating temporally and multi-view consistent videos which precisely match the given bird's-eye view layouts control. Specifically, the proposed framework leverages a parameter-free spatial view-inflated attention mechanism to guarantee the cross-view consistency, where joint cross-attention modules and ControlNet-Transformer are integrated to further improve the precision of control. To demonstrate our advantages, we extensively investigate the qualitative comparisons on nuScenes dataset, particularly in some most challenging corner cases. In summary, the effectiveness of our proposed method in producing long, controllable, and highly consistent videos under difficult conditions is proven to be effective.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A five-aspect, 24-metric benchmark, a 26K human-annotated dataset, and an AI evaluator show that today's driving world models cannot simultaneously look real, respect geometry, and behave safely.

  2. Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A joint video and LiDAR generation framework for driving scenes, conditioned on shared scene layouts, VLM captions, and BEV features, achieves SOTA generation and downstream perception gains on nuScenes.

  3. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  4. HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in th...

  5. LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

    cs.CV 2025-06 reject novelty 5.0 of 10

    A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...

Pith tools