Pith. sign in

REVIEW 6 cited by

Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03570 v4 pith:WZGZNGUY submitted 2024-02-05 cs.LG cs.AI

classification cs.LGcs.AI
keywords diffusionfuturemodelofflinedatadynamicslearninglong-horizon
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce Diffusion World Model (DWM), a conditional diffusion model capable of predicting multistep future states and rewards concurrently. As opposed to traditional one-step dynamics models, DWM offers long-horizon predictions in a single forward pass, eliminating the need for recursive queries. We integrate DWM into model-based value estimation, where the short-term return is simulated by future trajectories sampled from DWM. In the context of offline reinforcement learning, DWM can be viewed as a conservative value regularization through generative modeling. Alternatively, it can be seen as a data source that enables offline Q-learning with synthetic data. Our experiments on the D4RL dataset confirm the robustness of DWM to long-horizon simulation. In terms of absolute performance, DWM significantly surpasses one-step dynamics models with a $44\%$ performance gain, and is comparable to or slightly surpassing their model-free counterparts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Masked diffusion language models, not larger autoregressive LLMs, are the better building block for text-based world models in agentic RL, improving rollout fidelity, diversity, and downstream task success.

  2. VRAG: Learning World Models for Interactive Video Generation

    cs.CV 2025-05 unverdicted novelty 5.0 of 10

    VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft a...

  3. Extremum Flow Matching for Offline Goal Conditioned Reinforcement Learning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Extremum Flow Matching estimates distributional support bounds from a uniform source and uses them as return conditions for offline goal-conditioned robot policies.

  4. Exploratory Diffusion Model for Unsupervised Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A diffusion-model denoising loss serves as an intrinsic reward to guide unsupervised RL exploration, plus an alternating fine-tuning scheme for diffusion policies.

  5. Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A survey reviewing how world models and agentic AI could be combined to give edge devices predictive, proactive decision-making, with a taxonomy of methods, applications, and challenges.

  6. Bounding Distributional Shifts in World Modeling through Novelty Detection

    cs.RO 2025-08 conditional novelty 4.0 of 10

    Attaching a VAE novelty detector to the DINO-WM world model and penalizing out-of-distribution predicted states in CEM planning lowers Chamfer distance on small-data robot manipulation benchmarks.

Pith tools