Pith. sign in

REVIEW 10 cited by

AdaWorld: Learning Adaptable World Models with Latent Actions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.18938 v4 pith:37P2GXVK submitted 2025-03-24 cs.AI cs.CVcs.LGcs.RO

classification cs.AIcs.CVcs.LGcs.RO
keywords worldactionsmodelslearningadaworldlatentacrossadaptable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

World models aim to learn action-controlled future prediction and have proven essential for the development of intelligent agents. However, most existing world models rely heavily on substantial action-labeled data and costly training, making it challenging to adapt to novel environments with heterogeneous actions through limited interactions. This limitation can hinder their applicability across broader domains. To overcome this limitation, we propose AdaWorld, an innovative world model learning approach that enables efficient adaptation. The key idea is to incorporate action information during the pretraining of world models. This is achieved by extracting latent actions from videos in a self-supervised manner, capturing the most critical transitions between frames. We then develop an autoregressive world model that conditions on these latent actions. This learning paradigm enables highly adaptable world models, facilitating efficient transfer and learning of new actions even with limited interactions and finetuning. Our comprehensive experiments across multiple environments demonstrate that AdaWorld achieves superior performance in both simulation quality and visual planning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

    cs.RO 2026-07 conditional novelty 7.0 of 10

    Enfold folds the internal computation of a video world generator into a current-only representation, enabling competitive robot control without executing the generator at deployment.

  2. PhiZero: A World Model Built Around Physical Language

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A self-supervised discrete physical-language bottleneck plus a VLM reasoner lets a world model predict state transitions before rendering video, improving physical coherence and enabling zero-shot motion transfer.

  3. Robot-Factored World Models via Robot Rendering

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Conditioning a video world model on rendered nominal robot trajectories (URDF mesh + depth) instead of raw actions or logged future states improves action-following and enables zero-shot embodiment change.

  4. Factored Latent Action World Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    FLAM splits a scene into separate factors, each with its own latent action, and reports better video prediction and downstream policy learning than monolithic latent-action models.

  5. Ego-centric Predictive Model Conditioned on Hand Trajectories

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Ego-PM predicts future hand trajectories and then uses them to condition latent diffusion video generation, jointly outputting actions and future frames in egocentric and robotic scenes.

  6. DLAM: Distributional Latent Actions with Temporal Constraints

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Diagonal-Gaussian latent transitions with normalized composition and reversal constraints improve reconstruction and π0 policy transfer over deterministic structured latent-action models.

  7. 3D and 4D World Modeling: A Survey

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A survey that defines 3D/4D world modeling, organizes methods into VideoGen, OccGen, and LiDARGen categories, and compiles datasets, metrics, and benchmark numbers.

  8. Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics

    cs.LG 2026-08 conditional novelty 4.0 of 10

    A quantum-inspired world model with complex-valued latents beats matched classical baselines on one-step cellular-automaton prediction, but its advantage decays in long-horizon rollout.

  9. Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

    cs.CV 2026-07 reject novelty 4.0 of 10

    A CP-decomposed, instruction-conditioned latent bottleneck (CompactNav) improves VLN-CE success rate by about 2% over prior state of the art on two benchmarks.

  10. From World Models to World Action Models: A Concise Tutorial for Robotics

    cs.RO 2026-07 unverdicted novelty 4.0 of 10

    World models are action-conditioned predictors of task-relevant futures; world action models couple those futures to robot actions via four paradigms: imagine-then-execute, feature-conditioned, joint, and auxiliary pr...

Pith tools