Pith. sign in

REVIEW 11 cited by

Learning and Leveraging World Models in Visual Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.00504 v1 pith:LDLQHR4B submitted 2024-03-01 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords learningworldimagemodelrepresentationsapproachjepalearned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Joint-Embedding Predictive Architecture (JEPA) has emerged as a promising self-supervised approach that learns by leveraging a world model. While previously limited to predicting missing parts of an input, we explore how to generalize the JEPA prediction task to a broader set of corruptions. We introduce Image World Models, an approach that goes beyond masked image modeling and learns to predict the effect of global photometric transformations in latent space. We study the recipe of learning performant IWMs and show that it relies on three key aspects: conditioning, prediction difficulty, and capacity. Additionally, we show that the predictive world model learned by IWM can be adapted through finetuning to solve diverse tasks; a fine-tuned IWM world model matches or surpasses the performance of previous self-supervised methods. Finally, we show that learning with an IWM allows one to control the abstraction level of the learned representations, learning invariant representations such as contrastive methods, or equivariant representations such as masked image modelling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A factored support-operation readout with injective binding recovers held-out compositional edits on Shapes3D and COCO but not on a rebuilt MuJoCo substrate.

  2. World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

    cs.LG 2026-04 accept novelty 7.0 of 10

    WAV self-improves action-conditioned world models by cycle-consistent verification of state plausibility and sparse action reachability, doubling sample efficiency and lifting policy reward by over 22% on nine tasks.

  3. Galileo: Learning Global & Local Features of Many Remote Sensing Modalities

    cs.CV 2025-02 conditional novelty 7.0 of 10

    A single multimodal transformer, Galileo, jointly learns global and local features from optical, radar, elevation, weather, and land-cover inputs and outperforms specialized models on eleven benchmarks.

  4. Separating Representation from Reconstruction Enables Scalable Text Encoders

    cs.CL 2026-07 accept novelty 6.5 of 10

    Separating representation from token reconstruction via a bipartite CrossBERT architecture restores scalable frozen text embeddings and enables high-masking complementary training.

  5. SR-JEPA: Learning Predictive Latent State in 3D Scenes

    cs.CV 2026-08 conditional novelty 6.0 of 10

    After deleting a whole object from an indoor scene, the frozen SR-JEPA predictor imputes a latent that identifies the object's class at 43.13% macro accuracy, 22.18 points above the strongest floor.

  6. WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video-based benchmark shows that frontier AI models lag far behind humans on high-level world modeling and long-horizon procedural planning.

  7. The Lov\'{a}sz Local Lemma: Foundations and Applications

    math.CO 2026-03 unverdicted novelty 5.0 of 10

    LEPA predicts geometrically transformed patch embeddings from context and transform parameters, lifting MRR from <0.2 (interpolation) to >0.8 while keeping competitive PANGAEA segmentation scores.

  8. Adaptive Bayesian Single-Shot Quantum Sensing

    quant-ph 2025-07 reject novelty 5.0 of 10

    An adaptive Bayesian variational quantum sensing protocol selects probe and measurement settings by maximizing active information gain, demonstrated on a simulated sawtooth phase-tracking task.

  9. Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A survey reviewing how world models and agentic AI could be combined to give edge devices predictive, proactive decision-making, with a taxonomy of methods, applications, and challenges.

  10. A Survey of State Representation Learning for Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A six-class taxonomy of state representation learning methods for model-free online deep reinforcement learning, with selection guidelines, evaluation metrics, and future directions.

  11. Toward Embodied AGI: A Review of Embodied AI and the Road Ahead

    cs.AI 2025-05 accept novelty 4.0 of 10

    Embodied AI today sits between Level 1 and Level 2 on a new five-level roadmap toward all-purpose humanlike robots.

Pith tools