REVIEW 11 cited by
Learning and Leveraging World Models in Visual Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Joint-Embedding Predictive Architecture (JEPA) has emerged as a promising self-supervised approach that learns by leveraging a world model. While previously limited to predicting missing parts of an input, we explore how to generalize the JEPA prediction task to a broader set of corruptions. We introduce Image World Models, an approach that goes beyond masked image modeling and learns to predict the effect of global photometric transformations in latent space. We study the recipe of learning performant IWMs and show that it relies on three key aspects: conditioning, prediction difficulty, and capacity. Additionally, we show that the predictive world model learned by IWM can be adapted through finetuning to solve diverse tasks; a fine-tuned IWM world model matches or surpasses the performance of previous self-supervised methods. Finally, we show that learning with an IWM allows one to control the abstraction level of the learned representations, learning invariant representations such as contrastive methods, or equivariant representations such as masked image modelling.
Forward citations
Cited by 11 Pith papers
-
Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
A factored support-operation readout with injective binding recovers held-out compositional edits on Shapes3D and COCO but not on a rebuilt MuJoCo substrate.
-
World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry
WAV self-improves action-conditioned world models by cycle-consistent verification of state plausibility and sparse action reachability, doubling sample efficiency and lifting policy reward by over 22% on nine tasks.
-
Galileo: Learning Global & Local Features of Many Remote Sensing Modalities
A single multimodal transformer, Galileo, jointly learns global and local features from optical, radar, elevation, weather, and land-cover inputs and outperforms specialized models on eleven benchmarks.
-
Separating Representation from Reconstruction Enables Scalable Text Encoders
Separating representation from token reconstruction via a bipartite CrossBERT architecture restores scalable frozen text embeddings and enables high-masking complementary training.
-
SR-JEPA: Learning Predictive Latent State in 3D Scenes
After deleting a whole object from an indoor scene, the frozen SR-JEPA predictor imputes a latent that identifies the object's class at 43.13% macro accuracy, 22.18 points above the strongest floor.
-
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
A video-based benchmark shows that frontier AI models lag far behind humans on high-level world modeling and long-horizon procedural planning.
-
The Lov\'{a}sz Local Lemma: Foundations and Applications
LEPA predicts geometrically transformed patch embeddings from context and transform parameters, lifting MRR from <0.2 (interpolation) to >0.8 while keeping competitive PANGAEA segmentation scores.
-
Adaptive Bayesian Single-Shot Quantum Sensing
An adaptive Bayesian variational quantum sensing protocol selects probe and measurement settings by maximizing active information gain, demonstrated on a simulated sawtooth phase-tracking task.
-
Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
A survey reviewing how world models and agentic AI could be combined to give edge devices predictive, proactive decision-making, with a taxonomy of methods, applications, and challenges.
-
A Survey of State Representation Learning for Deep Reinforcement Learning
A six-class taxonomy of state representation learning methods for model-free online deep reinforcement learning, with selection guidelines, evaluation metrics, and future directions.
-
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
Embodied AI today sits between Level 1 and Level 2 on a new five-level roadmap toward all-purpose humanlike robots.
Discussion (0). Continue with ORCID to comment.