Pith. sign in

REVIEW 10 cited by

IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.00785 v1 pith:2SGHLGDV submitted 2024-10-17 cs.RO cs.AI

classification cs.ROcs.AI
keywords igorlatentactionhumanmodelrobotsspaceacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Image-GOal Representations (IGOR), aiming to learn a unified, semantically consistent action space across human and various robots. Through this unified latent action space, IGOR enables knowledge transfer among large-scale robot and human activity data. We achieve this by compressing visual changes between an initial image and its goal state into latent actions. IGOR allows us to generate latent action labels for internet-scale video data. This unified latent action space enables the training of foundation policy and world models across a wide variety of tasks performed by both robots and humans. We demonstrate that: (1) IGOR learns a semantically consistent action space for both human and robots, characterizing various possible motions of objects representing the physical interaction knowledge; (2) IGOR can "migrate" the movements of the object in the one video to other videos, even across human and robots, by jointly using the latent action model and world model; (3) IGOR can learn to align latent actions with natural language through the foundation policy model, and integrate latent actions with a low-level policy model to achieve effective robot control. We believe IGOR opens new possibilities for human-to-robot knowledge transfer and control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Foundation-model HOI work is organized into eight geometric, semantic, and visual sub-priors that enter six reconstruction/generation tasks and three robot-transfer routes.

  2. ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Cross-shadow prediction on appearance-resampled video pairs yields a unified latent dynamics interface that transfers demonstrated actions across environments better than prior latent-action and interactive world models.

  3. Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

    cs.RO 2026-07 conditional novelty 7.0 of 10

    Enfold folds the internal computation of a video world generator into a current-only representation, enabling competitive robot control without executing the generator at deployment.

  4. Is Diversity All You Need for Scalable Robotic Manipulation?

    cs.RO 2025-07 conditional novelty 7.0 of 10

    In robotic manipulation, task and scene diversity improve policy learning, multi-embodiment pre-training is not necessary for cross-embodiment transfer, and expert speed variation confounds imitation learning, so debi...

  5. Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Three lightweight LAM fine-tuning objectives (foreground-weighted reconstruction, primitive contrastive learning, zero-transition calibration) debias latent actions and yield stronger, cheaper robot-action following i...

  6. PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    PhyGDPO uses groupwise direct preference optimization with real videos as winners to make text-to-video models generate more physically plausible videos.

  7. Playing with Transformer at 30+ FPS via Next-Frame Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Next-Frame Diffusion combines block-wise causal attention, consistency distillation, and action-based speculative sampling to generate action-conditioned Minecraft video at over 30 FPS on an A100 with a 310M parameter model.

  8. WorldEval: World Model as Real-World Robot Policies Evaluator

    cs.RO 2025-05 conditional novelty 6.0 of 10

    WorldEval conditions a video generation model on a policy's internal action embeddings (Policy2Vec) and shows generated-video success rates correlate with real-world robot success rates.

  9. CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CoMo learns continuous latent motion self-supervised from internet videos and uses it as pseudo action labels to improve robot policy co-training.

  10. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A convolutional residual VQ-VAE action tokenizer trained on over 100x more data than prior work improves OpenVLA success rates and inference speed on several manipulation tasks.

Pith tools