Pith. sign in

Paper Citation Record · LEDGER

IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2411.00785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.00785 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:29:09.438211Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:09.783700Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 547818d4-8ee6-4805-ad5e-831fc3e2775a · inbound

VideoDPO: Omni-Preference Alignment for Video Diffusion Generation cites this paper.

VideoDPO: Omni-Preference Alignment for Video Diffusion Generation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:29:09.438211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:29:09.438211Z digest=sha256:a7008f221dd1b55908bfcbdd8c8f8bd3aaf1fe8ca2fec2a9fe6bfda191446065

Observation 9b581b68-ccdd-45de-b531-0391d1bd19c7 · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:38:11.313568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:33e7713e72839a27c43fc7a783d159009d97eeac9fb7032a93231c74e5133464

Observation f78bd75c-6e8c-4692-8777-791cbb8c8c34 · inbound

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent cites this paper.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.887809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.887809Z digest=sha256:ac063c4af2d8f909697574c796cea091d64e94d5829fb490213777cfb51d2dea

Observation 767b7ea1-b712-411a-9dc9-a8626f3ec87f · inbound

Latent Action Learning Requires Supervision in the Presence of Distractors cites this paper.

Latent Action Learning Requires Supervision in the Presence of Distractors IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:18:44.893686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:18:44.893686Z digest=sha256:51ec432c4d81e87d7d6a83ea30fead0daa5cac3a1f0ef42b15d3166ceb705cdb

Observation 91ebf4bc-00de-4437-80f8-de0c067b9d74 · inbound

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions cites this paper.

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:28:06.967896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T15:28:06.883492Z digest=sha256:a96934af869ceff34a2b2ded9b4369b8adebd4d4a839ba3af00598d6e13e9cf3

Observation 87d28d03-a0ad-4bca-a374-5a7b333b7764 · inbound

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning cites this paper.

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:33.139916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:33.139916Z digest=sha256:a71b5670648cddbfccd938a385ec1b70cd5547b9e2b566b5db852cdeb4614bb0

Observation 6d8aa1ed-3ea4-4adc-9999-2e46881603c9 · inbound

WorldEval: World Model as Real-World Robot Policies Evaluator cites this paper.

WorldEval: World Model as Real-World Robot Policies Evaluator IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:41.835024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:41.835024Z digest=sha256:a6c1b79a1755ea90e787c0c16e5336c2b529807ef1530ec4c18d16ee217882c2

Observation 794491b1-9a61-44bd-8399-e6d5650962a6 · inbound

Playing with Transformer at 30+ FPS via Next-Frame Diffusion cites this paper.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:42.207829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:42.207829Z digest=sha256:d94cbee9b6e2d454d1c99746f2fe279a9f5a528eeab1172f9d6dbf9d01f1c945

Observation 7cbd0345-3ac5-4cbe-85e4-6bda7755a4c8 · inbound

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation cites this paper.

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:50:29.231047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T07:48:33.832017Z digest=sha256:507f2c2580855d4d915d3199add2c19813e6b0868b86bf296ec674189cdda782

Observation 3827fc8f-d211-4da5-b32d-9e48b3d2b97c · inbound

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers cites this paper.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:22.759775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:22.759775Z digest=sha256:4dc284437b0b5161a00e46d6c33bb8abd7d6e0d5045a395e6ed282b95f6bb011

Observation 368c1dee-66e7-4e08-a7c8-d6b3dfddfcbb · inbound

Is Diversity All You Need for Scalable Robotic Manipulation? cites this paper.

Is Diversity All You Need for Scalable Robotic Manipulation? IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:10.758757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:10.758757Z digest=sha256:2fa18f8dde3d6507f110cccaf233f782e16f95926f627008705af10c3e0f847e

Observation dd2ef776-c9c6-459e-ad1a-8a7c1a38296a · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:52:02.980898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:31f93c99ad69ed16b1318d0445622337be3ec168f019d6379b4fa3d5b9e211a8

Observation 6efd5a22-aa1e-4ed5-92a7-f8b23806f8cd · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:44.036005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:44.036005Z digest=sha256:bc197f582a505acfa280da11699c5121c75491c1264e4ddf2beb5a3df0de61d5

Observation 30b947d9-96c0-4a33-b2c0-49c890709df5 · inbound

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos cites this paper.

DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:02:34.146487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T17:02:33.997887Z digest=sha256:d627b437df3dc313d17e40d5a7bc82929e90f8bef5121b7421d04348c5dfe3de

Observation ad1e5935-e2cc-422b-9efb-6a396a4b0f5c · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.665714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:f08ee571866d8f25eb35b287b843a5f44f32a6c4a8d657dbe82904192b3cb509

Observation 64c0227a-a4b0-48ff-b393-3a78a56503d8 · inbound

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training cites this paper.

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:36:07.668999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T21:26:26.540403Z digest=sha256:8b2b9b4cded298285384904940afdeb44a99950958efc1b60e0c4f0dcd764d48

Observation ba7bbf18-1bc3-4ac1-84f5-bfff8fdfe94b · inbound

GazeVLA: Learning Human Intention for Robotic Manipulation cites this paper.

GazeVLA: Learning Human Intention for Robotic Manipulation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:36:12.527612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T11:37:30.784513Z digest=sha256:eeb7cc925334fe9e8c7b6d497f532798cfaf61d55595d53ab418573a2807dadc

Observation 46bce8bf-d9c5-4aaa-aebb-c581e90d3716 · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:11.673687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T02:51:27.662262Z digest=sha256:ab4e67e897e8cb5e0f04d7415610ddab2f322356261ed6075d1764a5b742a53b

Observation 8de5b719-e340-4891-9f79-461ecb42440f · inbound

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation cites this paper.

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.238755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T11:16:58.104663Z digest=sha256:1818692ea8a4cdce28abc59e82a3369edb24ee688e5db340e732428846fad1d6

Observation 965afae3-47c1-41c4-a357-a712ee39d9ae · inbound

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models cites this paper.

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:16:09.928505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T17:43:00.453164Z digest=sha256:bc9b2cc9ed40ef075714c10f1134ca084324c43fd7a2c074b644f2ecf74ee103

Observation 7331027b-87d9-40cd-9f0f-fd2ab52caec9 · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:26:10.252214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T09:11:21.715023Z digest=sha256:98623186ccaf497059decb2cf17429496da3db6ce0161998102ae2a4913f0b1f

Observation 5594f8e5-fee1-41e7-86b4-ce891a6aa4e8 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:27.062999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:14:54.885244Z digest=sha256:34fa14b3fe0f81ae72d466535a91ce744df62ec2088b0959371cf5c10839f241

Observation bf2bea7b-7d9b-48df-aab2-66738c7dc589 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:59.505091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:14:56.501485Z digest=sha256:e46e0d60739a395ad4df5817ab9d2f25e1387b5746f1651be8113ba55ae4b32b

Observation 3efa652d-3532-4adf-8404-f23d6b9d4d17 · inbound

RotVLA: Rotational Latent Action for Vision-Language-Action Model cites this paper.

RotVLA: Rotational Latent Action for Vision-Language-Action Model IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:49:23.120637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T17:48:06.734816Z digest=sha256:3fcaeb654a90b2882ed8d771e57de61c455013aff41391ff570304311611b3b4

Observation ec9a79a5-9e97-473b-bd2c-63a9685361e0 · inbound

DiLA: Disentangled Latent Action World Models cites this paper.

DiLA: Disentangled Latent Action World Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:38:56.388138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:35:37.527479Z digest=sha256:1233602311634dbbce80f6cce78ce940f9639330f742bd3bdfab597ab1ca7626

Observation 3163ce5d-1ad9-47d0-966b-5d5c7d73387d · inbound

Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model cites this paper.

Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 3

Resolution
malformed identifier
arxiv_id, observed 2026-05-19T19:37:43.756786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T19:36:17.173279Z digest=sha256:06b07d61c1712a9250f8fd532ab25adb1ecd3343ce59fa6094c571138c687889

Observation e9adc3f5-8256-4c5d-b50d-54b231aa71d3 · inbound

UAM: A Dual-Stream Perspective on Forgetting in VLA Training cites this paper.

UAM: A Dual-Stream Perspective on Forgetting in VLA Training IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.099978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:24:57.339949Z digest=sha256:7125140edac708e7eddc0d411f46f07a55c59875621b26ebc5cdd0d2aa328fdd

Observation 32c9b733-a34b-4e5d-bc26-1b09967633af · inbound

Why Latent Actions Fail, and How to Prevent It cites this paper.

Why Latent Actions Fail, and How to Prevent It IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:34:05.330376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T08:33:02.745886Z digest=sha256:47ab8cbe38ea3c0704f3d20616ec7dc4b30e00edd81d6d43e531bd088412509c

Observation 172dd323-3d20-42be-9cab-eb70a0bdf684 · inbound

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data cites this paper.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:55:00.167016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:aeec8bf6eeb83e110c49e69496faf154ae2d388d55da3a9f042fe131020e3e83

Observation 8344fbfe-375e-40f9-94ba-00c48daaa38b · inbound

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization cites this paper.

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:36:29.583480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T09:55:00.402411Z digest=sha256:ff7fad45381237b12af87e941d10d46579514d526872b6c48cec512c6214b535

Observation 5c0ca728-7f52-46cb-9560-31f73155c281 · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:09.499733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T22:25:17.522240Z digest=sha256:0b90c3a015bb7a1050c6a8681104d739096c4133480827724f50e36352e314b8

Observation 77d76d54-83c7-4ceb-9f3a-5e1da62f7a68 · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:35:34.492973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-01T07:17:42.045939Z digest=sha256:2fbb1328129e6e1d160edec4cee3e74c4693c0af93ce245977dd6d6ad88740a9

Observation 92fadd4d-7535-4bbc-a91d-d257a94dbd83 · inbound

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos cites this paper.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:04.235028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T21:53:27.323112Z digest=sha256:64c22648eacdc3c83143f1588da1cea05b03aee5a3f62ca3300f9c773bbf5278

Observation f2cded77-958a-4ab6-a065-0eeb889d2748 · inbound

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos cites this paper.

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:01.900671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T23:44:35.238493Z digest=sha256:a92b25eeedd96d5a2eebfbe23e7491a81048adb82d5f9c9bd508aee3db311c85

Observation c276495b-7d4f-4f51-a882-7d22f3eeb120 · inbound

Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models cites this paper.

Imitation from Heterogeneous Demonstrations using Grounded Latent-Action World Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.568483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T14:07:31.139419Z digest=sha256:d01febea41f1547a0d61a95867a920537b8a07a810664bd41c2893ab0ac55034

Observation ef32b57f-519a-4253-b24c-a2b511622482 · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:00:09.785386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:cdc85d0f6a1a399c7eb1a8bf888290e2602f279888a95b4871c874f037e97f98

Observation db4fb4c8-8d36-4d37-864c-cb06c02743e2 · inbound

Causally Debiased Latent Action Model for Embodied Action Conditioned World Models cites this paper.

Causally Debiased Latent Action Model for Embodied Action Conditioned World Models IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T04:50:59.096166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:50:59.096166Z digest=sha256:c5ecf8621d05ef6c5efa02022c5967d470327256e96650268015f79e36430149

Observation 9d99e890-b6a8-4652-bca8-e6f8c2abf07d · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T11:27:05.835624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:27:05.835624Z digest=sha256:095a29ea590b20cf58ac85ba5f363cadd0bc414da3d0ebb0a9386e107a8c46a9

Observation 1756cd0f-e81c-4df8-a802-d10edc9e8d97 · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:48.964275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:48.964275Z digest=sha256:c118d24364ad152b662681b1c1c9a19bb6f791569c30c97a68f53643a440e07d

Observation 58573bcf-2373-411b-884b-74b3d872dc1c · inbound

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow cites this paper.

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T09:42:43.644737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T09:42:43.644737Z digest=sha256:d807d92c627b587bd41955b45d3c65104bb117f182bef1f601661b306bbf2035

Observation 720ef27e-345d-4bce-b680-ae4155102304 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 216

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:28.134244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:28.134244Z digest=sha256:caf3d1ee63bb55ad9778f959df6a85579c6a4e7be7f8e4d5506e8f56d4044e1b

Observation dc6f7df7-6c3e-44a2-8987-523be07b7926 · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:09.888871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:09.888871Z digest=sha256:fc3942723e8c555e4e4b0fbc57177007bafa127008ae5e5779ca83c5ba146a8f