Pith. sign in

Paper Citation Record · LEDGER

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

As of 15 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 10 inbound Pith citation observations for arXiv:2412.11974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11974 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:29:18.351588Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:20:10.174839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:51.757758Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b6cd4505-8830-4dc7-a541-50cac6a930dc · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning RT-H: Action Hierarchies Using Language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.232853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.232853Z digest=sha256:aacd5d57af9bc3f199c19e979884684855499dd14fca5600448c28dc5f7ee6bf

Observation 7cc30af0-d5ea-43f4-8c55-f0707043de11 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.238803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.238803Z digest=sha256:06b50fdcb19a4ba501f39fbe07c7cb9a0f95fa8f980ef1e6055bf8b3f974a7b7

Observation 09aa5587-cdb5-4475-9ffb-36ffbe75bf78 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning RT-1: Robotics Transformer for Real-World Control at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.245483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.245483Z digest=sha256:cda91e8538931bc2fc25c0f3e64afff8848f93bcb00ad2ccad21d570d321a3e1

Observation 1a1bd2f0-a86e-4a1b-8549-d7bc9ab342d9 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Diffusion policy: Visuomotor policy learning via action diffusion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:18.690276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.250810Z digest=sha256:3a53efbee4ea7bc37736a26eed2e85e61a109a854dde122335566e087e8679a5

Observation 997481be-b3c8-450c-b094-7504d99584f2 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.255358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.255358Z digest=sha256:37ac544132cf9ecba9163bbc322e402a49ab567187cb4fec7de7b72acafce502

Observation 3ba60856-e6be-4c2f-8b02-5ab8f6016880 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.260717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.260717Z digest=sha256:55120bbd600a35a5d0a56d26ee00f9c73307dde84d3317c533d53583b360075b

Observation 9bc56b6b-f5c4-4e6f-97d8-9a5d8a95c42b · outbound

This paper cites Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.265138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.265138Z digest=sha256:41ca0902b5eb2842a9d6fa8ea7a509d8089cf48ee8057c948b2773707d261b91

Observation 9d47a63a-c54c-460b-9ced-dd4bcdf31b2d · outbound

This paper cites Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.270169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.270169Z digest=sha256:7f101a7a4d56bd24dda94b7462ace90c0b79aa758702e0041c7c022e11682583

Observation bc07015f-0303-4d66-be00-602563c8e055 · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:18.668818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.275027Z digest=sha256:5fe461d9808e0e12d16a3fd531ddadba7b59b0bdeb0a4490280b11708e2a4aa0

Observation d05f4355-9f83-4d30-a77d-5381e2f5360d · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.285160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.285160Z digest=sha256:89deab1b776b99954d110a08f73b2721a9ed1ccfad7b46af62a7eb55e960d12f

Observation 62882544-299f-4151-848d-7cfd6eb8f4d7 · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.289576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.289576Z digest=sha256:0fbd49458ad9974c672890f854d6732e301d301b259e23be89a5280f1912166d

Observation f2d26b66-27c2-44b7-9c9f-ba91e5247067 · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Code as Policies: Language Model Programs for Embodied Control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.294678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.294678Z digest=sha256:ba79ce7263f401359168387ba0d9ae1d76c5e7ddd602f3f367560353fb3d1098

Observation ff73361c-77d5-40e7-ad17-1cb3178839e3 · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning A Survey on Vision-Language-Action Models for Embodied AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.299220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.299220Z digest=sha256:a355b8f7ec92d4be6e23b94469d8c4af851019d5e8b33124fb3aef42fe0a1f13

Observation 152a4107-bba9-45a2-88e1-8b2b896d502b · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:18.635734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.304480Z digest=sha256:b244e27be0f9904ded93799265505402f1c8a43f27b1ae3af641fcd03867147f

Observation 5cf0247a-68a6-4ee6-8e7f-b0f75cc66735 · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:18.621317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.309031Z digest=sha256:20706281b6439027384a55e797c1aab23798905dc9a2d3ff6bf05768937421b9

Observation ed080214-a205-4f30-8ce8-e07eaa0244d4 · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:18.608278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.313697Z digest=sha256:55ee5015294fa65e0de04ab46a546275c373087d9f349623e54d737d55b0a492

Observation ca46fc1f-46f3-43de-a4cf-2f120463032a · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.318541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.318541Z digest=sha256:5d09ccc2cca042f68d9f8a34265ff511944b9fc02b18765001f1c2dee9e43ede

Observation 44fbcea9-22a1-43bb-a169-ccfa06a8a2ac · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.323759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.323759Z digest=sha256:2adbcd2fd2fc7e7b7d896d6b9f8b0fc88fef8e7d24b502e855c05b1a006116ae

Observation 0825053a-6adc-4eb5-b5d1-16ac5b2ade44 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.328466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.328466Z digest=sha256:913bae6056d5750ba81e6a1d96f9e7a95b33aae6d1ce9c12f47e51368cfe28a1

Observation 35caf18f-76c1-4f83-b99f-4e7b9232e13e · outbound

This paper cites Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:29:18.584779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.334051Z digest=sha256:c011561f9215287e300fd459af97e24e974e6d782bd4e082808957e14d172e44

Observation 8b47159c-ad58-4ede-a848-aeb7bf82cd65 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.338102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.338102Z digest=sha256:9ab7bba005f0d54b0dd26986f82a023dca06b616c2e75c43b4ebee094ee7fff1

Observation 6e35bab8-d661-4ec9-b5f3-c118085e30a2 · outbound

This paper cites an unresolved cited work.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:29:18.568277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T14:29:18.343328Z digest=sha256:48312bfdc80bb813fa454a60f00a1fa1b6f23443abbf733b7925de6bb2d5f40a

Observation 81a0533a-290a-414d-a544-51cb7f8b0f3b · outbound

This paper cites online" 'onlinestring :=.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning online" 'onlinestring :=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.347048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.347048Z digest=sha256:835ce2e4b9658c7007d93592af3b006e27bded9f114ca4d03c3a903561a707b5

Observation f774aba1-3ce1-4470-85ed-34d7e6748668 · outbound

This paper cites write newline.

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:29:18.351588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:29:18.351588Z digest=sha256:4db79c46cdfa884ef6d2695b1ca82aafbc51b1d68016da47bdad71e1821442ee

Pith citing papers

Observation 63cbf8f5-25ec-4bc6-83dc-d99c703edc15 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.222889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:0dde9580935f7ca9b78bc1df4bbbda162ec2f1918563cf36ff960103b9b85ee3

Observation 40d1d015-218d-4320-b434-92ee724f9dcb · inbound

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks cites this paper.

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:53:29.482119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T15:53:29.412890Z digest=sha256:65b733ea3ae3d65b53ff507820342c7a89402cf9e996487f65a859cd97058d38

Observation fb77c175-999b-480a-96f4-23524523c667 · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:05.281797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:53:44.901684Z digest=sha256:fc17c300a0bcf30b776afe470d1f49941fb639ee0b7f256b9b3ca0129cdd748b

Observation e8c2aec0-2777-49a3-871b-a3f5f472a521 · inbound

MolmoAct2: Action Reasoning Models for Real-world Deployment cites this paper.

MolmoAct2: Action Reasoning Models for Real-world Deployment Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:57.288841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T00:59:54.787472Z digest=sha256:4e08c4505471aad208565363be0ce3e03fe633f4e9dc6957728a072d64b6895e

Observation 7c69d778-3eff-407a-95fb-8fce7d7dd0aa · inbound

TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models cites this paper.

TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:28.475024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T10:09:08.056968Z digest=sha256:99913a375d1ead71ff007f9d0e7fff8c2200d5c0d0524627d5776f156c178d41

Observation d85cff84-3b21-4767-8162-8ae6d8ee21fc · inbound

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation cites this paper.

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:46:33.195770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T09:40:04.685274Z digest=sha256:b9f49cfb65e1f8e5e66a72d0a0e629066decbdc9923947ede115ceeb9812df2a

Observation 71091751-41ad-47aa-b4b0-07c0d5d1d1f9 · inbound

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation cites this paper.

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:59:51.759123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T04:48:28.868562Z digest=sha256:5783134d5e178f3bf459810ef44e9b73995737e4bbd37d5aedd2ff9fca2dcd9b

Observation 10575531-55ff-41fe-be5c-9ac41b018e49 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 271

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.547106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.547106Z digest=sha256:de026e103baeb7d6bec24ce0447dedbfca8bdec6ee4951717c53055c97d04dba

Observation 77a40b28-65d4-491e-98f1-f2e9adcb9a84 · inbound

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction cites this paper.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:19:03.467276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:19:03.467276Z digest=sha256:8ab92abd9ea50b5c15678e3f2d27e93394bf7aee1626ac7a5986ff293e26eb7c

Observation 4578b24d-9fc7-49a8-816b-69b7f5d34e25 · inbound

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction cites this paper.

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:20:10.174839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:20:10.174839Z digest=sha256:1dd91f028530d97f57f1c5906801788f962d97e5f943dc419e8b2d2428da4ef4