Pith. sign in

Paper Citation Record · LEDGER

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2407.05996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05996 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:28.458480Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.161878Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d349d7c-ca9e-4355-a268-b399094cd15d · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.447400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:38d36180fb23587da39f228f17a37fc392e2706b9b9f25769ff112c18acf4a2b

Observation 2ad859e1-dc7e-4e85-928d-49d597bcbacf · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.308512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:951b5a533ae2ebe2e1ece39b450e54a9c6d3c89f94d36436c00f7a1310cbdf93

Observation f8e1ec0b-6798-440c-a082-68b7a88a3788 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:32.782607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:aba8ed3d21eb87af19d831d1e1acb9a571f25f1ac3d802147bff431f5cc2439d

Observation d5626ecc-b478-4453-9d54-eb8a36d32af6 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.258877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:83640f2264ab6dac886f075def510e18627066e54c13ad9fb9341aa0ad24596d

Observation 20674eb8-9ca8-460f-8c2a-7d9c7a584ea1 · inbound

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt cites this paper.

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:28.458480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:28.458480Z digest=sha256:55dc5137e131186598a55bcbbe2f105ebafdc4023bcef90bdb8786d41b59b84d

Observation ca04556d-919d-4dd7-b074-3d4ea17824ee · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T08:04:12.547881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:ac31edecdabe58b06ce076b5b312763c583a625c46701ed1d295e1d60bb72a99

Observation 76e7b3a8-92e5-4153-b977-6f651ae96d02 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.058916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.058916Z digest=sha256:17073689c8450e24c24ef1ac6df31c8fa56991c3325d9c018c1479585578f5d1

Observation fd48f158-af31-4e1f-9d87-0fe1f2dc01fc · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.525433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:d283e21f49f073df7602b5aa7fd408ffeb331804b53863b78f9a3b86a6ed580a

Observation ffee3e7b-3f88-4282-9ee9-d2b050e62ed1 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.635769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.635769Z digest=sha256:38dc5f5df6901f8d37d1a65b69f265de35f64dea9f0ede8e57be7376178d626a

Observation 7951d0f7-f695-444d-90ee-be9ed1e47dac · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.986035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:e4303c8eb2634e356e0be146d2dc6ece94d51eea3da262f58134933ee92835e7

Observation ee7f89a2-ab41-4a9a-b748-95bfdec19392 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:52.285243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:52.285243Z digest=sha256:dde3e884f0d1a07b7a6532d700086ad951af985f3edfe3ede42a20ce537970d1

Observation fc590509-8576-4970-bb43-2010eedc415b · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.002359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:fe01ead7bb55c0d06e964340c3108503c3923fb056894c230ad12743d012edd0

Observation 15fbe715-5e6c-4768-9716-468af86a2bc5 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.396182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.396182Z digest=sha256:f6baf61ba8717b7a917779c99911a38f30689073148a6bb6ffea30a289df4be1

Observation 54761429-a8f5-4208-9360-2821397c809c · inbound

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation cites this paper.

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:14.643503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:14.643503Z digest=sha256:96ca6d711977fef11b2e55c6816b37e073d09b7bed2e26a6a23a555f981b5e0c

Observation 3bbb0ec4-6d9c-496f-9a46-98a8775494e1 · inbound

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory cites this paper.

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:49:51.103217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:49:51.103217Z digest=sha256:d6a0b9166f7e373c0869bcc1675f9421bad711098cbcb262082cccb28f411a9c

Observation 254afa48-c502-4b35-ae74-3827c9e0af7f · inbound

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning cites this paper.

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T23:00:31.128534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:00:31.128534Z digest=sha256:6d785f9880d0d21ba5ec5b2662f6b42595e4e7ad84a3dd20e92ce779b598d19a

Observation cdeb824a-dafd-4038-a29b-6688a7a3bf9c · inbound

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories cites this paper.

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.377605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:08:16.186576Z digest=sha256:28c6f2f494c127d6092728b4fb18189bba9f3cd7270098f5d512d0af5d0c9c19

Observation 9508c971-13d5-46a9-ac4b-f393b047b750 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:51.465622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:776b80d9a1369cc1d029c12527f97ff2062cc0cc26101433dedd9c2517a6060a

Observation 931e186a-74ee-431a-a788-864391bcccf0 · inbound

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies cites this paper.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:17.564383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:6d918d51e0af61b66335530cfcd94d7298d510cfe453813f44c5c24734ed8711

Observation 9ea12bec-3613-4844-bf3b-0cae5d759fef · inbound

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training cites this paper.

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T01:03:18.366526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T01:03:08.765350Z digest=sha256:373355d16d7b23e117095ac5b24cab14f06db22c5dc1e3ad8a2e00c6c48c662a

Observation 685b7205-5126-4cb6-b2cb-1f1abe86f475 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.600695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:46fea206921d32cd68d7efe4f68a7f2f570916f61163f5db22cd0ca24e09eb0b

Observation ad7ffa25-6642-44fc-b6ec-fe6c10308560 · inbound

What Are We Actually Benchmarking in Robot Manipulation? cites this paper.

What Are We Actually Benchmarking in Robot Manipulation? Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.333611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T09:31:39.345777Z digest=sha256:fbdd2ec94d96c814611ff8554dea72afb901aafd2eb34e6dc5d6022eb6321dc8

Observation 558b8666-e791-44fd-980b-315c4bfa1be3 · inbound

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling cites this paper.

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.506820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T14:27:40.585782Z digest=sha256:7c87cff2f9ef31a36e515a28cea1980cb97a767ea9e1b4f23b10484849290291

Observation d84f9e18-49a6-4d60-812c-133f99a2d089 · inbound

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure cites this paper.

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.163333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:32:52.472573Z digest=sha256:52408f6e207d4df1345b85d12a8290e6104618443af01df3d872f553cb9cd7e8

Observation 369bdd64-4844-48fb-95e8-985d5619ca7b · inbound

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging cites this paper.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:2c0d4b9154009c234f341204b927e97453f452c458358dbc6b8fdb71cfc21d36

Observation 1ae464e7-d62a-4c6c-83a4-c5bc459552ef · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 213

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.162534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.162534Z digest=sha256:2f8c3687825aa1c3f824637fff38599b33630c8c82d6c38d57bdd0caafac7a8a