Pith. sign in

Paper Citation Record · LEDGER

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2407.05996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05996 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:51:28.458480Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.161878Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d349d7c-ca9e-4355-a268-b399094cd15d · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.447400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:d78fc1b69be360f4bdefdf2f3f897278cfa2b35aefc11354687453c98f879bf8

Observation 2ad859e1-dc7e-4e85-928d-49d597bcbacf · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.308512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:b3573caff7255a77ca2c30323d97321845f4f4a7bc65ea1366d5cca0f21621d6

Observation f8e1ec0b-6798-440c-a082-68b7a88a3788 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:32.782607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:b135fee791c9ada134a22abcaa6acc334f301892ac95ed7b903212cfd9a4b9e3

Observation d5626ecc-b478-4453-9d54-eb8a36d32af6 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.258877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:9f61e89c1450909c7563a1d369a62426c2aad14637f743e33c8a0cd7a1dc7ad0

Observation 20674eb8-9ca8-460f-8c2a-7d9c7a584ea1 · inbound

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt cites this paper.

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:28.458480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:28.458480Z digest=sha256:55dc5137e131186598a55bcbbe2f105ebafdc4023bcef90bdb8786d41b59b84d

Observation ca04556d-919d-4dd7-b074-3d4ea17824ee · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T08:04:12.547881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:a1f31fbdc1bebdfab25f36dde57aa71b91951b1b5127aa643b6060148f10f4f1

Observation 76e7b3a8-92e5-4153-b977-6f651ae96d02 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.058916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.058916Z digest=sha256:17073689c8450e24c24ef1ac6df31c8fa56991c3325d9c018c1479585578f5d1

Observation fd48f158-af31-4e1f-9d87-0fe1f2dc01fc · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.525433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:9d5c297ccae1928747a330b31bbd07a4df6f0c8c954f2962711bc43f5837c7d9

Observation ffee3e7b-3f88-4282-9ee9-d2b050e62ed1 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.635769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.635769Z digest=sha256:fd49da79ad2a831e0996b187343661e3ca64e0e5eaa80d5bab61f33fee3bd9d9

Observation 7951d0f7-f695-444d-90ee-be9ed1e47dac · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.986035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:ceaa5469a9a50dff581d1f582f0c902278c8b923acf8dabea0ab8a85dc441d21

Observation ee7f89a2-ab41-4a9a-b748-95bfdec19392 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:52.285243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:52.285243Z digest=sha256:dde3e884f0d1a07b7a6532d700086ad951af985f3edfe3ede42a20ce537970d1

Observation fc590509-8576-4970-bb43-2010eedc415b · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.002359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:9f442fc6355f3efaa4e1837fc1409d62d62087705caf38890a858fa29e006a86

Observation 15fbe715-5e6c-4768-9716-468af86a2bc5 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.396182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.396182Z digest=sha256:f6baf61ba8717b7a917779c99911a38f30689073148a6bb6ffea30a289df4be1

Observation 54761429-a8f5-4208-9360-2821397c809c · inbound

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation cites this paper.

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:14.643503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:14.643503Z digest=sha256:96ca6d711977fef11b2e55c6816b37e073d09b7bed2e26a6a23a555f981b5e0c

Observation 3bbb0ec4-6d9c-496f-9a46-98a8775494e1 · inbound

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory cites this paper.

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:49:51.103217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:49:51.103217Z digest=sha256:d6a0b9166f7e373c0869bcc1675f9421bad711098cbcb262082cccb28f411a9c

Observation 254afa48-c502-4b35-ae74-3827c9e0af7f · inbound

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning cites this paper.

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T23:00:31.128534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:00:31.128534Z digest=sha256:68a99738b01628c307ae323941e644f6872d462b75ec8172b368aa68cd2f2e5b

Observation cdeb824a-dafd-4038-a29b-6688a7a3bf9c · inbound

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories cites this paper.

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.377605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T00:08:16.186576Z digest=sha256:bc78980efa8f0b782423f511a9de7d2e60806f2f91e42e80f2b2a32bc6982a92

Observation 9508c971-13d5-46a9-ac4b-f393b047b750 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:51.465622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:5a9e6fad2aa829e52bcbd1a62c7d8000d1bc10c0c95f8bcdbda000612470f52d

Observation 931e186a-74ee-431a-a788-864391bcccf0 · inbound

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies cites this paper.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:17.564383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:fc3d7494c2722f11e58ac1f2411de7a533705c96c961f3beada246e062dbeed0

Observation 9ea12bec-3613-4844-bf3b-0cae5d759fef · inbound

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training cites this paper.

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T01:03:18.366526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T01:03:08.765350Z digest=sha256:0980cdba8335a5c673e89ad23fe76fcc966684c3fbe78d5d1b77364ec6511e60

Observation 685b7205-5126-4cb6-b2cb-1f1abe86f475 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.600695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:ae127eac0584218251ec471c45b24ea2c2f0a0d595a923eae6176de4b9798118

Observation ad7ffa25-6642-44fc-b6ec-fe6c10308560 · inbound

What Are We Actually Benchmarking in Robot Manipulation? cites this paper.

What Are We Actually Benchmarking in Robot Manipulation? Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.333611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T09:31:39.345777Z digest=sha256:06f5fbf6b5ecac215019838bf69b20aef83b469c0dd1f0096ce25874aac2d9f0

Observation 558b8666-e791-44fd-980b-315c4bfa1be3 · inbound

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling cites this paper.

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.506820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T14:27:40.585782Z digest=sha256:16af2cb0bcc8a0940e0c7c128e75885c2fbd14f7b6be4a980eb6e39c00b420ba

Observation d84f9e18-49a6-4d60-812c-133f99a2d089 · inbound

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure cites this paper.

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.163333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:32:52.472573Z digest=sha256:6960aaa238812c55cb721ce14d44b0a34f8b1ebdcbce4016836caad8cd23ee39

Observation 369bdd64-4844-48fb-95e8-985d5619ca7b · inbound

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging cites this paper.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:2c0d4b9154009c234f341204b927e97453f452c458358dbc6b8fdb71cfc21d36

Observation 1ae464e7-d62a-4c6c-83a4-c5bc459552ef · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 213

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.162534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.162534Z digest=sha256:2f8c3687825aa1c3f824637fff38599b33630c8c82d6c38d57bdd0caafac7a8a