Pith. sign in

Paper Citation Record · LEDGER

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.23909.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23909 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:48:46.688889Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b8ec4c4-1967-49d9-9ac9-c6d5269d1e55 · outbound

This paper cites R., Finn, C., Fusai, N., Galliker, M.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling R., Finn, C., Fusai, N., Galliker, M

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:44.751279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:44.751279Z digest=sha256:52ca8c950897713072daa93d0f8a0c6c6ed151e17f85f1f6d870ad3269ed066a

Observation 45153cf7-ec1c-45b1-bf07-30562e48f280 · outbound

This paper cites R., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., Levine, S., Li-Bell, A., Mothukuri, M., Nair, S., Pertsch, K., Shi, L.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling R., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., Levine, S., Li-Bell, A., Mothukuri, M., Nair, S., Pertsch, K., Shi, L

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:44.810098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:44.810098Z digest=sha256:b6440eaed0bd03751a2ec3f38f5c31d77eb7bce95d03eed26a83145b8eef5a60

Observation ba3868d5-24da-466b-82d7-0f207c555b02 · outbound

This paper cites WorldVLA: Towards autoregressive action world model.arXiv preprint, 2025.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling WorldVLA: Towards autoregressive action world model.arXiv preprint, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:44.902169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:44.902169Z digest=sha256:ef5a358386069d0bdfb3558b223044945999916eb5ccbf91a4f0fffb88fa9556

Observation a7d8b37a-888e-4a5d-8373-bbd259ebc0f9 · outbound

This paper cites Unified diffusion VLA: Vision-language-action model via joint discrete denoising diffusion process.International Conference on Learning Representations, 2026.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Unified diffusion VLA: Vision-language-action model via joint discrete denoising diffusion process.International Conference on Learning Representations, 2026

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:44.996042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:44.996042Z digest=sha256:e563f942ec013a233c062d867a967b3072d3b844c06f8a608bb60af14ef70fa6

Observation e17122ab-644a-4a89-a488-75c14e2f4b2f · outbound

This paper cites C., and Song, S.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling C., and Song, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.080272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.080272Z digest=sha256:dcf771caa3b3057666f27bbc7c2ed6082b7c5312d6b8c8a217c1002978f59733

Observation 8acfe0e2-1c4e-4599-be77-2487a6f9598b · outbound

This paper cites R., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., Luo, J., Tan, Y.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling R., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., Luo, J., Tan, Y

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.144789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.144789Z digest=sha256:4d90c1d450cc1ecc43cfd50040b8ef93c0269f96bdf72617ebff9d12e363828c

Observation 2c35f8f0-e87d-42cd-8646-a0b7c4f5fbdd · outbound

This paper cites VLA-0: Building state-of-the-art vlas with zero modification.arXiv preprint, 2025.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling VLA-0: Building state-of-the-art vlas with zero modification.arXiv preprint, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.222319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.222319Z digest=sha256:2ab67697c22df8d4b1b5dc0473521e893e5f393ea4fff30b0a4bd8faa7283dfd

Observation a27de74a-2002-4171-9da0-991a3d8d8f81 · outbound

This paper cites Diffusion Transformer Policy.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Diffusion Transformer Policy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.306314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.306314Z digest=sha256:b2ba4c788336e9a5e1b25538335b51bc3526870ee000048717bb3961e766f571

Observation 6f0571ee-3b41-4206-a5c3-2c11a65a3369 · outbound

This paper cites Paris: A Decentralized Trained Open-Weight Diffusion Model.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Paris: A Decentralized Trained Open-Weight Diffusion Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.352980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.352980Z digest=sha256:747237baff35d9a7b7ce1ab4562dcf41a61856212e9b8b9e176820fc655c0a2e

Observation b633aedd-3632-41ad-b5d5-2e1586d01c46 · outbound

This paper cites J., Finn, C., and Liang, P.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling J., Finn, C., and Liang, P

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.439223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.439223Z digest=sha256:e73f05cdb7066e4ed6a09ff4f4f2206bc1be7b2c64059c9733ce6c7fd5bb44cf

Observation 29e7a71a-5549-45db-b779-d31b8e816787 · outbound

This paper cites J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.509611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.509611Z digest=sha256:ae6fb5abc2561a8f087189ff0ceb808db5a2b61181b75a70541c2e33db46c6b6

Observation 6bb4fb0c-7272-49b4-b31f-7e5a816ff1d5 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.578890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.578890Z digest=sha256:f302b6dfb13ef4b6028728c7973e02c057c2919c042f3edc959a215ffa49c572

Observation 259c3a9a-e7df-4788-bb35-7a53dc814f5e · outbound

This paper cites an unresolved cited work.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.624419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.624419Z digest=sha256:068db4a499d06e5e62f96bd2d870d8ee6f72e34b52f461a9462efad0573f92cd

Observation 38b3f130-845f-4347-ad07-99edb075dc36 · outbound

This paper cites LIBERO: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling LIBERO: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.692505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.692505Z digest=sha256:23e7b0c8a60f8bb805e17b5983894dc1f78b923a8b229f49739a9db881ef357b

Observation 3e3173bc-89e1-44be-a1a6-7269122dc9a8 · outbound

This paper cites Mmada-vla: Large diffusion vision-language-action model with unified multi-modal instruction and generation.arXiv preprint arXiv:2603.25406, 2026.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Mmada-vla: Large diffusion vision-language-action model with unified multi-modal instruction and generation.arXiv preprint arXiv:2603.25406, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.757721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.757721Z digest=sha256:7ac14bcab082883c4d5a519d32f0593decae956fca9f0e01121c2d751fe6611c

Observation 2ea3062c-d515-45de-b4b4-5db61554f78d · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.846693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.846693Z digest=sha256:77a696029c4798331417c5cdf93207b8490ecf89ce8d3a03eaea533ebbfe344c

Observation e6558121-bae0-4e49-8ab2-877168dd385e · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:45.943231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:45.943231Z digest=sha256:c26b88fe5961cfcb403918e35ea50e535b98c4a26747f6a4712442d6089dc2c1

Observation 41ca5027-2759-4b18-bdcb-81c2a97b4aed · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.003376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.003376Z digest=sha256:49bffac24da102b1713c5be83c255d031f83c03ca9ef108020362ea15631e59d

Observation dbb895a7-64ab-4688-b576-2323741e7465 · outbound

This paper cites Paris 2.0: A Decentralized Diffusion Model for Video Generation.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Paris 2.0: A Decentralized Diffusion Model for Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.099526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.099526Z digest=sha256:751c1e1bdfe479e8ccb8d2d378c44a78de6334627f4fad0243100c5fa871cc67

Observation bd56e4d3-86c5-4ad8-b951-b70afd058b7e · outbound

This paper cites MemoryVLA: Perceptual-cognitive memory in vision-language-action models for robotic manipulation.International Conference on Learning Representations, 2026.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling MemoryVLA: Perceptual-cognitive memory in vision-language-action models for robotic manipulation.International Conference on Learning Representations, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.172625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.172625Z digest=sha256:38ec0a88d421dd22bc2ec3bac03bd911ad0d9843b4af9f7297c69e26b9710549

Observation 1ecd454f-c79d-4a3c-bb95-3fafd6fd2eb6 · outbound

This paper cites Predictive inverse dynamics models are scalable learners for robotic manipulation.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Predictive inverse dynamics models are scalable learners for robotic manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.268663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.268663Z digest=sha256:0ae69c4134a631cc9fc5da1050f1a5d0b0ebc3ca6e96db21e31b12b712e3971e

Observation 8604dc53-75cd-4b93-8ac6-651988149aa1 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.Proceedings of the AAAI Conference on Artificial Intelligence, 2026.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Vla-adapter: An effective paradigm for tiny-scale vision-language-action model.Proceedings of the AAAI Conference on Artificial Intelligence, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.335039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.335039Z digest=sha256:e792fb0fede8a95ca9ab502b984034cc40c44d003fac0da658923aa1dae16c74

Observation 8de45522-a48c-44fc-8e40-d2efca39f49c · outbound

This paper cites an unresolved cited work.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.394629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.394629Z digest=sha256:c04884ca822ef5f40c6ecf1deed36a47f1e301432c14e063e74c32dc08d7b593

Observation 98bf2814-8257-451f-ac1a-c06101b9c781 · outbound

This paper cites Dreamvla: a vision-language-action model dreamed with comprehensive world knowledge.Advances in Neural Information Processing Systems, 38:24195–24228, 2025.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Dreamvla: a vision-language-action model dreamed with comprehensive world knowledge.Advances in Neural Information Processing Systems, 38:24195–24228, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.449826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.449826Z digest=sha256:121f228f95fdb63db045aa098aefe18506ea2ba8f9301e290720f7324a1b4005

Observation 31f267a0-cefa-4dc3-b896-cb6c365b63c3 · outbound

This paper cites J., Fu, Z., Zhang, Z., Wu, Y., Li, Z., Ma, Q., Han, S., Finn, C., Handa, A., Lin, T.-Y., Wetzstein, G., Liu, M.-Y., and Xiang, D.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling J., Fu, Z., Zhang, Z., Wu, Y., Li, Z., Ma, Q., Han, S., Finn, C., Handa, A., Lin, T.-Y., Wetzstein, G., Liu, M.-Y., and Xiang, D

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.540715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.540715Z digest=sha256:5fdb3f326b23375e706c3e6578708b77983c3a984e86ee42413a8ffd81ab79d9

Observation b6391398-3ceb-43ec-8adc-a1d69bb62fcc · outbound

This paper cites Acot-vla: Action chain-of-thought for vision- language-action models.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling Acot-vla: Action chain-of-thought for vision- language-action models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.620534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.620534Z digest=sha256:e64f0121d1c21e478bd5a79a0728a51ab8a5c887f947e73e515f441c7760e4db

Observation f3a24a34-a7e9-4782-a231-0710918a6d66 · outbound

This paper cites FlowVLA: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269, 2025.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling FlowVLA: Visual chain of thought-based motion reasoning for vision-language-action models.arXiv preprint arXiv:2508.18269, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T01:48:46.688889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:48:46.688889Z digest=sha256:545b480afacb7421f693755b6406c1ec273fd8518080f21717870836706eabfe

Pith citing papers

No inbound Pith citation observations are available.