Pith. sign in

Paper Citation Record · LEDGER

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2503.19757.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19757 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:10.945513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T14:25:47.243057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45883420-8b1b-49c5-b289-cdd0952cddb7 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.245817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:8af03626f3a11a3f539cc810e19c92651809658b7a2e69cdce6d76ca796b7be1

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:2433f9eb86928e8ce2a36d78f708881e96faa5c78ec0f28e9a391fbf7338d787

Observation 81c2056a-eaf7-420e-95d3-022c4fd77128 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.968090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:2e72ac7cb114ee3e47dab2a7048257e696231ca245d96381ec77565d2f921404

Observation 529a2236-55b9-43e9-bbd7-244fb6f8dbf0 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.584985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:821ca147b3f213803b780ee8625d22c5ecdef11e0398dea01e9ee9f51c8a7598

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:5538d6bb614ebc6e29ce3afe187ffd641a1c9c631674e8e4880b2d81db1d16a7

Observation cd0ec015-300b-4053-aadc-25259e758238 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.793652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.793652Z digest=sha256:3fdf50ba714c8bf109b83ad0e52a604da46941b3514c310a533fae9ef3519410

Observation 77f6fd8f-3d92-47dc-981a-a7cbc7c3b8b9 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.247428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.247428Z digest=sha256:99860b5bcb20af4bc95117c06a03aff386f9cbddad97ac47a72697d2db129388

Observation 9c4cbf47-44f7-46b8-8c6b-00ce61b3bda3 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.628932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.628932Z digest=sha256:bb0c90568df3a73624cc24824f58cc04db3e1286ef49248709c6e43945dec5ce

Observation 1d7a26cb-bc20-4393-9bc1-12fc57289878 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.274529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.274529Z digest=sha256:ef28204a02261748458f7f6a43a067e4b79f647fe2dbbf1f0da3c44f858a57d0

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:1a1ed42b514ba1ad3a840e1d417c9e3e5051b3894cc4335b435224f179ab2256

Observation a3f25050-fd4b-489d-9bea-fcd5afbff5ed · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.245994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:8b5a78b640bfdd8dd1c7103f639fa89c025007a1509a3da76c615eb9c236fe51

Observation 867a5f79-f679-49e5-9187-90e85b92fdf8 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:12.928751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:5baeba6b024c99e60f5dd7902a6a320a76931cca07ac6ef9b8221877d963237b

Observation c4026ade-74dc-4276-8eaf-1bf3f6496f0c · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.338230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.338230Z digest=sha256:4ddd4f6e58370ad8bcb4b5aec4c7559e5a0e0a9aad04886f00e04444ef3d8bb5

Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.038317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.038317Z digest=sha256:5545e0c0121bbe10cf0441292a69925716f7ee584665e880bc66fa9a3f367564

Observation f42c87c8-2ccc-48bc-9c1d-ecf8e8e44b1c · inbound

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception cites this paper.

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.529823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T22:00:06.298685Z digest=sha256:80af43be011d36cb934fac505a56b259222913f97114d92cbb0212225c4a1444

Observation ecf86412-32d0-4d38-a605-6e5ac4a440c0 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:e241532f54ad6fe99ecb2fabe6da3284d6ac1cbba8326d4dc404671ade59596b

Observation a2ccbabb-8d6a-4879-9fec-c936219a8c24 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:53.215480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:687bda7d6b60f08b64bb40df74270542db27a2368c9587b7aa4c769790f882be

Observation d53dc7d5-3131-4113-b37b-d60799f5d88c · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.139142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:f700de5bbfb08ca8045b334577ef03de3685b10e538d43b8ef1e5db70c84292c

Observation 9c652f31-f114-43ee-b75b-03749e6410f7 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:51.313611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:51.313611Z digest=sha256:ef0de4928d1f0d690993f0d250d9aec920ae546c725c29b56380be6f10c034cc

Observation 87a5084d-c753-4d62-ab24-8d4b5e328164 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.973759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:903a9aecd3d0c6a7e035828a553c395158207777779755bd0b7ac3300ac7ad42