Pith. sign in

Paper Citation Record · LEDGER

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2503.19757.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19757 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:10.945513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T14:25:47.243057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45883420-8b1b-49c5-b289-cdd0952cddb7 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.245817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:b1addcf44efaee0777ad8185226aff1be6a7d362176e024581b1d2eb486fe88a

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:0068fc8e5a872c7ef5b111d7f5263fbdc6d11409b8ba08d3bc0b3c7ed28fd291

Observation 81c2056a-eaf7-420e-95d3-022c4fd77128 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.968090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:c20fc477241e7cd6c5d1b137bb3a7925d802b58602c68a960e8c21ce7212d93a

Observation 529a2236-55b9-43e9-bbd7-244fb6f8dbf0 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.584985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:f3be8f3a4111a72d42f8eef6a13d3675653c9ef6b910af82ab8a23dcd795583f

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:b0fc838139d4af4746e1fb0cfab0a0060ebad6dbdff423d593edecae01966d4b

Observation cd0ec015-300b-4053-aadc-25259e758238 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.793652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.793652Z digest=sha256:87f37f78b4e5eeccfc71c88bb9d83998659fb1327b45fa07c55b3fb09f64d93d

Observation 77f6fd8f-3d92-47dc-981a-a7cbc7c3b8b9 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.247428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.247428Z digest=sha256:410ed1878710d49eac71cecdf606d4c10f08ac92b2e39e34acc8bf66e025f295

Observation 9c4cbf47-44f7-46b8-8c6b-00ce61b3bda3 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.628932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.628932Z digest=sha256:f8ba0f7890b948c3eec8e954c8d103e340b500f2a57ed3fe9236171e1fdd3769

Observation 1d7a26cb-bc20-4393-9bc1-12fc57289878 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.274529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.274529Z digest=sha256:fe61c03adaa52c34b44a2a54044b60ab951aaa68da2ac9f56be767e73387e121

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:658662948a9037755348654291111a951af0190f7b87133f750fa1f116fab423

Observation a3f25050-fd4b-489d-9bea-fcd5afbff5ed · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.245994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:71e8a84c324fcc49ba0baca83ee12429c845c0a3fee99d85764a8a879d0b62b4

Observation 867a5f79-f679-49e5-9187-90e85b92fdf8 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:12.928751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:ac121b3d19ca0f998124fc0f7e0bdd43c75e48f93a065c807ef9a00c133eee4e

Observation c4026ade-74dc-4276-8eaf-1bf3f6496f0c · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.338230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.338230Z digest=sha256:7aeb2e75b28932447302ba5ef5e0005d54918f28a3f42df2630f30b397a50258

Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.038317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.038317Z digest=sha256:fe6a9b3e000d0e42a785fd30832739ad13912d71f533195fc1c3feb4bc39dec7

Observation f42c87c8-2ccc-48bc-9c1d-ecf8e8e44b1c · inbound

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception cites this paper.

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.529823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T22:00:06.298685Z digest=sha256:97501f53eea88426f79e49e0875dac3fb07c094ffab10504c0d8557adccf2860

Observation ecf86412-32d0-4d38-a605-6e5ac4a440c0 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:13b885cf76a048103375fcf7529f90ac31759404a90199db04914880e7c878e0

Observation a2ccbabb-8d6a-4879-9fec-c936219a8c24 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:53.215480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:a41748aff04bce81c4fe9b4d0a29d499c202308075956fa98d22b1506eb5ccbe

Observation d53dc7d5-3131-4113-b37b-d60799f5d88c · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.139142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:c07f272dea742a7f7cf6312e7b2d7fc51bc8f275cb05aa640762db6e11cf6f12

Observation 9c652f31-f114-43ee-b75b-03749e6410f7 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:51.313611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:51.313611Z digest=sha256:3c77a8848b7f333834be816d09669ec0b0ff38181e1bd217999727829d60698d

Observation 87a5084d-c753-4d62-ab24-8d4b5e328164 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.973759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:02f011cd1ca171a8f70dcb31769eb0f490148f9494a72985afd689e3a1bec769