Pith. sign in

Paper Citation Record · LEDGER

EdgeVLA: Efficient Vision-Language-Action Models

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 10 inbound Pith citation observations for arXiv:2507.14049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14049 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:14:51.731026Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:09.241377Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.333762Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcc61c4b-b9e6-4380-a4cb-ea1bec758bb6 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:52.114122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.643918Z digest=sha256:2ee3f70a046eae538e9cae9c0905abfa1b2ddb8ec79a86d75481ad028ad80099

Observation 7cabeaff-8ea3-401f-bf07-ce88bd2fdb57 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:52.093783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.649176Z digest=sha256:d5eca06238226898f06ddc7ef1ba492ba994b1a4227caa14dcf451511e817e46

Observation 4b55040a-e5e3-4fa9-bc2f-ea89ea386e74 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.653392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.653392Z digest=sha256:5cd6d879cfef8b19df3056e0651d4e5c7468644775c4cfa6ca89b3c9c0208cd4

Observation b4339069-68af-4cb4-90d1-244af27c6055 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:52.062527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.658018Z digest=sha256:4c7c3e43ad2bf2053f3d20cc4f4f1e431ca583710e2f7202ca1050feb78ccb33

Observation 1d2cb8fd-f9c3-4f45-9572-544ca6f370f6 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Flashattention-2: Faster attention with better parallelism and work partitioning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.663390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.663390Z digest=sha256:e65d914ccfa26bc5f39c614f572be560fff48743345dd5f239dee1538c517a4d

Observation e58dc51b-1e4d-4ff7-a6ae-16b7d2198660 · outbound

This paper cites Zhao, and Chelsea Finn.

EdgeVLA: Efficient Vision-Language-Action Models Zhao, and Chelsea Finn

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:52.030724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.668497Z digest=sha256:5b9b9b7a3439b98d33d3ca707a5d64e9ffd037824da6365be4db55f5b6cfa0dd

Observation 017e0517-d1c4-42a9-828b-8ecd93c5d542 · outbound

This paper cites Flexattention: The flexibility of pytorch with the performance of flashattention, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Flexattention: The flexibility of pytorch with the performance of flashattention, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:52.006801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.673240Z digest=sha256:45fb321fa51fedaf40d36e9892cb2105e2232323ff5951d8cd31067bbb5b8e27

Observation ba47cdc3-8731-48ee-8c30-bbe99711d169 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.978880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.677745Z digest=sha256:492002afc78ce4de96ce147e99c3ea9f940d22a52242adae48666ee8dafb0d15

Observation 7cb397e5-b48b-4eb7-82f4-501a6295911a · outbound

This paper cites Openvla: An open-source vision-language-action model.

EdgeVLA: Efficient Vision-Language-Action Models Openvla: An open-source vision-language-action model

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.960675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.681757Z digest=sha256:00bfd35cecc0a34a3bcbd23ebb520cf876808e8334c5bff41ffd1293249066c6

Observation f1bd67e5-551e-4282-8c46-95159e0ca72b · outbound

This paper cites Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto.

EdgeVLA: Efficient Vision-Language-Action Models Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.943638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.686592Z digest=sha256:94cf1d0acca23c4ccff3a7c5fdfa486ce02ab618b461340fd6beba2076d081e2

Observation 724075ea-6e53-4700-bb1e-5d6264548882 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Video-llava: Learning united visual representation by alignment before projection, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.691041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.691041Z digest=sha256:d336251e5ae8e71ec036a5d8f5719dafb23c30549138a7db317efc8c74515228

Observation 5e746a9c-9553-41a7-a8eb-28ae70598b29 · outbound

This paper cites Dinov2: Learning robust visual features without supervision, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Dinov2: Learning robust visual features without supervision, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.916406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.694992Z digest=sha256:30634dc42594cdff561173b76ade706ce8b68b88f7bd6de5970c7d8c9eb713ea

Observation b33b3167-bea9-4fad-9bea-eae5ff169748 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:51.895106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.699660Z digest=sha256:d2d05c0d208a51e31e8be2ef2635b92bfeb9778429c7ffd2940cb793604ef200

Observation bb728f67-722d-45f1-a62c-cd881427c286 · outbound

This paper cites an unresolved cited work.

EdgeVLA: Efficient Vision-Language-Action Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:14:51.878082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.704415Z digest=sha256:af44777b1ead6295400b29cebe33b3d455e3244da86eb38ee064ca7e845d75f1

Observation 18e69dac-09a5-407c-b9f2-59e05049e314 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Bridgedata v2: A dataset for robot learning at scale, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.861172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.708513Z digest=sha256:91dda93343d31d8d68c7ddf1198d0b591263e6ca4b3dd6c99f96685883c540f9

Observation 26f78f88-05c1-42de-913f-ab7c8125e026 · outbound

This paper cites Bitnet: Scaling 1-bit transformers for large language models, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Bitnet: Scaling 1-bit transformers for large language models, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.712660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.712660Z digest=sha256:74ac5747978f457449e9c0490b89e56a547b3a1b9eb11c4e1e597212ed86fda9

Observation 259752fe-5b07-4191-99a8-1d85cd21750f · outbound

This paper cites Qwen2 technical report, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Qwen2 technical report, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.716782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.716782Z digest=sha256:ccefe7f7ecf9b5761b037a7d41a72bf43a8808d7d72bc11720ff90f057e08916

Observation b5798ae5-359a-4d01-b392-842f0dd672d9 · outbound

This paper cites Homerobot: Open-vocabulary mobile manipulation, 2024.

EdgeVLA: Efficient Vision-Language-Action Models Homerobot: Open-vocabulary mobile manipulation, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.814516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.721214Z digest=sha256:956bd4611224d6f9b1075785fd13e13a0c501e626bc3b1d2f440f7b5af51ed30

Observation 7d9cf56f-60fe-4b63-a7e9-e3e93e31a559 · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

EdgeVLA: Efficient Vision-Language-Action Models Sigmoid loss for language image pre-training, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:51.725819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:51.725819Z digest=sha256:88896711ae30576e7d7dcb95d54bafb7ef408be4195d1918fe5d5942bc94a976

Observation 56a71ae0-c734-48bb-9384-25db405ab0a4 · outbound

This paper cites Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn.

EdgeVLA: Efficient Vision-Language-Action Models Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:14:51.774308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:14:51.731026Z digest=sha256:6fa32dd9ab260bbc0e95422da9a1230b4e639ccfde3c50654dfc6ee1b2702f26

Pith citing papers

Observation da9aa6c4-60e5-45c7-9ff8-4f003dadc063 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning EdgeVLA: Efficient Vision-Language-Action Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.552950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.552950Z digest=sha256:83f43e2db11862ab7c964dbda8999aab48b47bc38568f9ab9be5e1ed1f90b93a

Observation 54c57a68-59f9-4697-a2a0-c616ee2cd437 · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models EdgeVLA: Efficient Vision-Language-Action Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.496820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:128ea752a551fc1a08e3991b629421c4b4424b7187c9e53647b323f3e3ab1468

Observation f7bbf31f-8fc3-4012-ace2-eedf8b7042b9 · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control EdgeVLA: Efficient Vision-Language-Action Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:41.041488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:41.041488Z digest=sha256:b4de9f06304e5da5bc82d7982e1ee290283312ea1f24d07bb1f69fda444d7c33

Observation d8f0805e-1214-469c-9032-b9c6bbcbe125 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models EdgeVLA: Efficient Vision-Language-Action Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.582101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:e3034e6b5eb1a271ee8529c5e3f4b335e0bb4893230d344c346b24613aead0dc

Observation 6bf5f007-1226-4aec-bbca-a9fb2570e77d · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness EdgeVLA: Efficient Vision-Language-Action Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.424179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:41357c1c763dfa194c7eac806bdf708ef086e12dda04b876207cc29ea8c21387

Observation 002d6968-f493-4674-985d-6952177d4333 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs EdgeVLA: Efficient Vision-Language-Action Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:05:15.140852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T08:02:13.188363Z digest=sha256:61c1699a7bbcb7f7dc5cfba660509485d26ac3873de1199efba9d626f307e07a

Observation 50cac3e0-43bd-4d76-9562-69ca0c71d509 · inbound

FASTER: Rethinking Real-Time Flow VLAs cites this paper.

FASTER: Rethinking Real-Time Flow VLAs EdgeVLA: Efficient Vision-Language-Action Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:50:01.129618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T10:48:56.280105Z digest=sha256:0f4a1d348f46221a2b322d5096137cf643e785e2d5b3cd7ab7444eb1126b47d7

Observation 763f57ef-d2b5-480a-bad2-867d8b36c257 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model EdgeVLA: Efficient Vision-Language-Action Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:50:51.536005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:85632043562121e033c5bb2e37de4c6d9b08e71a4e5649626b8083780b91a820

Observation 9c12160e-3380-4ccd-b26b-127e60ada1c1 · inbound

PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models cites this paper.

PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models EdgeVLA: Efficient Vision-Language-Action Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.335102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T10:27:00.283283Z digest=sha256:ec4dbbb7d347531f0f96b67bf73514806ede4a3cdcbf2ce8a40d8abe154d3d19

Observation 38a9c4d0-2bbf-41ba-83b6-b07b40aa5dc0 · inbound

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving cites this paper.

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving EdgeVLA: Efficient Vision-Language-Action Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:09.241377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:26:09.241377Z digest=sha256:0d49cc546ae8d57b80bd27e7742b4d7b7180dd044bbf7d8130ab3f2e6f58a845