Pith. sign in

Paper Citation Record · LEDGER

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

As of 7 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 18 inbound Pith citation observations for arXiv:2505.21200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21200 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:43:04.537493Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:36:11.929173Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 42200279-e32f-4538-8d46-af1268e082e5 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.620579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.620579Z digest=sha256:322a03fefe8d4ead99c26a2d081f9ec1177a86927a41051aa953a88c194bf210

Observation 220d2467-d9a8-4f82-af99-bfdd8cdf092a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.759097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.759097Z digest=sha256:66555baf74dfc3ee0ad5a1e96f89b2c079f2d36c4c5bc2d4ecf424d4a6ed6a09

Observation f040c164-af62-4b84-a709-3bca321d088b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.929949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.929949Z digest=sha256:cfa1d08669f7b2e8cb3056b70fbe6c1f75f21cfde0c90856d577b5e1ec09aa9e

Observation b33ab496-7c76-4034-bbbe-d09dfbaccd17 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.076821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.076821Z digest=sha256:35d68130c2a225f5d875a378edb95a2a8d5cc758214d51f6bbb49385efc1ee17

Observation b38532cb-ba0d-403a-a6c3-4d68fef0db36 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.173764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.173764Z digest=sha256:c181885a18d24204f7f93218cff46a180179b6adb53f025a44fd9e92e5908658

Observation 262519e2-caa4-4cbf-b62e-00dbd37ca820 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.768722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.237915Z digest=sha256:594406bf8b59e74cffa654811f1a621c730c9c436a27adfe612197d00ecd32bf

Observation 5552c08d-1b32-43ac-aee0-2431f622ec04 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.651756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.492720Z digest=sha256:3fcf2e94a5b839d5b8ebf92fca2e3ccba5bdd9b845602f4a3240a47fcb71c69e

Observation ef69b40b-65a0-4648-8698-b9ac0637225c · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.466170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.647436Z digest=sha256:cc77f36ad97c6067e37789da677a48ddf991635eddec1cee97b0ad9cfeeefcce

Observation 2e7eda50-524d-4190-a527-f414ba586b84 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.302972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:42:59.712646Z digest=sha256:8204f2750933c1ee11baf896df5c1734378a8e3413097c910092d067ef8545a9

Observation 14bfbcab-b152-44de-b9fe-5e4af94d01e5 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.879271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.879271Z digest=sha256:3c6637ee29f0210f2aef4013314f9c0f8b4a1b786b427b89ca2bcbaef284ec2d

Observation af051aea-0e19-4b7f-96ae-f55f54d1c16c · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.008350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.008350Z digest=sha256:29763a079b2f2a3941b9b00e114ba2041260e4a8168ca33483eacf851bf3332d

Observation 8df6e184-76dd-4822-a79d-9717c2273671 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.230985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.230985Z digest=sha256:54ffa597812a0e162cf0b6bc367cb601dd18ccc2f762ab9b6bf3bac00075da94

Observation e6e1485c-b068-44a3-b2bd-e5fe14d36143 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.431971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.431971Z digest=sha256:9f531d2fa0009350f40788ddc1733f4eecec1a14b2a4ef81936cb48620efb51d

Observation 55fea985-5ab4-45d7-ae53-ef1212805730 · outbound

This paper cites Diffusion Transformer Policy.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion Transformer Policy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.597211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.597211Z digest=sha256:cfbfdef2a2b9e02968d8174dc47e839ab443e3e637f4ff16f16223106192793e

Observation 9e645849-c1fb-48f3-96b6-e46d5bcf4fe1 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.735532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.735532Z digest=sha256:cf0815f1b28f6b2ccc6fdd38dcf2d41223b9f34faaa973a96e99fe18b6d98c1c

Observation 5f2091d7-02ca-4580-a627-db9cc2403b6e · outbound

This paper cites Jiang, X.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Jiang, X

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:06.120893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:00.879965Z digest=sha256:cc0fafae02efb4b743f4adb61d0d7e951be0ac9d7cc2800c1901da06e0a4e8d1

Observation 9d404b8d-14f8-44d9-86af-fd0f71ed1923 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.083162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.083162Z digest=sha256:02f244db9ad0064affddd35440779fda512837277206a6a7bc39c3167e4038c9

Observation 4050d5a9-723a-4ca2-aeaf-cabc24b5d3b5 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.279474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.279474Z digest=sha256:035ca0f613bc3498270dcbc972b3b95a8cf59dbb5be908e290feccee01bba734

Observation 9b4c248c-1307-4904-93bb-016fe5ee8e86 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.351486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.351486Z digest=sha256:de95be4e6b49677cb7324a426f647dbd87651890f97c4ecf9ee069966e4ffc56

Observation 664d0ca2-60ca-40e3-9239-9fe0ecb9f715 · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:43:05.160857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:01.359705Z digest=sha256:f67e0c2187f3d6f013f72ed3bb27d445982dbb70c9e58aa16d4a49d2b96f09e6

Observation f06cc0f4-5215-4d09-866f-98dd11e095e7 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.480814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.480814Z digest=sha256:cbae0a5c57aafb26eb09b237252789833841140ffa0d80510a749af9ae7d5122

Observation 34ebeda1-0fb0-4921-8ac1-4c0902d3e029 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.638742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.638742Z digest=sha256:838d551bc3f7b987275d14ada991646d3f73ce4cfca9fa38d08e366b12fef71f

Observation a517b294-5f40-42d8-90fe-a0e1c789012c · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.760587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.760587Z digest=sha256:22ff9ca369ccb5b6cd19fb3fc3aa941d54b527656e0fde0b75cc60479d5358d0

Observation 83926484-51c0-4457-93df-48ae6ccede47 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.916518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.916518Z digest=sha256:404f32b5471a5cb421ddc078215f61daea73fa2a43d6e5993778b030f89fd109

Observation d0bc70a6-f547-44cb-8a61-00668dc6b4f8 · outbound

This paper cites RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.033013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.033013Z digest=sha256:99825d5888ec92a3a46af3225a45a1b8dcd7b8729bd0c22472a18b0e5e13bc85

Observation 8e3fbcf7-60c1-4579-bac9-cf9dbd6b82d4 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.178483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.178483Z digest=sha256:c31181af269c2ce7739c898b5c4f4419788035f34f7160aa8e249032ef743163

Observation 6ba2346a-6f98-45c9-a41c-80c1c81e0659 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.340523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.340523Z digest=sha256:5e48541005816f51d7754a170b1560c8cb1e06aeaeab7a0b589152a66fe8a88a

Observation bede8346-9c03-4645-b231-b8d865439650 · outbound

This paper cites Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.506958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.506958Z digest=sha256:8811e274fce0325d5f45acbf6cc5a087fb7333aa74b5875b1291be1174c4483d

Observation 7b4e0d74-7f2c-4ab2-8c46-f9f0c3947d2c · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models R3M: A Universal Visual Representation for Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.629364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.629364Z digest=sha256:beada860c702d3e75daa94ec2bf8c61be69ede8ab24fd87917a29af0ed5bc2ff

Observation 47a2d80b-83af-4431-8303-138fafda7b9b · outbound

This paper cites Oquab, T.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Oquab, T

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.739636Z digest=sha256:9a311a96606ee6d98816f9ddcf1b1b748d0eeda65e4e1d11d0197e49ad37a2ec

Observation be8158a6-1d9b-4d35-af86-66d8b09a0150 · outbound

This paper cites O’Neill, A.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models O’Neill, A

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.744927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.744927Z digest=sha256:e933e0887608280bd385e70325944f583f15b5ef7b3d16bb0ce9d6ac803fc9d0

Observation d8b20428-4c1d-403e-9387-b86e1b507580 · outbound

This paper cites Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.750457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.750457Z digest=sha256:aaf97b5b228b7bc6603f3167470ce1e6efb71a87568c2f580ebc6051223a0736

Observation bae20548-c5b9-4c76-a6e5-58533845c6b4 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.850761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.850761Z digest=sha256:b5c89f6ea4af42761055066467848c7518beeb2203287130d5fea49adaaa7a83

Observation 5cc97333-604b-4e29-9a59-121840267127 · outbound

This paper cites Singh, V.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Singh, V

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.924852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:03.002933Z digest=sha256:dfb874b00c44319bef64b633c495e35215425bc9231f12e3a0c9b28ea5422546

Observation bc2b18de-deae-4f08-ad3f-df8ae4ddf3ed · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.201022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.201022Z digest=sha256:7d8b5c022e313653c38efbee5698cca609af6785bf80716ab71b6dd90cef7d12

Observation 849551f8-8e3d-41bb-8fd4-fde8f3c680d5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.323929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.323929Z digest=sha256:75da09b62627d803b945abd1596748fe675548cfc0d486bf7e5f7e51fb621b35

Observation d9b01b24-1d17-4365-b681-446034d39b78 · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.455619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.455619Z digest=sha256:f9f7ad93ad1331ad3757a6a6ddd639a1b4d9ce710c50e9ab279e80b4cf685c0e

Observation 0fc3c282-6fd8-4637-930c-58cea05178ab · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.601909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.601909Z digest=sha256:a3ccf6b42ce31ee8d9dbefba1655a514a25d53aeb397fee32d0ff6203ffe3253

Observation d9482747-4466-46a6-ae62-9b5a24e150ee · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.749499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.749499Z digest=sha256:7a5f8859ae83676b3fa2feed0e9c4f1b749f9b029a297974cd3d1516ac23fb4b

Observation f8fc4f28-c74a-48c7-881f-8a7e7109521e · outbound

This paper cites DNAct: Diffusion Guided Multi-Task 3D Policy Learning.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models DNAct: Diffusion Guided Multi-Task 3D Policy Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.817109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.817109Z digest=sha256:e440d41a13c05b6d93a113f81a4965dc598f228e76c3992b59bb5e01d8aa0473

Observation abffa18c-c2c1-47ef-8a27-82f0e2f3eccf · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.910286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.910286Z digest=sha256:5d47f04ca79ed8ec4a0c6374a5ff3609f60c6958d379b0b0350a62cb9c09774d

Observation 8c50a432-9f19-4125-981f-ee77a805da2e · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:05.796712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:04.018067Z digest=sha256:d3b74c396d044c61505507872495eab5772ef17ca6073f8eddf0b15ca0bc1bb0

Observation c769917f-87c4-458b-bc29-fa86fccecb79 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.147405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.147405Z digest=sha256:2fdb5850c798d7d4e638452e4e9d95bfc5f8c3ae8e1d311007e197e6d09ab552

Observation 4e812d10-ea7d-45f4-80f4-3e28783cfb9e · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.258730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.258730Z digest=sha256:b86ec2c674583ac5afb19f721a232a666edebaa581aedf912cb5003365912ffa

Observation 87e0c8c6-c14a-4a5f-9e62-32f80e215ed9 · outbound

This paper cites Zhang, Z.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Zhang, Z

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.638348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:04.323508Z digest=sha256:8dcbb99e25ebb9c3cd643129774fcbce4e30100212e7671cad1358e42ebf152e

Observation 549af06b-7b60-4c27-9aa4-e09d009bb627 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.407246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.407246Z digest=sha256:9cf80bd9ba022975ae271cfa1f0b7e2a1e7068564f3c708a8b782f1c9c3f6c31

Observation 5de09fb8-0c6e-4ec0-b728-13a0d27c0ed1 · outbound

This paper cites Zitkovich, T.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Zitkovich, T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.434319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:43:04.537493Z digest=sha256:e57326e60d20d0dbba37881ec4229e93264f6c3fd96ab733529575eefe15305e

Pith citing papers

Observation 27e0a52e-db45-4355-a05b-7d132211f3f4 · inbound

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models cites this paper.

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:36:11.929173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:36:11.929173Z digest=sha256:181feefeed82d8c821b17fc8fb7e06fcf212233731cc6cb8cfb3f28ccc37efde

Observation e9ad02eb-0de1-4f98-89bb-7bfaae0e40a2 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:28:16.251395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:2e8782876d77851ed7aac161a7ca8a287ab1e0dc2f793f858acbc4b31c9424b3

Observation 5cc907b8-c95a-409f-af39-091b8910feed · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:12.680018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:12.680018Z digest=sha256:654657ea8a45eaf895b6c0a636d0adf943fe24de471f6b9594f021335af05cd8

Observation c7669fec-08b0-4352-868d-3ff350b648b3 · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.433198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:5cadf5bcfb34bd2bf1b84059f8caf738aa469aa88175b88d9131e391aab2c929

Observation d75ee6f8-1593-40a5-aff7-9f64f760f4ba · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:39.309018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:39.309018Z digest=sha256:50f87ec84f6ff2d97f7dcaaefed2bb3839e874068c040e062131132e1cbb59c5

Observation cdeed0c5-557e-4128-8df2-49523da42a74 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:01726b2569f6960a82cfa2e13b0368b649cd522f39fe511a9a30e6b9bdb66c10

Observation 566736a1-5b5a-44d2-9456-127d74258f25 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:50:03.934133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:7d86cec5e48a0571821175f6fa90d93785f8553610904ad7cdc3fd13fc9a904f

Observation ebda3e5c-1bb5-4920-babb-7c268ddb6849 · inbound

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success cites this paper.

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:50.142804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:33:26.708966Z digest=sha256:a117f12fd5e7567d2501b6b6a388b64858795f2edf9bb7788e04bf7a3a2069ed

Observation a65763df-9030-4369-a569-943e7c27f4f5 · inbound

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models cites this paper.

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:22:06.687892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T02:19:39.604140Z digest=sha256:01bd8040fec358ea5e3b4426c82b99b72d25a7bb3f0e7798f7fc72657a29acd5

Observation 2a3220a9-f94f-44ad-8e8d-4501e7d6aa96 · inbound

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models cites this paper.

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:07.104557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:13:12.442859Z digest=sha256:3e259fb725da03e898efa7f54ef578b7ba5df6ab79fc54f55fcaefc6ccdeea7d

Observation 346d9813-74f7-4774-a93f-75f41c53f5bf · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:47:37.220705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:43:08.029165Z digest=sha256:3c07ca1c097400825c03cf8468b536c19002a4799600217a65cd51a781769310

Observation 012a3ec7-7c05-4075-a2b5-c77670f3a936 · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.902573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:42:05.592911Z digest=sha256:7d6974b5ef546c27682001043d80bddb00b8c4621e5bcb1390e37baad639ff83

Observation 207fda53-dc32-48a8-a34f-e553e199d3a8 · inbound

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models cites this paper.

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.041269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:16:27.229603Z digest=sha256:aaf5eddb6414b1fe5be7b6ebcfbcdf8b091c34f62b2f4c54cab2076c29c89114

Observation 43146e1b-4c59-4bc5-a8d3-d2e9ed35e07b · inbound

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models cites this paper.

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.619734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:11:10.246045Z digest=sha256:9ea8262621b7017210f95acb802f0b39c125b7390e0d44487f16146fefb7821e

Observation 10a0d5dd-2dfe-427d-b614-7934d9e4435a · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.661540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:c10b7b3d3d1776e12ae411e8ac3b882110d32d9b98ad288259871170d0288f67

Observation 8bdac25c-b34c-4d14-943d-f0123da50301 · inbound

Improving Robotic Imitation Learning via Trajectory Standardization cites this paper.

Improving Robotic Imitation Learning via Trajectory Standardization Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.564349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:37:26.864436Z digest=sha256:716ca3ef33fe0436d997b6952e2f19db9d70939af3b3d6cbe948edc155583d1f

Observation c89b7deb-90ab-4cc3-b7c5-789c8b0edc79 · inbound

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? cites this paper.

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:52.300386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T04:56:00.919208Z digest=sha256:a277dfef061cd041d9595f590c398972d048cfde7eaf6c3dad8ff984e9dbecd0

Observation 6ee57c48-dd03-412e-ba93-5ae81aa47fbb · inbound

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation cites this paper.

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.864902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:07:45.792132Z digest=sha256:934fbc8ad9149fb80ccc9b95f9c22bf33d69e25f23cf4d5b1087dbed0e3355a5