Pith. sign in

Paper Citation Record · LEDGER

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 19 inbound Pith citation observations for arXiv:2505.21200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21200 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:43:04.537493Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:09.074689Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 42200279-e32f-4538-8d46-af1268e082e5 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.620579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.620579Z digest=sha256:a13bef2ba99e5b20395d67bd97a013dbeef2c6a6913cf93bc0ec235f4be1d0af

Observation 220d2467-d9a8-4f82-af99-bfdd8cdf092a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.759097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.759097Z digest=sha256:07f8047f9ae17fbd1f6d714516cd3d290f11fb54c564fc16166eaf985b25a68e

Observation f040c164-af62-4b84-a709-3bca321d088b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:58.929949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:58.929949Z digest=sha256:f39b979566ec91de15154217a15ceb51cdcdae424e879e0728153897a5a6a205

Observation b33ab496-7c76-4034-bbbe-d09dfbaccd17 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.076821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.076821Z digest=sha256:e4aec69e0e4b59f7eaaf70b9c6573930ff8ae8b59b2deb2613703a8d6de73e19

Observation b38532cb-ba0d-403a-a6c3-4d68fef0db36 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.173764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.173764Z digest=sha256:449c9cf9c752a99cf98461473359c2581f0258d78362fcfeb1e3cb7ab93dcf61

Observation 262519e2-caa4-4cbf-b62e-00dbd37ca820 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.768722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:42:59.237915Z digest=sha256:e7e514526cb3a551b6828e79e0ac7c51b9e04563bdc9e903ee31b659f4740e6a

Observation 5552c08d-1b32-43ac-aee0-2431f622ec04 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.651756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:42:59.492720Z digest=sha256:b12e1cd243b7e3df6881e0bb4f23d81de538b35030dc55745322924288b085fc

Observation ef69b40b-65a0-4648-8698-b9ac0637225c · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.466170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:42:59.647436Z digest=sha256:03c9f9a572d6996a9a3999ac5f8975cd794f8d3f5b1aeeba496ec42596e64a8d

Observation 2e7eda50-524d-4190-a527-f414ba586b84 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:06.302972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:42:59.712646Z digest=sha256:15ade526e9c2f788a3808cbb7c15149aab81056f8f42ba328a2273a2a8b9c8ef

Observation 14bfbcab-b152-44de-b9fe-5e4af94d01e5 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:59.879271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:59.879271Z digest=sha256:5b4964ba3b93f8577d1f1f99d59b18c03883d097d67b89bf3af222e66476e908

Observation af051aea-0e19-4b7f-96ae-f55f54d1c16c · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.008350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.008350Z digest=sha256:68e3ec437030258b3e301a051c2e071079e1ffbfdce1b30a55c5a29ed958e40b

Observation 8df6e184-76dd-4822-a79d-9717c2273671 · outbound

This paper cites Manipulate-Anything: Automating Real-World Robots using Vision-Language Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.230985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.230985Z digest=sha256:6db4b308fd6b00ae19d7d5cf9e9817ebd14b8d6a5c517e063a39211acf02ccca

Observation e6e1485c-b068-44a3-b2bd-e5fe14d36143 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.431971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.431971Z digest=sha256:4979fe489d6c5dd4df3fcd3c99cd8a861de82a57de863d12b660581d32bfbac9

Observation 55fea985-5ab4-45d7-ae53-ef1212805730 · outbound

This paper cites Diffusion Transformer Policy.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion Transformer Policy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.597211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.597211Z digest=sha256:39e11694b3cf3af3d180e72c46a1d269dbb745afbc194ddb5b8c36d2d96a7576

Observation 9e645849-c1fb-48f3-96b6-e46d5bcf4fe1 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:00.735532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:00.735532Z digest=sha256:2a1ada6c052be162cf0bc89e66ff3b0a2a26e520650fa17ddd62fe0a1a9d943e

Observation 5f2091d7-02ca-4580-a627-db9cc2403b6e · outbound

This paper cites Jiang, X.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Jiang, X

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:06.120893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:43:00.879965Z digest=sha256:d389a452b8c4c3f52097a7b818a1314646bbc3fe8a413b798c31dc193b948876

Observation 9d404b8d-14f8-44d9-86af-fd0f71ed1923 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.083162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.083162Z digest=sha256:6aa0e12eded93e80377eeaffda3a6ba0cf888faa4f6e307d29d2136434d87413

Observation 4050d5a9-723a-4ca2-aeaf-cabc24b5d3b5 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.279474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.279474Z digest=sha256:c478e80ad0210a4bc9ee7c13925065b24805ec737d69be507d4cadf08cacbf2b

Observation 9b4c248c-1307-4904-93bb-016fe5ee8e86 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.351486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.351486Z digest=sha256:bcfa0d59ace3c22b010d2644f5af3f297e404d1d3ac71860988c25a40990d02f

Observation 664d0ca2-60ca-40e3-9239-9fe0ecb9f715 · outbound

This paper cites Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:43:05.160857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:43:01.359705Z digest=sha256:704fe00fbfc9aaf18a90618ec319c5dc6d71816001c35fc96488fc58dd5fccb0

Observation f06cc0f4-5215-4d09-866f-98dd11e095e7 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.480814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.480814Z digest=sha256:05d473f8604905595a4ca1be8d7bd1a84a5de687a69f25ab65f5c0ae4ce46b94

Observation 34ebeda1-0fb0-4921-8ac1-4c0902d3e029 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.638742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.638742Z digest=sha256:090893b55a80d80ed368fe37a0e102ef423a29a21f7ca61cfd6cd81b1aafc705

Observation a517b294-5f40-42d8-90fe-a0e1c789012c · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.760587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.760587Z digest=sha256:6e15de80d0bab738ced7f4f62527f088f039109815238cc61e1d037cea805521

Observation 83926484-51c0-4457-93df-48ae6ccede47 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.916518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.916518Z digest=sha256:f94943568e401700965209a6ed4126ff41ca0e3cd784e3005ef1c91768f21b50

Observation d0bc70a6-f547-44cb-8a61-00668dc6b4f8 · outbound

This paper cites RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.033013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.033013Z digest=sha256:03c7e380c5d0a358b7372c19fe45fbb69cc71360e2ed3e2c9cb009005d2665e8

Observation 8e3fbcf7-60c1-4579-bac9-cf9dbd6b82d4 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.178483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.178483Z digest=sha256:e81cf09f8a1889a41a9c63dc64dc19c5ba923eb77e96300d2f82c501d3eb23d5

Observation 6ba2346a-6f98-45c9-a41c-80c1c81e0659 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.340523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.340523Z digest=sha256:dee91750fd5618d67b245b53e42c174cb872e3380b0ddbc1575d75777ba2b4dd

Observation bede8346-9c03-4645-b231-b8d865439650 · outbound

This paper cites Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Bidirectional Decoding: Improving Action Chunking via Guided Test-Time Sampling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.506958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.506958Z digest=sha256:a351fdd99fd333091d4be579f12438d8fe7d0e69202b8e2428ab26cb2e07346b

Observation 7b4e0d74-7f2c-4ab2-8c46-f9f0c3947d2c · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models R3M: A Universal Visual Representation for Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.629364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.629364Z digest=sha256:086a9dc7eabc75dc3f639d4e9d197d455f87d9816033320d4ba0816ad8b1c8a1

Observation 47a2d80b-83af-4431-8303-138fafda7b9b · outbound

This paper cites Oquab, T.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Oquab, T

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.739636Z digest=sha256:c018fc684b02880dfbd8d567a324a8d467c5825b83418c28ff6c7a0006208f7a

Observation be8158a6-1d9b-4d35-af86-66d8b09a0150 · outbound

This paper cites O’Neill, A.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models O’Neill, A

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.744927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.744927Z digest=sha256:c4bcd7e44f3b6dd49efa96facb0e4b41ed92276aead5c874cc2a866359b834ff

Observation d8b20428-4c1d-403e-9387-b86e1b507580 · outbound

This paper cites Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.750457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.750457Z digest=sha256:d6837713edb1baa4c9e4e9a776f686b70badfa7614bf30c1e82e2481e88e58ea

Observation bae20548-c5b9-4c76-a6e5-58533845c6b4 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:02.850761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:02.850761Z digest=sha256:e34b081e3d66b0e03360d8b29c86dd104c4d151ea8713de3c7203e983f1cf697

Observation 5cc97333-604b-4e29-9a59-121840267127 · outbound

This paper cites Singh, V.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Singh, V

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.924852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:43:03.002933Z digest=sha256:7481e7bf67cb8776067e1d5ef079d22f6f0c4c2273d2bda0090024acc2df0de6

Observation bc2b18de-deae-4f08-ad3f-df8ae4ddf3ed · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.201022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.201022Z digest=sha256:e6040bb148cb38d06e129225808fcceeb350c7845f124726ea6fbe902f4dda76

Observation 849551f8-8e3d-41bb-8fd4-fde8f3c680d5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.323929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.323929Z digest=sha256:aed75b77753ae24ee41a3fa6debd82727b9cf1556f4dac16bd983b2354e0f6b7

Observation d9b01b24-1d17-4365-b681-446034d39b78 · outbound

This paper cites Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.455619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.455619Z digest=sha256:5c50f348c685581f7466f0db9854353a6cea6f0bc79c635cd5a4f55cb2328480

Observation 0fc3c282-6fd8-4637-930c-58cea05178ab · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.601909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.601909Z digest=sha256:53cf5e904bc82cd962b61f43319feb18c050e70296c559d4f57a0dd09d336cba

Observation d9482747-4466-46a6-ae62-9b5a24e150ee · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.749499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.749499Z digest=sha256:21c57d350896d0227358b67d76a204689ebc267952443d08029cfa2b51da6973

Observation f8fc4f28-c74a-48c7-881f-8a7e7109521e · outbound

This paper cites DNAct: Diffusion Guided Multi-Task 3D Policy Learning.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models DNAct: Diffusion Guided Multi-Task 3D Policy Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.817109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.817109Z digest=sha256:cc7fd9947bcb7b25bef7230935032d08cf7ba398e764ec9952447da1ee500cf5

Observation abffa18c-c2c1-47ef-8a27-82f0e2f3eccf · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:03.910286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:03.910286Z digest=sha256:3a16cb586ec50fa5fc478706bae42ee4a0e92acb5272adfc16364d5ee9211e49

Observation 8c50a432-9f19-4125-981f-ee77a805da2e · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:43:05.796712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:43:04.018067Z digest=sha256:afb51078229891be054462c902d3d7e2c7cb43cd4aed83d4174a2efb0931a2c9

Observation c769917f-87c4-458b-bc29-fa86fccecb79 · outbound

This paper cites an unresolved cited work.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.147405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.147405Z digest=sha256:f46ca06ef6daf6bfabb052b261153945886205919dca1dec0a60998d7418acd4

Observation 4e812d10-ea7d-45f4-80f4-3e28783cfb9e · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.258730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.258730Z digest=sha256:29ce31e42d74909809ca3c4cb8df51efa2d1adc7a6714a1dd05d0e2c7687f574

Observation 87e0c8c6-c14a-4a5f-9e62-32f80e215ed9 · outbound

This paper cites Zhang, Z.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Zhang, Z

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.638348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:43:04.323508Z digest=sha256:b72c4ec905d5d151395c2f639befe332dd9872a439a877401821acd85c135837

Observation 549af06b-7b60-4c27-9aa4-e09d009bb627 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:04.407246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:04.407246Z digest=sha256:9902384b072b0d03be12448c082af30b93c58b70aab02df8391680ac4cc45b31

Observation 5de09fb8-0c6e-4ec0-b728-13a0d27c0ed1 · outbound

This paper cites Zitkovich, T.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Zitkovich, T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:43:05.434319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:43:04.537493Z digest=sha256:12c58d151216e03ab0274ea83cd3012429a03392b0fb221ef9afd4dc6ffc6905

Pith citing papers

Observation 27e0a52e-db45-4355-a05b-7d132211f3f4 · inbound

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models cites this paper.

RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:14.854316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:14.854316Z digest=sha256:2ab39c1df141f7827556c3229d4cd02b9ff6abdd7708b42c1d514a1364ed2875

Observation e9ad02eb-0de1-4f98-89bb-7bfaae0e40a2 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:28:16.251395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:5ed8417ded7b6316ea06940dd0ac43a1c792c488ecaac3b5a7586afc86c0bd5a

Observation 5cc907b8-c95a-409f-af39-091b8910feed · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:12.680018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:12.680018Z digest=sha256:e676849e7e7cde942bc565f3e097ea4cedf0d3017d28420ff3f5fc1f831a96e5

Observation c7669fec-08b0-4352-868d-3ff350b648b3 · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.433198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:848ca3cc57a05b8272c22cea5d8dcb5631bb58c03e0a6195725c61dfed284090

Observation d75ee6f8-1593-40a5-aff7-9f64f760f4ba · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:39.309018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:39.309018Z digest=sha256:e455a6fc82e0479a11693891976b172097f9a0a24511900e050c6e780000d525

Observation cdeed0c5-557e-4128-8df2-49523da42a74 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:f9eeace023790bd6a64a8b0c47b1063aff5380a3dd9b3fc5b690619ddc55be74

Observation 566736a1-5b5a-44d2-9456-127d74258f25 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:50:03.934133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:f9df77a8d79cce1b42fafd5f2f208c12506ebf895e959dd8f0ae0f9ca99c97a4

Observation ebda3e5c-1bb5-4920-babb-7c268ddb6849 · inbound

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success cites this paper.

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:50.142804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:33:26.708966Z digest=sha256:fd30a8938f1e3e6604bc7e677d50fbda7dd5b87697fc8f51f6be0a940020d649

Observation a65763df-9030-4369-a569-943e7c27f4f5 · inbound

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models cites this paper.

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:22:06.687892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T02:19:39.604140Z digest=sha256:210403850fca93d983fd4bd0f9731181508fbcc219ee889e3caf7ab0f52980c8

Observation 2a3220a9-f94f-44ad-8e8d-4501e7d6aa96 · inbound

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models cites this paper.

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:15:07.104557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:13:12.442859Z digest=sha256:3d60adb8c4fb89e44c622854ef65ba36cf536d924a43916402602b53a43528fa

Observation 346d9813-74f7-4774-a93f-75f41c53f5bf · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:47:37.220705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T18:43:08.029165Z digest=sha256:6a6bb02f794cb7122d3bb99bc01037e0ddaef5f3fb5047866188ef4db05e1a15

Observation 012a3ec7-7c05-4075-a2b5-c77670f3a936 · inbound

AttenA+: Rectifying Action Inequality in Robotic Foundation Models cites this paper.

AttenA+: Rectifying Action Inequality in Robotic Foundation Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:45.902573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T21:42:05.592911Z digest=sha256:a750b91b39fdd16d5e2288e6f71ba78611d1680012deecff1a440c9adb906ed0

Observation 207fda53-dc32-48a8-a34f-e553e199d3a8 · inbound

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models cites this paper.

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.041269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T07:16:27.229603Z digest=sha256:199afee3a904d8c25ed68fb8e71af82a6d4795a8a6838fcd0ed35acd54e8d036

Observation 43146e1b-4c59-4bc5-a8d3-d2e9ed35e07b · inbound

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models cites this paper.

GeoSem-WAM: Geometry- and Semantic-Aware World Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.619734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:11:10.246045Z digest=sha256:15110cb2a4cea5c9dedd6b0037e52cb951e5d98018df5ed3ac5c162e6a4d3ddd

Observation 10a0d5dd-2dfe-427d-b614-7934d9e4435a · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:47.661540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:2b8f45a33b6233e8f9d37c6eb051458cf6292bd525e1b4102f03b7ee4c371ff3

Observation 8bdac25c-b34c-4d14-943d-f0123da50301 · inbound

Improving Robotic Imitation Learning via Trajectory Standardization cites this paper.

Improving Robotic Imitation Learning via Trajectory Standardization Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.564349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T08:37:26.864436Z digest=sha256:f01794a4a8088912f5ffbf9079d88b551cbe2971ca3850ef0aeb617104579056

Observation c89b7deb-90ab-4cc3-b7c5-789c8b0edc79 · inbound

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? cites this paper.

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:52.300386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T04:56:00.919208Z digest=sha256:6d7b2ba2fde33a2ff3fea7e8883f99cf15d078a3d46983947040e38ebc5a632d

Observation 6ee57c48-dd03-412e-ba93-5ae81aa47fbb · inbound

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation cites this paper.

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.864902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T05:07:45.792132Z digest=sha256:7c4ab01fe03274cb96ca5e4b9dec19b7e1f63ba11a8afcf7fff66b0972822481

Observation 9426d060-5cdb-4d25-b5cc-493c87fc90f6 · inbound

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving cites this paper.

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:09.074689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:26:09.074689Z digest=sha256:b3748ba11baf1cc57d5ac7645eb53c94cad19064c116c6e16500b498f2dc01b8