Pith. sign in

Paper Citation Record · LEDGER

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs

As of 17 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2606.29350.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29350 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:24:59.159037Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfa6a222-4c88-42a1-b377-6bdb8a55bb5e · outbound

This paper cites The Llama 3 Herd of Models.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs The Llama 3 Herd of Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.124779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:4b648efabec2f61eaf4df8a1220ae07889ca603bedd7652e34d8795cb990de3c

Observation 8dd48735-492b-43e9-ab99-d003ff7e30dd · outbound

This paper cites GPT-4 Technical Report.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.130859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:4b0c84452390f479db0d1e18a29578054e902e297e3755ba02c3482afa886295

Observation 36a530f2-00c9-4206-99d1-53172469ba99 · outbound

This paper cites DeepSeek-V3 Technical Report.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs DeepSeek-V3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.117016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:05946c5f28581dd2909cfd7a2ef241bbfaa822b1616315c7d53626d7e53e42d0

Observation cddc2954-e1c9-4168-aa06-11ed219cb9fa · outbound

This paper cites Qwen2.5-VL Technical Report.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.119869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:9866316a703dd28f61dde97772d0c832bd0fd603dd54609b40c83291b7c1c44a

Observation a485c24c-cd2b-43d6-89fc-803775ff22c5 · outbound

This paper cites Qwen3 Technical Report.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Qwen3 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.122440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:0ca4e6ecf47772df3598ee5f7d59b6e7b47b44f38e123468a9a2bb6dc8e8e1b6

Observation 4d214fce-f917-4415-bbd7-0676650c77ae · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.109277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:d9ef2f87a56a93e61f825362978af4d303fa2db8cab18d9c9ce08f52feb5d31c

Observation 8eed4ad4-d269-46e2-8a6d-def0eb2588bd · outbound

This paper cites Impedancegpt: Vlm-driven impedance control of swarm of mini-drones for intelligent navigation in dynamic environment,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Impedancegpt: Vlm-driven impedance control of swarm of mini-drones for intelligent navigation in dynamic environment,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:8b4e466b72c51dfd96de8e2b219063c89f3dedbce8514aab80f15cc6add9842d

Observation 1d74d197-929c-4f4f-b5a8-ebe1865af5ee · outbound

This paper cites Rod-vlm: A framework of real-time robotic perception, reasoning and manipulation,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Rod-vlm: A framework of real-time robotic perception, reasoning and manipulation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:fe77cd0464dcd1882f0a21985b0acda107b83666dbafc3abb7898da8e21d5428

Observation f196455a-34bc-402e-8edd-3e933462884b · outbound

This paper cites On the safety concerns of deploying llms/vlms in robotics: Highlighting the risks and vulnerabilities,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs On the safety concerns of deploying llms/vlms in robotics: Highlighting the risks and vulnerabilities,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:e5723b2b53a0ccf2d575e50b9a799aaf134568096860e92f80d2045ca19dfa9b

Observation 679fb846-7f80-4646-a552-f5bf90cc467d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs OpenVLA: An Open-Source Vision-Language-Action Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.099014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:a006dfd2dcfe9c529fbff4155158c971403720ea314bb690d03801dc192bb2ba

Observation b985be5a-df5f-464f-8385-fcc934e8f8fa · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.117350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:98b3727323cad43e9a2221bcb765f1e0d2c6d3591addfd8f6b2964ccc1dc4bb8

Observation df498cf9-7197-4728-acf5-237dc11be59e · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:7004233f792365f708c248afa9b0e2e8d5385584075a91ff3a8234b15be91558

Observation 48d86d1c-a1b4-44a3-a713-47a48858a4f7 · outbound

This paper cites Video Token Merging for Long-form Video Understanding.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Video Token Merging for Long-form Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.136566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:9457a129f40f5f49fb68d53aa36e0381a82a8cd89e4feb776f15efb700c72c39

Observation 2887f1e4-bb4c-4ed7-a3f7-6017f8d40248 · outbound

This paper cites PuMer: Pruning and Merging Tokens for Efficient Vision Language Models.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:34:22.100839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:fe71873f7e6a4f52aeaeef6f4583d01f001d12a40d58bf775015f2314b5f2cd6

Observation d0289b44-1d3a-46a9-8a8d-6673629f30e3 · outbound

This paper cites Token Merging: Your ViT But Faster.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Token Merging: Your ViT But Faster

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.122264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:d5c0efc9737b7614d9693da25b0694a589c711b07e0e538b8ed347ac9b95865f

Observation 94d41bf3-1808-4809-8b85-88f15ee10a97 · outbound

This paper cites Boosting multimodal large language models with visual tokens withdrawal for rapid inference,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Boosting multimodal large language models with visual tokens withdrawal for rapid inference,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:4600d007bea624e8d9001ad3dddae3c8037b092acf5402e66227b6d0949a9320

Observation 54ad8726-5061-49d9-9931-32ef63431cf8 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug- and-play inference acceleration for large vision-language mod- els,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs An image is worth 1/2 tokens after layer 2: Plug- and-play inference acceleration for large vision-language mod- els,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:44ad0e9024c65348018f2e0f7eeee6a4d89656dce0ceb01c3f28b7359a953254

Observation 694edf93-7b38-4903-9800-dee449e9f99c · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large vision language models,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Framefusion: Combining similarity and importance for video token reduction on large vision language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:0649844d6eceb956330948bf246f53d9fc0df058f92696567c8783e69049831b

Observation 4d57363d-6ea0-4bef-ac28-0fac05256c08 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Libero: Benchmarking knowledge transfer for lifelong robot learning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:baaff78caf07919fc73ad1d51708a68386d4fe2b87fc2db07f24a342a946b4a1

Observation 7589ee9a-6b38-4a4c-9a9f-ae5e7ff2e5ca · outbound

This paper cites Visual instruction tuning,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Visual instruction tuning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:43be2d38fb24c0151a14e259d957a7911cb4a117409ffe16ea89f6185c752d35

Observation f609e80c-90cb-48e1-b0ed-e518164bf6b1 · outbound

This paper cites An advanced driving agent with the multimodal large language model for autonomous vehicles,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs An advanced driving agent with the multimodal large language model for autonomous vehicles,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:d4831abe7e2e65c52ca9495a3877f428b3e270c0a839b17200d877722161713f

Observation e350805a-9678-45df-8f28-6ada83bf8da5 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Cliport: What and where pathways for robotic manipulation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:fd5f1f4c556cd7cf6488945f5a57c7061571d8381ac4d76535adcf67ac61d215

Observation 9af89bed-4449-4cba-a163-110c6480b37a · outbound

This paper cites Vima: General robot manipulation with multimodal prompts,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Vima: General robot manipulation with multimodal prompts,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:977fcbf02df8ab27fa246a972e960bcf374b36540e573d06db35f104da701bb8

Observation 5da6386a-2f46-444d-8228-249548131f3c · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs PaLM-E: An Embodied Multimodal Language Model

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.127845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:a6c773c70a5aca90160cc01788e7102bc71f5e1c3ce34405532ae3b4f2e47b8b

Observation dfcfe20a-1bea-43c1-b89b-a4d09fdd5557 · outbound

This paper cites Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:fe6313d7551ce978a3affa6df00a8734f8aad54808c4eb36ce0f15643eb3662a

Observation d3606b2e-4ec4-4c78-8724-030d0242b4c4 · outbound

This paper cites Physvlm: Enabling visual language models to understand robotic physical reachability,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Physvlm: Enabling visual language models to understand robotic physical reachability,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:9c66804813a5c13d39a90fd6135f274e8f45e1cc337f9769eb7e2e4fca642513

Observation 7f57ac76-391f-4f57-80c1-f0f548b67f96 · outbound

This paper cites VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.139716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:32eaf329b58711c7963bdf0b3e2342d3dc9b200cd9ac0f53222b94d642dc0c62

Observation 3d45dbb1-deca-423c-ab4f-72ffd7a2537e · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-30T07:34:22.124969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:2107cbfd566526bc9c6eb6a95ab7ab755318da5d151d70176a80656e6bb49c1a

Observation d89824cc-af80-446b-9624-1b0aad4f29e1 · outbound

This paper cites Drivelm: Driving with graph visual question answering,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Drivelm: Driving with graph visual question answering,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:f8d18926959c0fe1873d943087ff0f50eb569f076dedd3e4c405e017cb0b6001

Observation 1fc8220e-365c-4d18-9a12-bfe228ac0780 · outbound

This paper cites TR-DQ: Time-Rotation Diffusion Quantization.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs TR-DQ: Time-Rotation Diffusion Quantization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.106711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:bc51cac9809b05a023368844efdeb58033cd50f1b5409f92dfcf13cfead4d7e5

Observation 1943de30-9f1e-4531-82f7-11f2f0aaf79b · outbound

This paper cites Learning to Merge Tokens in Vision Transformers.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Learning to Merge Tokens in Vision Transformers

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.142467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:b87486fcbace5270dd0a98b289c2bd2fda2d99714e454ce392e750aaf7f9f9ed

Observation d2cf0d65-e021-420f-91f0-187a65f4a9f4 · outbound

This paper cites Learned token pruning for transformers,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Learned token pruning for transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:5814feb6c203059ac53abe78e84303d81a4b91f769b89b9f3083f669653ba089

Observation 8156aec3-1cac-436e-b95a-8a6e7177745b · outbound

This paper cites Exploring token pruning in vision state space models,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Exploring token pruning in vision state space models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:c4f5b4f2e09b6064a79bad6449dddecbda1bf2d37f65e1557bc92b79f36d1653

Observation ba041b35-a96a-45a0-82d9-820d8b7e45b5 · outbound

This paper cites Topv: Compatible token pruning with inference time optimization for fast and low- memory multimodal vision language model,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Topv: Compatible token pruning with inference time optimization for fast and low- memory multimodal vision language model,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:9fa270bddf92af21b8e42183a75f332bfa272738b85c7318d6ee85974cc2ac1a

Observation b0d3ecb4-de88-436a-9037-80c60653df10 · outbound

This paper cites Flashat- tention: Fast and memory-efficient exact attention with io- awareness,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Flashat- tention: Fast and memory-efficient exact attention with io- awareness,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:e099ea2131949a78d8b162266026b6a721303ac84bc78cb385ce39e64eef7c1f

Observation de8e5f54-191c-4a60-b11f-cd5b14913f5f · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.133596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:896b30f5e3428207e2719abd894375d20c8952e40d3b0a431b03df1a2b8beedf

Observation 9f0e35cc-af81-47e0-aae5-3d857b5bd176 · outbound

This paper cites Rotary position embedding for vision transformer,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Rotary position embedding for vision transformer,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:583e88676935186a1a12d72cd0c8c30011301067deee87f2a883a361d90e41a5

Observation b9d7fb34-5c93-4fd4-bd81-52d1de2fb150 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:c9398a5760df1ec5bf7ed99f5fb425406147e2350536bc5fe85f264b418f77a6

Observation 978f05a2-2528-4dbe-a1c6-ecdf976c33fb · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual- linguistic tasks,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Internvl: Scaling up vision foundation models and aligning for generic visual- linguistic tasks,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:50b4970ee633dcfb0586c815a56406818aa863a7040f47e5a2b88556a42c79cb

Observation 9a3e28a1-7ced-4dd7-8430-ea9ba2297cde · outbound

This paper cites Roformer: Enhanced transformer with rotary position em- bedding,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Roformer: Enhanced transformer with rotary position em- bedding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:6cd96477b0d9a2f7a5d558bfccaf28a0d68c206618b93a970535b47080dadc01

Observation 405e9bd1-0db9-4bfc-8bf8-6a6e5c2d1341 · outbound

This paper cites TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:22.145236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:cfa60705fb7100ca921129ceff0b7828d7d4ca74821c9752c43e00102034bdaf

Observation 13093a61-3d58-48db-bbe6-6d046b56c1a0 · outbound

This paper cites lerobot_ π0.5_base,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs lerobot_ π0.5_base,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:3f52ea03b0d256e917b2b27c70888040d801bb9c9035503f9c6285c088052edb

Observation 5fcc7b26-bdb0-4e32-a349-3a8cab0f3b41 · outbound

This paper cites A shortcut-aware video-qa benchmark for physical understanding via minimal video pairs,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs A shortcut-aware video-qa benchmark for physical understanding via minimal video pairs,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:7bf3e2af774b309a1abe3e904a224ebbe04315cb761585a15ce2aecb4b646957

Observation 9d026fb4-d253-4ab5-aecb-16b2f679ee76 · outbound

This paper cites Lerobot: State-of-the-art machine learning for real-world robotics in pytorch,.

Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Lerobot: State-of-the-art machine learning for real-world robotics in pytorch,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-30T07:24:59.159037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T07:24:59.159037Z digest=sha256:3b5717938548a5c730c820c59f7fbb7846b162404957d9ab222f81459321f79d

Pith citing papers

No inbound Pith citation observations are available.