Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:24:59.159037Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2606.29350.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:24:59.159037Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dfa6a222-4c88-42a1-b377-6bdb8a55bb5e · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs The Llama 3 Herd of Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8dd48735-492b-43e9-ab99-d003ff7e30dd · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 36a530f2-00c9-4206-99d1-53172469ba99 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cddc2954-e1c9-4168-aa06-11ed219cb9fa · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a485c24c-cd2b-43d6-89fc-803775ff22c5 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Qwen3 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4d214fce-f917-4415-bbd7-0676650c77ae · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs LLaVA-OneVision: Easy Visual Task Transfer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8eed4ad4-d269-46e2-8a6d-def0eb2588bd · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Impedancegpt: Vlm-driven impedance control of swarm of mini-drones for intelligent navigation in dynamic environment,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d74d197-929c-4f4f-b5a8-ebe1865af5ee · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Rod-vlm: A framework of real-time robotic perception, reasoning and manipulation,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f196455a-34bc-402e-8edd-3e933462884b · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs On the safety concerns of deploying llms/vlms in robotics: Highlighting the risks and vulnerabilities,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 679fb846-7f80-4646-a552-f5bf90cc467d · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs OpenVLA: An Open-Source Vision-Language-Action Model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b985be5a-df5f-464f-8385-fcc934e8f8fa · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df498cf9-7197-4728-acf5-237dc11be59e · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Diffusion policy: Visuomotor policy learning via action diffusion,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d86d1c-a1b4-44a3-a713-47a48858a4f7 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Video Token Merging for Long-form Video Understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2887f1e4-bb4c-4ed7-a3f7-6017f8d40248 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs PuMer: Pruning and Merging Tokens for Efficient Vision Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d0289b44-1d3a-46a9-8a8d-6673629f30e3 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Token Merging: Your ViT But Faster
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94d41bf3-1808-4809-8b85-88f15ee10a97 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Boosting multimodal large language models with visual tokens withdrawal for rapid inference,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ad8726-5061-49d9-9931-32ef63431cf8 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs An image is worth 1/2 tokens after layer 2: Plug- and-play inference acceleration for large vision-language mod- els,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694edf93-7b38-4903-9800-dee449e9f99c · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Framefusion: Combining similarity and importance for video token reduction on large vision language models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d57363d-6ea0-4bef-ac28-0fac05256c08 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Libero: Benchmarking knowledge transfer for lifelong robot learning,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7589ee9a-6b38-4a4c-9a9f-ae5e7ff2e5ca · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Visual instruction tuning,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f609e80c-90cb-48e1-b0ed-e518164bf6b1 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs An advanced driving agent with the multimodal large language model for autonomous vehicles,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e350805a-9678-45df-8f28-6ada83bf8da5 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Cliport: What and where pathways for robotic manipulation,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9af89bed-4449-4cba-a163-110c6480b37a · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Vima: General robot manipulation with multimodal prompts,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da6386a-2f46-444d-8228-249548131f3c · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs PaLM-E: An Embodied Multimodal Language Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dfcfe20a-1bea-43c1-b89b-a4d09fdd5557 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Robospatial: Teaching spatial understanding to 2d and 3d vision-language models for robotics,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3606b2e-4ec4-4c78-8724-030d0242b4c4 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Physvlm: Enabling visual language models to understand robotic physical reachability,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f57ac76-391f-4f57-80c1-f0f548b67f96 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d45dbb1-deca-423c-ab4f-72ffd7a2537e · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d89824cc-af80-446b-9624-1b0aad4f29e1 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Drivelm: Driving with graph visual question answering,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc8220e-365c-4d18-9a12-bfe228ac0780 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs TR-DQ: Time-Rotation Diffusion Quantization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1943de30-9f1e-4531-82f7-11f2f0aaf79b · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Learning to Merge Tokens in Vision Transformers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2cf0d65-e021-420f-91f0-187a65f4a9f4 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Learned token pruning for transformers,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8156aec3-1cac-436e-b95a-8a6e7177745b · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Exploring token pruning in vision state space models,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba041b35-a96a-45a0-82d9-820d8b7e45b5 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Topv: Compatible token pruning with inference time optimization for fast and low- memory multimodal vision language model,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d3ecb4-de88-436a-9037-80c60653df10 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Flashat- tention: Fast and memory-efficient exact attention with io- awareness,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8e5f54-191c-4a60-b11f-cd5b14913f5f · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9f0e35cc-af81-47e0-aae5-3d857b5bd176 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Rotary position embedding for vision transformer,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d7fb34-5c93-4fd4-bd81-52d1de2fb150 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Llava-med: Training a large language-and-vision assistant for biomedicine in one day,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978f05a2-2528-4dbe-a1c6-ecdf976c33fb · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Internvl: Scaling up vision foundation models and aligning for generic visual- linguistic tasks,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3e28a1-7ced-4dd7-8430-ea9ba2297cde · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Roformer: Enhanced transformer with rotary position em- bedding,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405e9bd1-0db9-4bfc-8bf8-6a6e5c2d1341 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 13093a61-3d58-48db-bbe6-6d046b56c1a0 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs lerobot_ π0.5_base,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fcc7b26-bdb0-4e32-a349-3a8cab0f3b41 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs A shortcut-aware video-qa benchmark for physical understanding via minimal video pairs,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d026fb4-d253-4ab5-aecb-16b2f679ee76 · outbound
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Lerobot: State-of-the-art machine learning for real-world robotics in pytorch,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.