Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:19:02.848690Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2608.03649.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:19:02.848690Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a9ac022f-09d5-4619-ada0-72b387edb7e8 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a72b7e-cfe8-4a6d-a7f5-4f4b45bf6057 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c58f92f9-2a10-4044-b4dd-ae0044801909 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware CARES : Context-aware resolution selector for VLM s
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a842a616-3f4c-4e81-9abb-82e1759158ba · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b1fc09e-7b6b-41d9-9f83-aa2e0de2ad58 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware ResAdapt : Adaptive resolution for efficient multimodal reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbff600b-51cf-4dc4-b64e-65e8b4e056c8 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware AdaptVision : Efficient vision-language models via adaptive visual acquisition
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b7b91bc-e4f1-478f-b1b8-e9ffb00a8c82 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 477b4d13-068d-49db-ad08-8707a305e08d · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware LLaVA-PruMerge : Adaptive token reduction for efficient large multimodal models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40a67815-76bc-4e7b-9c45-59bb231065ae · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Towards VQA models that can read
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03fcfec7-111b-4c18-8d98-5ee76158c127 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e56e5f-2d4b-4c0a-ba66-b94aabab9a82 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a8e9594-1b6f-4d93-94d8-584de078c6d5 · outbound
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.