Pith. sign in

Paper Citation Record · LEDGER

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

As of 7 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2608.03649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03649 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:19:02.848690Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9ac022f-09d5-4619-ada0-72b387edb7e8 · outbound

This paper cites Qwen2.5-VL Technical Report.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.785738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.785738Z digest=sha256:a95ea34b9e80b890e16f7276a5fc1dbac17b8a11ef9da7ed6ff7d19d9f3e7dee

Observation 12a72b7e-cfe8-4a6d-a7f5-4f4b45bf6057 · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.790931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.790931Z digest=sha256:0368e4a615e7f6ae60d702610480f41c8e418c5a98cb1e643ff4004cc8b5af61

Observation c58f92f9-2a10-4044-b4dd-ae0044801909 · outbound

This paper cites CARES : Context-aware resolution selector for VLM s.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware CARES : Context-aware resolution selector for VLM s

Reference 3

Resolution
verified exact
doi, observed 2026-08-05T15:19:03.166692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T15:19:02.799318Z digest=sha256:300df74d92126042c8054bc011697b74a5abb033235b6ce1d8415dbe97b8b3ae

Observation a842a616-3f4c-4e81-9abb-82e1759158ba · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.805044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.805044Z digest=sha256:b15a4116bde40363f353f576a42c33c573ffba284fbdc617fb6d4e9d936a5086

Observation 8b1fc09e-7b6b-41d9-9f83-aa2e0de2ad58 · outbound

This paper cites ResAdapt : Adaptive resolution for efficient multimodal reasoning.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware ResAdapt : Adaptive resolution for efficient multimodal reasoning

Reference 5

Resolution
verified exact
doi, observed 2026-08-05T15:19:03.130483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T15:19:02.810877Z digest=sha256:17d0de3d3079efa1d120eedcbca4f723a6cd67fe2a4bbb15654aaac2f72890d9

Observation cbff600b-51cf-4dc4-b64e-65e8b4e056c8 · outbound

This paper cites AdaptVision : Efficient vision-language models via adaptive visual acquisition.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware AdaptVision : Efficient vision-language models via adaptive visual acquisition

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:03.342589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T15:19:02.815910Z digest=sha256:dd23473b1bfca1820539906415ae5986a60a2a75b8e078b9b4f4f1c79e88d5c7

Observation 4b7b91bc-e4f1-478f-b1b8-e9ffb00a8c82 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.822062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.822062Z digest=sha256:f961b8b6dbd1038ee278da974a838cd2c1b647cc298249797fcedad9d048f1dd

Observation 477b4d13-068d-49db-ad08-8707a305e08d · outbound

This paper cites LLaVA-PruMerge : Adaptive token reduction for efficient large multimodal models.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware LLaVA-PruMerge : Adaptive token reduction for efficient large multimodal models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.827885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.827885Z digest=sha256:c28326615a740d7439d0328978a6e894c2797e25620635052be519972d8d5dfd

Observation 40a67815-76bc-4e7b-9c45-59bb231065ae · outbound

This paper cites Towards VQA models that can read.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Towards VQA models that can read

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.833221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.833221Z digest=sha256:e2d3262c18e1558ff1104b5ca880747252a55ce69707101ba153bbb799e200df

Observation 03fcfec7-111b-4c18-8d98-5ee76158c127 · outbound

This paper cites Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.837652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.837652Z digest=sha256:dda5f83433768e9b96d56fc5e72d24b3766745ba9502318eb5ffc0738250b12f

Observation b8e56e5f-2d4b-4c0a-ba66-b94aabab9a82 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.843428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.843428Z digest=sha256:449de116d7ccf4febfd6a465e2bd08782824abd1dc6dcd0d041186a71f1eeaa9

Observation 4a8e9594-1b6f-4d93-94d8-584de078c6d5 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.848690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.848690Z digest=sha256:f1bd107c491224d47044229eabd39af976afc28dbcff8e84ba3d553d299a7e56

Pith citing papers

No inbound Pith citation observations are available.