Pith. sign in

Paper Citation Record · LEDGER

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference

As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.16326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16326 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:42:16.353153Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22e10980-24c3-4109-b9ec-fe6d9a67ee10 · outbound

This paper cites A survey of state of the art large vision language models: Benchmark evaluations and challenges,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference A survey of state of the art large vision language models: Benchmark evaluations and challenges,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:13.962870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:13.962870Z digest=sha256:9471b62507868b02da0b0ac2090eefad1c265ad44f65285f321e174920e40423

Observation a2e7303d-8d63-4d0b-a2ec-b2d66a03a22c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.029230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.029230Z digest=sha256:a28751dd666a5985d233c8886d8c0f712d0badaaa008556939b629b9ab28da7b

Observation ffc841a0-501a-4e7a-8a12-df7467060223 · outbound

This paper cites [cls] attention is all you need for training-free visual token pruning: Make vlm inference faster,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference [cls] attention is all you need for training-free visual token pruning: Make vlm inference faster,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.088093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.088093Z digest=sha256:2f4ec7aee33ab8b1e8b081dbcf315eade76640c96eb1927a9e78fd8e07e5fcf5

Observation cb80ec59-6fda-4a39-a281-9eb459c21c3f · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.167493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.167493Z digest=sha256:6bb77f8c524dc8793ee228ac0b2f724398a5b0552b281ecd3fe66f469e3e88af

Observation fcb18766-f096-4f16-9ae0-344f2ddf65fd · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Visionzip: Longer is better but not necessary in vision language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.225585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.225585Z digest=sha256:847d605a8b39005d2fa1ad98c8b13cac1c0b1a41e1b7476074d3e19db62e7244

Observation a3f840ac-ea02-43e2-9cf5-8fcc758c8011 · outbound

This paper cites Divprune: Diversity- based visual token pruning for large multimodal models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Divprune: Diversity- based visual token pruning for large multimodal models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.268413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.268413Z digest=sha256:497b1ceb58653f2bf4c424c54b6a2747e8839b83f8521e5b3c363519621d84cd

Observation b2747e3b-6ba8-4f41-84f1-43088a6a4d92 · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.328001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.328001Z digest=sha256:b95a2124103d9d60ac8a9a23ecbc128b6ba4d2126c83070a0bf92356f6480f0c

Observation deb8fc81-c2ea-47af-a648-ca7099b13bd2 · outbound

This paper cites Which experimental design is better suited for vqa tasks?: Eye tracking study on cognitive load, performance, and gaze allocations,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Which experimental design is better suited for vqa tasks?: Eye tracking study on cognitive load, performance, and gaze allocations,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.387140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.387140Z digest=sha256:c6f52a61657d9ff559609895b959145196ac0040a6444074802e603792bf867b

Observation cf119c7b-0c1e-4101-8c49-904514bb2413 · outbound

This paper cites Making the invisible visible: Verbal but not visual cues enhance visual detection,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Making the invisible visible: Verbal but not visual cues enhance visual detection,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.451408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.451408Z digest=sha256:f0ae869907b59cab1700b9d4147ab71ec9d80cdc1f2b9fcc6dae0efb5ae82e81

Observation 84348696-6677-4790-919f-ea30c9956231 · outbound

This paper cites Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.510468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.510468Z digest=sha256:f88287fc40cccce52e7b35f9579d7185d36d07884431d5d89407ec016bb92d06

Observation afcfed36-0ac4-44ac-8e3d-7e56bab9427b · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.568395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.568395Z digest=sha256:7cee76c7df88262346af1748397a3ee6f835f3d20cd31f8862a3551b2ed273e9

Observation a87a12c8-fe66-4e61-b471-5a5b9fb3f6f8 · outbound

This paper cites Improved baselines with visual instruction tuning,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Improved baselines with visual instruction tuning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.614845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.614845Z digest=sha256:70820c5b5cd0c2c9549bb5fbf2a2b8bfb76046a6de09b1ad4f167a5439c7db4b

Observation 0f349ea5-b27c-4a40-aabb-1980a500fb4f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.672964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.672964Z digest=sha256:fac227765fbe8dc238abaf0452fd22a268b97e409b0eb9164347d557b15cbc7c

Observation 4b78eb07-ce38-47e4-906f-82d0d86ed7e6 · outbound

This paper cites A-vl: Adaptive attention for large vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference A-vl: Adaptive attention for large vision-language models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.747206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.747206Z digest=sha256:c3e7097a2c875f9d022a36499b404f9ad945923d70645be77a4249acb7bcd33c

Observation 43bb17fe-eb23-4fad-ba9e-8ac6e71a801b · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Llava-prumerge: Adaptive token reduction for efficient large multimodal models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.810397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.810397Z digest=sha256:14208c570150fc58d853d79043cae26b1856dd2cb60235af2abc88c5b4383917

Observation a6c1fd54-da0e-4dc4-85b6-278d6826381c · outbound

This paper cites Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.869453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.869453Z digest=sha256:22e712b5f59694271324b3bfc80e7081ce50d1b0562131fdf43b7875a9f4791f

Observation 7e8570fb-3b5f-477d-a0d0-4f7da959d9b4 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.917470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.917470Z digest=sha256:1235a77ddc1519bda56f64dec830d65e287d5159816778f90a2be83bb37a2103

Observation e70aee6f-94b2-4f0e-893a-450614b14839 · outbound

This paper cites Token merging: Your vit but faster,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Token merging: Your vit but faster,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.968369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.968369Z digest=sha256:e89d105b3bd1b969a044fca936ed368a4039fe0f14b8eac59cb45aa8450b14f1

Observation 9be3d9bc-25e7-4233-9703-6ef794af667e · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Less is more: A simple yet effective token reduction method for efficient multi-modal llms,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.973153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.973153Z digest=sha256:7e4dd2a27a547d8143c2d1217e12b50d4a90058e2e35fd7ece805a2ab2d769d8

Observation 2cac17e1-2c7a-41fb-9ea6-fe26463ff5d8 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:14.989028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:14.989028Z digest=sha256:77100c73862f9cc2397c339c4e26a5ee55ee60a2a8e0d926ae8e5cfd2e5a3ffa

Observation eb6895fa-9c7e-47b1-948a-5c12c1dfccec · outbound

This paper cites en core web sm: spacy english small model,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference en core web sm: spacy english small model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.064940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.064940Z digest=sha256:0fb6ed278b3052daa95e2cce8114581af55dd08e0b595453e6f245b3537eb63a

Observation aa5edcd7-3e17-4763-95cc-a265e3619c42 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.224141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.224141Z digest=sha256:c121056b28ef76b4c310c575c09c51e3b0bb58e64661e228226cfaeb1f7ba50c

Observation 6ee3588a-2ed4-48ec-a572-b6008311a5bb · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Vizwiz grand challenge: Answering visual questions from blind people,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.385887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.385887Z digest=sha256:ce70697abd642b706f303a10a048097c29eb3b43209bdce4e3f3228b1c14e602

Observation 0b291d16-1ac2-4445-9add-45ac222e2926 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Gqa: A new dataset for real-world visual reasoning and compositional question answering,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.538454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.538454Z digest=sha256:dab770cb6ca97d2dbc7141b49bfd85b9f2e4929a6329e4a92df891f9c79fd6e8

Observation e23e9de6-19a2-423b-9be8-701099206de4 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Learn to explain: Multimodal reasoning via thought chains for science question answering,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.652552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.652552Z digest=sha256:e9bbadbbc81ecda348b9e9e88980bf3bfc5a1fbf9608c7416afe1333a21aa8f5

Observation 6ac73fdb-ead5-4f4c-a603-1df1de942469 · outbound

This paper cites Towards vqa models that can read,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Towards vqa models that can read,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.811440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.811440Z digest=sha256:f2b4e2e18fbf9ba81ef0c60344fc8ea96bcaea641701560e3507a0e9ea00c2ec

Observation a0dc2b26-e124-494d-8b8a-f0c40c015cc9 · outbound

This paper cites Evaluating object hallucination in large vision-language models,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Evaluating object hallucination in large vision-language models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:15.972178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:15.972178Z digest=sha256:a256677a05b5bf4ef2abe13e02fd6fe9fbd312e7e5f57c6409672a4b564a157e

Observation cf4a3d85-1401-493d-9782-227dffb6d8ee · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:16.140370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:16.140370Z digest=sha256:96222235edd0fefd65867040202f9af7a6a6e7ce3deda95f2743c35318acda3b

Observation 9ccc400f-956c-4576-9587-b505246e1706 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Mmbench: Is your multi-modal model an all-around player?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:16.307600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:16.307600Z digest=sha256:ec61b19a31f511db9777930372b15d9226cfeb65d4ba4b34fc461a9be564beb8

Observation 6462f743-c873-4b9f-8e30-c2f19f1637fd · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities,.

CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Mm-vet: evaluating large multimodal models for integrated capabilities,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T03:42:16.353153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:42:16.353153Z digest=sha256:b63220586bfa5dab286da04f3e9c1ec044f78acf4144bb22317e80a0f1af1721

Pith citing papers

No inbound Pith citation observations are available.