Pith. sign in

Paper Citation Record · LEDGER

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2412.01818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01818 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.534196Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:26.122002Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d590f45-b32e-4d8f-9566-7aa62d9e1872 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.826191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:f7e4e02a6dfb33927841fb05ec5e2c165380ae957745512c50ca0701efa5c918

Observation a152e2cf-71b3-4838-ae5a-24f22fbedff1 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:10.568820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:10.568820Z digest=sha256:f4d60198e39fd0be87d51738dece77b944ff487ddb980f5c9e35e0e5f5bbe62d

Observation dfa4e81f-0ea0-49af-aad9-41a331e93e78 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.721952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.721952Z digest=sha256:d3c3433b8fa13ab06e89e19277d6997141bcc4f646773b90c7ba3314d81752ca

Observation 71ec673b-1526-4118-a103-06c1a1fd7bbf · inbound

Do Concept Replacement Techniques Really Erase Unacceptable Concepts? cites this paper.

Do Concept Replacement Techniques Really Erase Unacceptable Concepts? Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:22.689471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:22.689471Z digest=sha256:0f5707714e289decb3d7672ea75fe0645105cdcd661bc29c7dc88ce906f5f744

Observation a1a373ab-ce2e-4b1e-8eac-7867ae7a5305 · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.731513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.731513Z digest=sha256:a0fa7bd57902e92ecfa7cca91e7a7ec374bee6d12e683f3821cd462eb302da34

Observation de3b0a21-a37f-471b-b485-d88fd6c34fe7 · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.833498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.833498Z digest=sha256:f787cbce0010e826222d9d720ea0a529e4d27f967db17883e57bc9c656fc7e0b

Observation 027bcc54-980e-4005-976e-4024bede4f92 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.069967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.069967Z digest=sha256:a8c168f48c8fc7808755d00a483aa7593f0cffed8e797d32bb11e81be235b6c1

Observation 9769bad8-5242-44f0-8796-3fb811209059 · inbound

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning cites this paper.

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.007759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.007759Z digest=sha256:da369f0af794d307c2450061aa7ba1305fd884e7a04f362b8a97b656e5db72f3

Observation db89bafa-0a57-4cdb-9ed4-0046b86f872e · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.989156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.989156Z digest=sha256:698229b9103ff427f8e971f91ad52a5045ba75a18cea9b45adb6ab0256972734

Observation 6e34a4b6-98fa-45e8-929f-7b2dce7e4b12 · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.501635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.501635Z digest=sha256:50fe3952d0a9800f6fa363137ae2f00409c3d8afd4202eb246b54b1788d5a1fb

Observation 18328f71-a4db-42db-a401-0e22267c7c09 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.802359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.802359Z digest=sha256:17802bc61c9bd8bc46a3624c74380adfb73cb32de8a5a59da4f3915876515aa6

Observation feaaf656-4822-432f-856a-1c299155dab3 · inbound

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models cites this paper.

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.455914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:37:35.455914Z digest=sha256:c37f0464839aed74f7740335e5279b429936777cb5cf6a1d9a704842b786af6c

Observation a8b6add4-2c32-4340-9e76-eb965d4975ed · inbound

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression cites this paper.

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:42:28.922349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:42:28.922349Z digest=sha256:3f3232c8dda59d78785d1d223cce644f4672c4c2fdf1c86d672bede338339bcc

Observation 8b8ff788-206e-4cad-bc21-8ee7d2dc6826 · inbound

ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference cites this paper.

ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:51.376207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:00:03.743899Z digest=sha256:7fe9dbff5b1f862047ccde328c1b7f4d1971a011e759451789cf5b70e178512e

Observation 7dfed331-4316-49bf-b658-c03c40981a1b · inbound

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding cites this paper.

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:26:01.376382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:29:49.681399Z digest=sha256:045e9eff76d5055466a4c6005940f1b3916f6c14d7960a155c6b29b29246a023

Observation 934549df-7e05-41c0-912b-9f82bb5c1873 · inbound

Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models cites this paper.

Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:37.432782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:29:22.152053Z digest=sha256:61511c182cb280f3ed04f20da000dfc45e555d5ce6e2f27760cbcec524d18acc

Observation 9f0a851d-c669-4ddd-b934-f862a7070ee4 · inbound

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning cites this paper.

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:25.369861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:14:04.738673Z digest=sha256:e553980b5ef399ac04d45183c754547ff1a0c0fbb6314dd936c972a7e75ce770

Observation a9079ea7-51fd-418c-a080-ce9acb14dd8c · inbound

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs cites this paper.

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:08:54.487547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:04:15.273932Z digest=sha256:58c5ca755250d89203acc361c067e17b5dcbb32e443587b18cb98acd23c886e0

Observation 37f6b381-e56a-4e33-a38a-e603e6bee3fa · inbound

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models cites this paper.

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:44:38.351501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T11:39:41.560725Z digest=sha256:3518e82e50cf9904e5f51d7b8ca3be4544c3ff10277daaa63c91ac2edef71c75

Observation 470d56ab-54bb-4b1e-87db-372c9e49ec2f · inbound

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs cites this paper.

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.123720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:02:59.403688Z digest=sha256:bfcaddbc8594cf14076ae506d6d56d975179c481473247048b3d4335ff5ce6c8

Observation 9afba30e-add1-4789-94e4-34aad4b59574 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.594522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:f178835bf5b509ed994bf7956055f9d5879cb3134573be646c20d08986d8b88e

Observation 31e0ae82-63bd-4e86-aa8f-789c4ea6cec9 · inbound

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs cites this paper.

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:50:32.497038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:50:32.497038Z digest=sha256:dc90fb36262f154428add25ebcc3b0fc74442888113aaac661a7bc06ecf4a170

Observation 9cab3c0a-6146-49ac-bb56-15432c60c487 · inbound

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models cites this paper.

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:26:49.636465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:26:49.636465Z digest=sha256:f313447c24ef7b9840c4b6d87ceedc40b239e1ff8ea851dca6c68bb056f09b98

Observation 2ff5e9df-6bdd-44d4-964f-969d81b032dd · inbound

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs cites this paper.

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T03:07:43.698719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:07:43.698719Z digest=sha256:38f29599ac34e737a99331e2c2811feff9afbaa37c0119b1fd0dee0abe31deef

Observation fe0ca0c7-afc8-4328-9ca5-3091f69bb5b4 · inbound

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models cites this paper.

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T10:42:21.341740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:42:21.341740Z digest=sha256:d718cff93f296b48d76f84f784e9a2d729f3d087f58c5c9c36aae899d4f7c403

Observation b8e56e5f-2d4b-4c0a-ba66-b94aabab9a82 · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.843428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.843428Z digest=sha256:100598642d305c0c7674cfb4f525299d1cf0d231d3f6664757649c808517a7a8

Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · inbound

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding cites this paper.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.534196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.534196Z digest=sha256:ecd25b953f1b5300c78cc066b91ff278fcbbbe0eba40a08092112dd968819ddd