Pith. sign in

Paper Citation Record · LEDGER

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2412.01818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01818 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.534196Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:26.122002Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d590f45-b32e-4d8f-9566-7aa62d9e1872 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.826191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:46f1f102f14f7d6b29f66e6393bce507483f44a3b7d9a45fc2349205e29e1644

Observation a152e2cf-71b3-4838-ae5a-24f22fbedff1 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:10.568820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:10.568820Z digest=sha256:f4d60198e39fd0be87d51738dece77b944ff487ddb980f5c9e35e0e5f5bbe62d

Observation dfa4e81f-0ea0-49af-aad9-41a331e93e78 · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.721952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.721952Z digest=sha256:c6a53b838af45a6cb701894c63b6a9f9388e595925ddcc2b917b1d14bf105321

Observation 71ec673b-1526-4118-a103-06c1a1fd7bbf · inbound

Do Concept Replacement Techniques Really Erase Unacceptable Concepts? cites this paper.

Do Concept Replacement Techniques Really Erase Unacceptable Concepts? Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:22.689471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:22.689471Z digest=sha256:815bc4620e436b0ebd6c72aca2ba4c0a7a3e4bf926041e0292b7cb2e8d2d72ba

Observation a1a373ab-ce2e-4b1e-8eac-7867ae7a5305 · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.731513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.731513Z digest=sha256:a0fa7bd57902e92ecfa7cca91e7a7ec374bee6d12e683f3821cd462eb302da34

Observation de3b0a21-a37f-471b-b485-d88fd6c34fe7 · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.833498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.833498Z digest=sha256:1ea88b865c05d0b86b982df89e37be3fd433675005aab348b46e9e0793ecc8fd

Observation 027bcc54-980e-4005-976e-4024bede4f92 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.069967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.069967Z digest=sha256:3a492d140eeb803c7c6e09ad74e77aa0a01dcaa6c50edce91c0947cb924a84d3

Observation 9769bad8-5242-44f0-8796-3fb811209059 · inbound

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning cites this paper.

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:38.007759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:38.007759Z digest=sha256:da369f0af794d307c2450061aa7ba1305fd884e7a04f362b8a97b656e5db72f3

Observation db89bafa-0a57-4cdb-9ed4-0046b86f872e · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.989156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.989156Z digest=sha256:6fcd2457961326f02a526c02f66707d40a62e405f68a3fdbd1cef55def8cc215

Observation 6e34a4b6-98fa-45e8-929f-7b2dce7e4b12 · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.501635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.501635Z digest=sha256:6a73b73c1597cd2b0f050505127a2f8a93118b18610b7ea4fd1d8a318303603a

Observation 18328f71-a4db-42db-a401-0e22267c7c09 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.802359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.802359Z digest=sha256:17802bc61c9bd8bc46a3624c74380adfb73cb32de8a5a59da4f3915876515aa6

Observation feaaf656-4822-432f-856a-1c299155dab3 · inbound

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models cites this paper.

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.455914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:37:35.455914Z digest=sha256:c37f0464839aed74f7740335e5279b429936777cb5cf6a1d9a704842b786af6c

Observation a8b6add4-2c32-4340-9e76-eb965d4975ed · inbound

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression cites this paper.

LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T13:42:28.922349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:42:28.922349Z digest=sha256:3f3232c8dda59d78785d1d223cce644f4672c4c2fdf1c86d672bede338339bcc

Observation 8b8ff788-206e-4cad-bc21-8ee7d2dc6826 · inbound

ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference cites this paper.

ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:51.376207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:00:03.743899Z digest=sha256:e0057425fbf955de17740e994e1154b1913229f027d4c0bc3cc24860df815faa

Observation 7dfed331-4316-49bf-b658-c03c40981a1b · inbound

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding cites this paper.

Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:26:01.376382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:29:49.681399Z digest=sha256:b6c6d5d195e1ffcf7e354dc893dcdb0633240f86f503c0a8866acb57662271cd

Observation 934549df-7e05-41c0-912b-9f82bb5c1873 · inbound

Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models cites this paper.

Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:37.432782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:29:22.152053Z digest=sha256:0af340f71e4477e77d781cc96ca81b5fb5f4aa3bc927267a6a07709464283000

Observation 9f0a851d-c669-4ddd-b934-f862a7070ee4 · inbound

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning cites this paper.

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:25.369861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:14:04.738673Z digest=sha256:ea88b35aa507450f1093b16456912acdfdbb43362cf8b33801c17ed4b343fc65

Observation a9079ea7-51fd-418c-a080-ce9acb14dd8c · inbound

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs cites this paper.

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:08:54.487547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T19:04:15.273932Z digest=sha256:535bc10e03400b5ef22804e0e51051b24f95938fc1324537abfb904e56afff2c

Observation 37f6b381-e56a-4e33-a38a-e603e6bee3fa · inbound

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models cites this paper.

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:44:38.351501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:39:41.560725Z digest=sha256:d9ce8b34bb0151309cc0f6f1aa280b2a4bea5d7622fd625bd0fa60ee139e1f3c

Observation 470d56ab-54bb-4b1e-87db-372c9e49ec2f · inbound

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs cites this paper.

Look Less, Reason More: Block-wise Attention Skipping for Efficient Multimodal LLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.123720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T19:02:59.403688Z digest=sha256:ae99404af846359dfd091cf44721f086a4611427f33d00567060cfd7266a5bcf

Observation 9afba30e-add1-4789-94e4-34aad4b59574 · inbound

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs cites this paper.

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.594522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:41:04.184461Z digest=sha256:ccdef389bd529d897038f716ba739bad459900744bb4d1a36bd45898e36731a7

Observation 31e0ae82-63bd-4e86-aa8f-789c4ea6cec9 · inbound

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs cites this paper.

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:50:32.497038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:50:32.497038Z digest=sha256:e30ca7e9c8b1b2d3c19b83fb65198738d7b612906b27b147dc8278c3edefc04a

Observation 9cab3c0a-6146-49ac-bb56-15432c60c487 · inbound

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models cites this paper.

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:26:49.636465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:26:49.636465Z digest=sha256:f313447c24ef7b9840c4b6d87ceedc40b239e1ff8ea851dca6c68bb056f09b98

Observation 2ff5e9df-6bdd-44d4-964f-969d81b032dd · inbound

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs cites this paper.

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T03:07:43.698719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:07:43.698719Z digest=sha256:38f29599ac34e737a99331e2c2811feff9afbaa37c0119b1fd0dee0abe31deef

Observation fe0ca0c7-afc8-4328-9ca5-3091f69bb5b4 · inbound

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models cites this paper.

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T10:42:21.341740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:42:21.341740Z digest=sha256:76d8e7ea819f740f5067ad42faf1c7ea3cf1edb6f3bc6e956a08c7c156c655e6

Observation b8e56e5f-2d4b-4c0a-ba66-b94aabab9a82 · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.843428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.843428Z digest=sha256:100598642d305c0c7674cfb4f525299d1cf0d231d3f6664757649c808517a7a8

Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · inbound

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding cites this paper.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.534196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.534196Z digest=sha256:328916199a916fd41662a7ebab353890a01792f64be3a1f0f4671f21fd65ea90