Pith. sign in

Paper Citation Record · LEDGER

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

As of 10 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2608.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06411 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:14.535753Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3fefcc0-c1e5-49a9-8b02-7fb1be112d50 · outbound

This paper cites Qwen2.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.184628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.184628Z digest=sha256:b2346891bde5823aef7c561f40bb3a357a7689298cd744797d7ab7bc8f066f66

Observation 5bf4cec4-7fbf-4622-8715-6616263e904a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.189388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.189388Z digest=sha256:9fcadfe6b1e6799ead88d2264b2a1b1608edf355eac58f95bc46772d6a87d003

Observation 6613c666-b789-4a1d-ab34-393cc274f899 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.193471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.193471Z digest=sha256:b239d0b0c2d3120a895d1e465471cfc7974cae87285dc3a421bcfc6bfed12c40

Observation c552ff0a-8265-45dc-a260-ef63ec3ca58b · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.197067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.197067Z digest=sha256:2ca185d1f90fa90e3fa90f9e91681df0e465883488add6f03cfdfab11349af96

Observation 9216dc4a-716e-4300-a6a9-785c5b82ebb1 · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.200763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.200763Z digest=sha256:76be501fd33aae1b3352fa99a8805fedbf5030ca9e1659b1f378dcc2a8ffd2f1

Observation 7f200349-2f41-498c-bbb6-00669a2110fe · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.204559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.204559Z digest=sha256:e4ab1c6a3e2e622b0f181063ba2c54b89dff82cd73fb8fdae8dd36ccb44502a1

Observation 655ff2fc-a489-418d-8f29-ccdc15ad5dab · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.208606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.208606Z digest=sha256:7a2fc08aa51e372f21ec56b53be29214dc91c3181075dd7a8d52c468c6bc05bf

Observation 788c189a-8fd3-4e98-b920-0f4319ec14bc · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.211519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.211519Z digest=sha256:2e810c6f7370bc7fe356ef1df19a4fbc102071a7638bba8a531625284a9fe622

Observation 36222e3f-701e-4743-b392-5722206979a5 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.214635Z digest=sha256:1f9ec60eaa3ae8154ab95011513b986b2473d234f9d637c4a76ebdf2ef4407ad

Observation caa82ddf-0c37-421f-a9fc-9152b55e2ff1 · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.217581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.217581Z digest=sha256:bbae82e76492c969ade81452e40e28c51f9af3f9d44aaec2696f7f71bb2fe955

Observation 456cabf5-7ad3-4338-a54b-94098b316d52 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.220359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.220359Z digest=sha256:bff6c7ad52c7a4056cd6bef5c7a777d5e9c54f3c71ff9e6cf225071eab206e44

Observation 4a5433c3-b826-4b3f-8dd8-2695fc491d56 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.223732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.223732Z digest=sha256:8858c63b73adc66cb551a837acfa40179e616ecb2a634511ece609cc5de85667

Observation e3ed016b-e8d6-4909-8094-f1906e58ba7d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.227340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.227340Z digest=sha256:8a7c98ca51a32ea990114f0f8701734764a6fea6b2006deee698c5c0325d0507

Observation 498c2365-7b92-4d8d-a347-8476b14d3987 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.230915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.230915Z digest=sha256:b590ec5e361200326da23dbff5b5b21a3e8b28c3e375feb3889d51bc93c26b26

Observation 83c95443-0f89-4b8f-b099-9c4efeda70fe · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.234665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.234665Z digest=sha256:a4e82e3be1e533701cff070d9344121b45eacc73db5917e902a18f758d0b76cb

Observation 04fcfba3-a3aa-4fc8-be22-c34dbc4866f3 · outbound

This paper cites Token Merging: Your ViT But Faster.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Token Merging: Your ViT But Faster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.237915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.237915Z digest=sha256:254f0d47cfbdbd0fa38256d6ade622c15a1eca8d44d4ac7cec52589fe179b395

Observation daddbc57-a323-435e-9fb5-7d783b1be326 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.241516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.241516Z digest=sha256:18ac68e11df9391beca4b07265675209f37350a3e06a8d198fcba24aa21dd799

Observation d0f3ba0b-a1e1-49bf-ac40-98a4e7670049 · outbound

This paper cites arXiv preprint arXiv:2403.15388 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2403.15388 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.244923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.244923Z digest=sha256:469d2ec8cc3ece352c40cf4714edaf03d6ae5b1d71850d46422a2d4476a61ed1

Observation 4cd2cbfe-3851-4f2f-9a68-ff4bc8485b30 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.248316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.248316Z digest=sha256:794dd09d4458e8fe5f09011ffbabf63391233428c128537b60149d4b7ca5155f

Observation c2530b53-5a83-4dd9-9946-b4b97b816b33 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.251891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.251891Z digest=sha256:7d62c336b0213a3bea92fa805c03c60bc9699588d072c8ac1f8d2b63fd6a38b8

Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.255625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.255625Z digest=sha256:60e17e7e161129d52c72abdb70250f4668efe7181432dbc35535f25b5a05d7d2

Observation 5c9d12e3-3039-410b-b995-177db0a4e534 · outbound

This paper cites arXiv preprint arXiv:2411.17686 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2411.17686 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.259308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.259308Z digest=sha256:fbae7eef54a48c5447f342f5a998f7edd22481aa36234a35c1ba1ff6d27c7ad5

Observation b05b1e3c-ad65-4f76-a39e-be00825a9b60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.263104Z digest=sha256:d1173c7b7c5f2c1aa0f852d535ee4c0080b174c73b7051cf39d814ea51fd6643

Observation 73fd6f05-f54a-4edf-a71b-92ff8ec78783 · outbound

This paper cites Qwen Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.266859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.266859Z digest=sha256:fdea88e5d508ba93812c6de0d41e2f15e0bfb50f881c72556b94bf47023acac4

Observation 6ed976d5-52fa-4f87-b350-6d6483148a3a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.270569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.270569Z digest=sha256:881d75d82c1fe2a9786b0669f0a6d72cffe82d42dec1d019a40d584d7a341652

Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.274367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.274367Z digest=sha256:93bf9e0cce9722821eb3b4f47650dac9101f5fa695db032d6baee64c1d7bf906

Observation 51982df7-5667-4d16-8a10-81e67ee8acb4 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.278391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.278391Z digest=sha256:297b1c8a151beec5128d2ccd75613c9e5674331356a5e71fc3107a27ca6233fb

Observation 5bc6395f-69e3-4fb0-b1e5-d9efa37ef0cf · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.281720Z digest=sha256:d772db12291941c43d286433549b12d5e714b3a2542c179822fe06fcb8fe686d

Observation 6164b773-31a9-42e1-a6bc-7d14aa21dccc · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.285282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.285282Z digest=sha256:2ce8170b275f971fe5ccc7fabe9c1c104acdd8cf44d4694adf618e3b1d0083f8

Observation e270aa12-acc4-49a4-8290-05c05ff6845a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.288837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.288837Z digest=sha256:3bc82cc7c4fa3a62f073a13c77a2e9f10f4672dd527625253893bf94e9e83e1c

Observation b888b749-faa7-4be3-b08e-3024b0e967d0 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.292670Z digest=sha256:14b1b2b13401598a942e907f133740533754042c06a546b96be2640ef8eecc73

Observation 5f172363-a22b-4c8d-b29a-583864f922bb · outbound

This paper cites ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.296106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.296106Z digest=sha256:80cf3ad0bef8ea835fc5a2836104a275c4794ad6032b50a483e0f64fe8ee3128

Observation f7e6a82c-d1ad-4b30-9c0c-272f3de478b9 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.299362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.299362Z digest=sha256:095b30afcfe8079fdf63db859fea92527715db0dd8796a34e5066de235a21895

Observation 608d509a-bded-40cb-8ee1-8cb0781c9a7d · outbound

This paper cites InternLM2 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternLM2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.302724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.302724Z digest=sha256:70af49bb47d944a61becf788aa05f1de60ff1205c86d56520e6a4b4d73ce2364

Observation 92b9e655-b4f4-4e0a-9a03-f79c0a52089d · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.305742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.305742Z digest=sha256:a177344fa035834aee92b845926e733386ece7882c73cd083a64df95a121c245

Observation aaa4e496-0e8a-42c7-b286-2b822f85099c · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.308686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.308686Z digest=sha256:4d80e9fd336fc9ecbd63abb92cea3953cce40ca629e77df7c440029ca68ab838

Observation 5cb99119-9302-4632-ba03-ef4f6964d551 · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.312135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.312135Z digest=sha256:fb58f64b39640207937aca9430a842503999ed3f9153246f8b25e3ff21475b25

Observation 678ccdbd-6575-4c75-9c72-c3410921ded0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.315460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.315460Z digest=sha256:26334832726a1692bd24f9983cf555d3a9e3e40899ce4a809cb5fbf2bd2eb48b

Observation 9567cdd9-78a5-40e5-b775-3c853645e88b · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.319438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.319438Z digest=sha256:9a4ddd7e0155536f0b8caa60a28c6697771fe1815d89099f5cc93dd3a22d90ef

Observation 8060f9d7-21bd-49d3-b775-04767abd0b3a · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.322726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.322726Z digest=sha256:7fd3273ecac7bb39870fe70fa0ddc1da2f25994e1db8f9f5b917334628e44b91

Observation c9bd4578-b93d-4a31-a85d-2ccab01855f4 · outbound

This paper cites EMNLP , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin EMNLP , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.326019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.326019Z digest=sha256:7850b3304d7879f72203980353d9c2b7f5bf869f3855e53512f33e10e1594428

Observation 27495724-e151-41ce-b800-bc8f69d17a57 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.329346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.329346Z digest=sha256:073e9c8c004541e18d8b4a1dcfa7b4c445f0a72d3eaad02070522924485deed7

Observation 3c74206c-3482-49aa-93da-85f0681ee83f · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.332660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.332660Z digest=sha256:d40c20b90afc683912ca92ad5654654f56025387a1ef05291b715d6e6f43fe92

Observation 8b16f436-314e-4cb7-879a-1c7375f87e85 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.336324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.336324Z digest=sha256:ac62ff3ba546bfd2819e18c8385efb4f10f2cd75646f3ce81b8865711d89ecfd

Observation c8e980ba-49ac-46a2-9ef9-9724a899f10d · outbound

This paper cites arXiv e-prints , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv e-prints , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.339687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.339687Z digest=sha256:9f611f4c7a4b9e32d05db77fff27aeb6881efd6110b32d0c7b102f10cf1aba1f

Observation cc095e69-dfcf-4ff8-96bd-0af283339d9e · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.343274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.343274Z digest=sha256:39f6f414603e9db0fa1e35bf5e56d1faf0b9771958815cb673765f3a6c39c070

Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.347023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.347023Z digest=sha256:fa724a6440af88201937f1147c170d4703e80279b26794ea55e7290d0705123d

Observation 0a18acc9-376b-4343-b92c-c026baf70eed · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.350680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.350680Z digest=sha256:429f634df87039fd9f37088c2fdd6c05c2ef60caece5e5986ec9407e248db767

Observation 20e53230-899e-417f-92db-a8dfd19d0703 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.354396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.354396Z digest=sha256:ea025ee6c83ef10dd0a661bb6b122538aa318e0413fb1c45b4ff249ac6009615

Observation d406f4c0-480d-4dd0-8802-74a953a56716 · outbound

This paper cites ECCV , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.357716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.357716Z digest=sha256:8fbf1a9db94304b3086ffd8fbe0e454e090ee447f9f579f26cfc3d21d91bf461

Observation ef15e32c-ce71-4f94-aa36-064ccacaabc5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.360982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.360982Z digest=sha256:e83a72004a7d0ce4fe1cbdd60380ed0f032b046f9e9361c643d4a37d90bda37b

Observation 18ef600a-d58c-4cbe-8cdb-1238e205f75b · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.364525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.364525Z digest=sha256:d4a2f5612c343d1b712aefe87998de5e3b771dd5ea9264ff24803d6b414e19e6

Observation 5bbe1032-10e0-43e0-96f6-cb7127e9fa02 · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.367925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.367925Z digest=sha256:62d2b0d65427ec9eb32e14cbb53869908ea71fda0e654133b3412f0a7135a1b5

Observation 57db35f6-fe3e-49e7-a9fd-989fbf4a5cde · outbound

This paper cites TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T04:30:15.167292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.371579Z digest=sha256:192dc96755b9977631e56b2138cea9a0d31d1abbc247f697962be78229d90e76

Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.375574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.375574Z digest=sha256:1eed4c703e20ab629a20f9bd31ea2f56796e34776c8e0970b18926c5e04e0e66

Observation 72812fb3-2681-4c41-b044-677deb2ad573 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.379313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.379313Z digest=sha256:244ac0d5b9224b913e730c884a37075730117740be9506b41ff4011b80a70124

Observation b7c065a8-5ccc-4a65-a1a6-4ab6c37f6b8f · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.382832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.382832Z digest=sha256:a666e4884483da7d4427c1ddd0120fa0704d082d8636a0ac07b324296ec09a27

Observation 19a85498-d981-42e3-8a41-06ba36b4a12c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.386155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.386155Z digest=sha256:8b085dba0a90e76b30f0277f7b196ef3bbade0de6eacb6dddccd4b45ff56f43a

Observation 868f1a82-2656-424b-b0a5-40c2d9dde5d3 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.856172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.390028Z digest=sha256:d57136694ec58918b49c4249bdfc8d2b6bfc49448793f03c8b81ff5a2e161b93

Observation 709f01b0-9b09-40de-bf1b-d9a8590fbc55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.393401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.393401Z digest=sha256:96c64a4033eaa7b170b10ac5db747cd8a386cecde0d2dc940a4a7b65c3fbecdb

Observation 444a931e-74ea-4114-b2aa-e45f19f5938f · outbound

This paper cites ICML , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , pages=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.845287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.396351Z digest=sha256:aa875e39072f8144be9f1a113beac7cb105949828bf15a568adb4aef85d4e04d

Observation d3efd223-49a3-4e6d-b047-506a4b0bddac · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.399358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.399358Z digest=sha256:bb3739baa7473e585812edad5701e39281bf9284956bce75f3c434c44001a856

Observation 035b1fdb-dc96-40b6-aa2e-e4ddf1c81ea2 · outbound

This paper cites ACM Transactions on Database Systems (TODS) , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ACM Transactions on Database Systems (TODS) , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.834253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.403709Z digest=sha256:c31e479a2aa11becb91671deb1046b5176f9408bb165d86fa0e52d201064b454

Observation 6adf513c-42ee-4652-8ca4-90b192613326 · outbound

This paper cites GPT-4o System Card.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin GPT-4o System Card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.406724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.406724Z digest=sha256:aa301311a6bc6f59f35c4c7f9104d57667d6b4669c4e8ca2a19247025e839227

Observation b8db08b1-f786-44c5-981b-5861db570550 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.409858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.409858Z digest=sha256:aef432eb682be9d5cb462575e051754dc7fc009d75596d27f7a163675e608946

Observation 2d548e17-b98a-459e-b572-53109317b777 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.414203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.414203Z digest=sha256:1692987ddbd5fef52b58aadeed6229d0e46818ea6b3ddbb6f56b1b85a02ad07d

Observation 4717ca4d-4ce6-4747-b553-94e6f052838b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.417640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.417640Z digest=sha256:c73d01d3d6a3a66dd69141c9dc0d6eee84e0f7fd211188a66afc6de39885a246

Observation 7d630164-4544-449e-8ca2-cd2eb7387854 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.421428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.421428Z digest=sha256:10d4832c04118b30d3b437129b4a08983727454f2c5a7c3610e73a03c95dea64

Observation 2bc0c6ba-b6c3-4fb4-a600-100a0fafb2da · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.424914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.424914Z digest=sha256:0938f3072cb91bfcbe8f1dd6c04cf00d382204a39fc7c0613e487fd467a6774a

Observation db8c182a-0f84-47e2-98b8-c2c81904cb62 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.800164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.428501Z digest=sha256:32cd7767bf2ccbf7cbf424e9f57f09a512f86cd2de787e897d34d997cf66f27b

Observation 457f16e9-e01a-42bd-a39e-d79d004b12fc · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.432345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.432345Z digest=sha256:62c4395b79d248335c8fc700e2a44158979f24b4c84dfa1baa74149707965532

Observation e148b1e9-b912-4325-bbd0-0e0adafa7149 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.436166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.436166Z digest=sha256:da3aea1316ea4961115de9fceb40d9ad0f57896904776e9453d358cda367f428

Observation 4f4ca747-1fac-483b-b8cc-40afb9b37933 · outbound

This paper cites ICCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICCV , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.788679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.440011Z digest=sha256:235b6966baf46427cf8350daba756677153b56d078edb716aa009e1bf94bfe6a

Observation 289f2c93-275c-42ca-99dd-849180301cb3 · outbound

This paper cites Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.443529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.443529Z digest=sha256:47101cccb24cbc3e00e2f04e170cefc6f6704b847981ca8db8c3e5662bb4b03d

Observation e7414c56-1099-4ff0-8fe4-97423e1612eb · outbound

This paper cites Qwen3-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3-VL Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.447236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.447236Z digest=sha256:5c07454ed32ae0b8d55d8b74b620bf10e2c1409e0b5896024a3d23e9d46c2650

Observation 34ee7d7b-0e4d-426e-92b7-98060254abd8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-OneVision: Easy Visual Task Transfer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.450759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.450759Z digest=sha256:6fd9c734b3435ee8dbae200f20d1f098156240fd23fef7becff96b8dcdce9b61

Observation 04c5c2e7-9d73-4d01-9880-688927c997e7 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.775868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.454179Z digest=sha256:4b8808b6e801ab6d7efb73cd252b27afc25f56f19995295ba9b577b3d9b3faca

Observation b7787d07-583c-4a18-81f7-f54d3e940750 · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.765249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.457768Z digest=sha256:a55919c87713cdf70379006f99e40de4084c293190e0bcdb6e2e113457ebc750

Observation 3ec38fa5-dacd-434a-b02d-69faf847e68e · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.754472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.461441Z digest=sha256:31992d1b8f1fafd65b659e3b07ba0e4f5cbf336c9184bb8b96ff552c796fd398

Observation 40a305b0-f104-48b5-866b-d91e4a27bc00 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.743459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.464845Z digest=sha256:210f0178ec0c78b4fe52bcd211d05aa49d3bb35bb85c88197d9f08aae9d88298

Observation f66abced-ddcb-4954-8520-3e15de978fef · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.731968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.468161Z digest=sha256:c10676bed64b4f1fb77987364aae59365751dd5f59c16aa697f6f2830012db1a

Observation caa523ce-41c7-476e-a371-1ba415be5e77 · outbound

This paper cites Seed1.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Seed1.5-VL Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.471570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.471570Z digest=sha256:cb78fa54155137a7cef62f25ff7e9bffb62b31fbf1453b4321023d887fb82bc7

Observation 3e278458-c443-4484-aac3-85b23c675074 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:53a39ed197cb5ecdd3601757ed303f6f09434d14821eac68316b83fe87efe083

Observation f3da428c-2a7a-49f5-ad92-2ef1b917e3e8 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.478903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.478903Z digest=sha256:3952d79213051109b26956932b94a72e638abd8c4e1b6b08e50258bb4ef0f303

Observation 33aaf4d2-def7-4a77-8995-12df540f570b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.482592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.482592Z digest=sha256:189e67169b86098c613b23c8e972fcef9f7c59b1246787392d5b1585665af766

Observation 019cdcc6-3c79-4eb6-aa13-02a89bf8fe65 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-V3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.486359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.486359Z digest=sha256:447ae3c7092f0ffb36b995a54b57f9f8b18d86f2f3fd05b8d84f4de86996a943

Observation a9fe5aea-29c6-4dcd-8398-07cc5861b1cb · outbound

This paper cites Qwen3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.490105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.490105Z digest=sha256:807015ba4efabb8a6bb3839f6a8ebc8f098be58527178912c4ab0c55df063ee2

Observation 4d21e384-dc0c-4128-b777-d704c90f14bb · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.712087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.493811Z digest=sha256:7b6f3df6eabc9feb8a5666891c5815f65a7de08e998977892f179a66306d16db

Observation 044e74ac-68f1-45b6-b553-57d2505b46b7 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.497187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.497187Z digest=sha256:182662601abf6d6d69d8aa283d86ac2dade584ff0731af289d1e2786c7937f82

Observation f2c659c3-613d-410a-a1eb-09ec16f00ad9 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.500265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.500265Z digest=sha256:cccf926f534b523477f610bee5040443d7c2244a3a815137c730a6d307ac2caa

Observation c403b17c-df75-49e6-a8f4-2bb4ad3f6767 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.503529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.503529Z digest=sha256:586a39c31e107c68ddca57cfbf2c70504337bdb5b815cdb5ee536c79d155b488

Observation 8ec82576-f1d4-4563-a1ce-99ee577f7eab · outbound

This paper cites SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.506806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.506806Z digest=sha256:9c6470ca963405d27d0b77546e981cee207dd296c31ad5d1996109b85220a004

Observation 204158c3-e71a-4351-8031-9963114faf95 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.510251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.510251Z digest=sha256:27944f78fafecc5ffbf64211b75157c5af90024783a65d7012e4968e5637f468

Observation 41a5184a-a68f-454e-a9cd-4422249b0d01 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.686593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.513961Z digest=sha256:f604dcbc01b1853767d8361643c39e8945557dcc2fdf25a7f59049ed0b8d4f1a

Observation 09aec212-b394-4a0a-ab5a-ec21b1b564b9 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin OneThinker: All-in-one Reasoning Model for Image and Video

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.517712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.517712Z digest=sha256:f48de8080e426d96787e2d25b8e23d575fb1da83ddfcd371c578fe6eee7afe59

Observation b8dbd646-f81e-45ae-9aed-dae7da1668b8 · outbound

This paper cites LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.521293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.521293Z digest=sha256:407177a0cd2ee668529ad1aef584721678c5cfc3aa61f17c6f4571bdc08b62e7

Observation f5d1790a-cddf-444e-b41d-4f6533f24bba · outbound

This paper cites arXiv preprint arXiv:2601.22674 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2601.22674 , year=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.525086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.525086Z digest=sha256:826dd0a5a4af7d9719b95501de7bdaea70873de2d58530456a9496de6fc254e3

Observation e7b0d891-b32c-481b-b396-e3e10cc6872e · outbound

This paper cites AAAI , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , pages=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.675107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.528582Z digest=sha256:2ce98c7b31f3e0286b80dde1e2a709a695c5bd194237f645ad5887cb01cb1cb8

Observation 13f1b919-3589-4789-9d9a-62e60ffdf5f7 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Findings of the Association for Computational Linguistics: NAACL 2025 , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.662753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.531924Z digest=sha256:2d8dd21b0cacd878bca5abb5a27e40977615c27fce5677e4915674bf0ef4f83b

Observation f250d59f-7c12-45d5-ac2b-c74a27ad96b2 · outbound

This paper cites arXiv preprint arXiv:2602.23699 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2602.23699 , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.535753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.535753Z digest=sha256:0175d9663b114d505cce633a909aa7b4686b4a293754e71ab9fa02a45762e4f4

Pith citing papers

No inbound Pith citation observations are available.