Pith. sign in

Paper Citation Record · LEDGER

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

As of 16 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2608.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06411 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:30:14.535753Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3fefcc0-c1e5-49a9-8b02-7fb1be112d50 · outbound

This paper cites Qwen2.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.184628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.184628Z digest=sha256:aca30c8734223d527ed23de3aac0073827a0c06169b0c840878bd4a67e80c131

Observation 5bf4cec4-7fbf-4622-8715-6616263e904a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.189388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.189388Z digest=sha256:32a680a91864ec03a0403db076d8b64014e3f03b87883cfeed158ebd3b213b9f

Observation 6613c666-b789-4a1d-ab34-393cc274f899 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.193471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.193471Z digest=sha256:27aa4bc000ae9e8d4d9e56e3c406221a5b4f638b7c43c192dc610f34adcd7cd5

Observation c552ff0a-8265-45dc-a260-ef63ec3ca58b · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.197067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.197067Z digest=sha256:0efd66b9358aa19a7c1f834c0f362e0c049fc9c2154dcc439a24402004fa8c81

Observation 9216dc4a-716e-4300-a6a9-785c5b82ebb1 · outbound

This paper cites ICML , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.200763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.200763Z digest=sha256:3da0b3f54179b63a44215c9848d567c28b7950c782d64568f3f212b5d73735c3

Observation 7f200349-2f41-498c-bbb6-00669a2110fe · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.204559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.204559Z digest=sha256:d17e02ff2ec50c89cd260c89689347ea810caa02daecea3798216315d52f657a

Observation 655ff2fc-a489-418d-8f29-ccdc15ad5dab · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.208606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.208606Z digest=sha256:1bf452faf9a29457f8a8251851702a79502d323a65af019c962b73f343bbf737

Observation 788c189a-8fd3-4e98-b920-0f4319ec14bc · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.211519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.211519Z digest=sha256:fcdafab8c1a46fb6864045b355fc908fbcbadaea6f56f680d6bf51d9247f83f4

Observation 36222e3f-701e-4743-b392-5722206979a5 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.214635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.214635Z digest=sha256:c945195bbcee94db47cca541029308334bdae6d7ccdd1f570b88d8be381b801d

Observation caa82ddf-0c37-421f-a9fc-9152b55e2ff1 · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.217581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.217581Z digest=sha256:e8abd3f7dd48b87b29623801a9cd66fcae89d5cf330c6c661c188590e7c7fb29

Observation 456cabf5-7ad3-4338-a54b-94098b316d52 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.220359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.220359Z digest=sha256:6927a2747bae3029141f11040dd91e87260f87e7d4a961daae7966d95e4f25b9

Observation 4a5433c3-b826-4b3f-8dd8-2695fc491d56 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.223732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.223732Z digest=sha256:f27131b7a02605360feb11c96871e25910f8a88c822bb87a6218043581d778be

Observation e3ed016b-e8d6-4909-8094-f1906e58ba7d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.227340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.227340Z digest=sha256:fed29975a289d9a51ad0a4bf6072819ff25bd18fdbc18b16ad107c11970bf8ad

Observation 498c2365-7b92-4d8d-a347-8476b14d3987 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.230915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.230915Z digest=sha256:537c0f2b79bc9cf963dadfe0a76c041d8ec0573539a8703beab6fa6f360b126e

Observation 83c95443-0f89-4b8f-b099-9c4efeda70fe · outbound

This paper cites an unresolved cited work.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.234665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.234665Z digest=sha256:9a87d44a4af20678ffcbb1c273580a973aab47d4c13e794d90ff4e618027f91f

Observation 04fcfba3-a3aa-4fc8-be22-c34dbc4866f3 · outbound

This paper cites Token Merging: Your ViT But Faster.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Token Merging: Your ViT But Faster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.237915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.237915Z digest=sha256:20848153b468997223eb657afdc72e9a07dfce12b9fae84e02ee3bb3feb91c98

Observation daddbc57-a323-435e-9fb5-7d783b1be326 · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.241516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.241516Z digest=sha256:d0d72cf85bf334e3c654b84d330727c75f89f62291a1231d856bdebad5ec78ec

Observation d0f3ba0b-a1e1-49bf-ac40-98a4e7670049 · outbound

This paper cites arXiv preprint arXiv:2403.15388 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2403.15388 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.244923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.244923Z digest=sha256:1aa3ea000766fd0d9180eb644b638d9db3820a6f8aa6ad61dd0ca596a2afccba

Observation 4cd2cbfe-3851-4f2f-9a68-ff4bc8485b30 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.248316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.248316Z digest=sha256:74557cd2cafcf8136595ab2a66576074f861203dadc345da6e9240eabb29515a

Observation c2530b53-5a83-4dd9-9946-b4b97b816b33 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.251891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.251891Z digest=sha256:f485a9cd840eb360dd397b7468bb7d963296b4a4ef0802c16b757a2417ca4cef

Observation 3c74d0f9-56b4-442e-91c8-51e2e8a3ec73 · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.255625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.255625Z digest=sha256:8a33702c621774b43b17594990d5d43fb8682f13cdacd289ca078ccd7abafc80

Observation 5c9d12e3-3039-410b-b995-177db0a4e534 · outbound

This paper cites arXiv preprint arXiv:2411.17686 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2411.17686 , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.259308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.259308Z digest=sha256:ba43e1db8a3be4a5bedccd67e24cb83586f3c8f3e9b1a6674216c2b583a6ddab

Observation b05b1e3c-ad65-4f76-a39e-be00825a9b60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.263104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.263104Z digest=sha256:1d7af7b3f3afc53884729b07e2482f94c0617571ff38844d680c06357efbd31a

Observation 73fd6f05-f54a-4edf-a71b-92ff8ec78783 · outbound

This paper cites Qwen Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.266859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.266859Z digest=sha256:d43a4776785e0559cc95f2eba9b04a2096c9af55bfe8fa8c4573b5910c91cd56

Observation 6ed976d5-52fa-4f87-b350-6d6483148a3a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.270569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.270569Z digest=sha256:677edd5e9ed193ff6d0c27b201655267404ddc49f9293d8ab026446216860fb7

Observation 100f7f4a-5ab0-4897-aeac-f15ce15c3ddc · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.274367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.274367Z digest=sha256:c4ee16435ca5249541ab661e619e041367ca3d351d6e2990d64b32120447798e

Observation 51982df7-5667-4d16-8a10-81e67ee8acb4 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.278391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.278391Z digest=sha256:3d26e122c192b1acc8a3d3cce6b663dd996b6b8b2a0b7775f47b470d706b2397

Observation 5bc6395f-69e3-4fb0-b1e5-d9efa37ef0cf · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.281720Z digest=sha256:012b2d1dbf768be099ed0ac55882c9376b686d4b3d99d5f7de758e2299e696aa

Observation 6164b773-31a9-42e1-a6bc-7d14aa21dccc · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.285282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.285282Z digest=sha256:c39892056c888c34c1f5fbe852f9ae7b96bf0757ad31ac80be596dca22a27276

Observation e270aa12-acc4-49a4-8290-05c05ff6845a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.288837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.288837Z digest=sha256:22b8f31c39e41f439cbd1741c8f91963469f4632844143aa79a7c85a36f94de4

Observation b888b749-faa7-4be3-b08e-3024b0e967d0 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.292670Z digest=sha256:6422929825c3be8056acb39ca4fadb5225b309980806bf39d597bd77a4904ab4

Observation 5f172363-a22b-4c8d-b29a-583864f922bb · outbound

This paper cites ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.296106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.296106Z digest=sha256:eadd4ac4bba7b12030ab546bd08ed8227dc34d9a9093893478bc4c118e8e5347

Observation f7e6a82c-d1ad-4b30-9c0c-272f3de478b9 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.299362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.299362Z digest=sha256:4b6abb3ff5d16ec943cc248eb46b95c364f3c6aff55248a841554f45dd3897f8

Observation 608d509a-bded-40cb-8ee1-8cb0781c9a7d · outbound

This paper cites InternLM2 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternLM2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.302724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.302724Z digest=sha256:b2bc356cbd5f0cc25785b1e6821be0ced8b5ea42ce1397cd94db66346d06de4f

Observation 92b9e655-b4f4-4e0a-9a03-f79c0a52089d · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.305742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.305742Z digest=sha256:e277c88f88e058801a639b6a71592b6243a9e8419ba98f567f4c233ed8cddb08

Observation aaa4e496-0e8a-42c7-b286-2b822f85099c · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.308686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.308686Z digest=sha256:4ed3da9cb418582bb4a71d5653c45d92e6a4659792a22e2dd12b1b72f1ccede0

Observation 5cb99119-9302-4632-ba03-ef4f6964d551 · outbound

This paper cites WACV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin WACV , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.312135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.312135Z digest=sha256:c7b49b08091f2e35117c0ac1e7a6f140ba5cacd465ced48b786906727d6cda56

Observation 678ccdbd-6575-4c75-9c72-c3410921ded0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.315460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.315460Z digest=sha256:bb687b95666ea94192ac4cd525b3e70610833227d98049bfc774c2d918afcffc

Observation 9567cdd9-78a5-40e5-b775-3c853645e88b · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.319438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.319438Z digest=sha256:5219b339752137c698beae4b89bbb12f5eb5a8eaf85447523d2273b37bc8f774

Observation 8060f9d7-21bd-49d3-b775-04767abd0b3a · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.322726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.322726Z digest=sha256:abd2faa5fd0c3c816df349592d8530c0db831c3a6ddc6e1b19531431cbd75adb

Observation c9bd4578-b93d-4a31-a85d-2ccab01855f4 · outbound

This paper cites EMNLP , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin EMNLP , year=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.326019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.326019Z digest=sha256:c1604b03bba54a846d866d92acd667e0fbfc1d5ad9114a624bd6e9fa1659344c

Observation 27495724-e151-41ce-b800-bc8f69d17a57 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.329346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.329346Z digest=sha256:27dbe8a87a520cc26b41b62c146b448ff9706fab9bff7070d79ca4653092e8be

Observation 3c74206c-3482-49aa-93da-85f0681ee83f · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.332660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.332660Z digest=sha256:eb6da3efd69b99055aebc96a747fa759bbe9db2e0d7ec59f25744020e7d85397

Observation 8b16f436-314e-4cb7-879a-1c7375f87e85 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.336324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.336324Z digest=sha256:c8a5947594a1885d3c5e28111a32a70912f4b50b2108e9a837ab66eceb5f8e83

Observation c8e980ba-49ac-46a2-9ef9-9724a899f10d · outbound

This paper cites arXiv e-prints , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv e-prints , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.339687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.339687Z digest=sha256:6896ca7d115f3f8cf3fd9807cdf7553596dd7e7d513f1bc2de23a3820f02f24b

Observation cc095e69-dfcf-4ff8-96bd-0af283339d9e · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.343274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.343274Z digest=sha256:7e311779e752a16daf02528f20028e4b1fe97a9bb9f7fcb4cc96a3fa9c79d504

Observation d44975fa-5d17-40c0-a7a4-11d836af12bf · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.347023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.347023Z digest=sha256:daf2d93154c46e21ac0139dc720d40edb2fc99df4f509654a30345d36afce2ae

Observation 0a18acc9-376b-4343-b92c-c026baf70eed · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.350680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.350680Z digest=sha256:cfd589daa5c48122a68686199fa620b490f48881627ca0e199a93af643235dc4

Observation 20e53230-899e-417f-92db-a8dfd19d0703 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.354396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.354396Z digest=sha256:0402aa375396831d647b89c927b4c3814a6bbb6c68edb62bb5b0041ee5a4eabf

Observation d406f4c0-480d-4dd0-8802-74a953a56716 · outbound

This paper cites ECCV , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , pages=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.357716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.357716Z digest=sha256:df91e2991ab4e261f22dfb6f804f203e0633e927527c3ebd66e01fbf7cb3cfb4

Observation ef15e32c-ce71-4f94-aa36-064ccacaabc5 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.360982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.360982Z digest=sha256:ef2ef15574e4a90e069626cd40677de81e63499f7060d92feee9f4c1b64ee481

Observation 18ef600a-d58c-4cbe-8cdb-1238e205f75b · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.364525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.364525Z digest=sha256:64e18689701f0be02030c3459ced4d96c960a6738918ac937bd9d628e0359446

Observation 5bbe1032-10e0-43e0-96f6-cb7127e9fa02 · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.367925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.367925Z digest=sha256:2962ee50558843def5c21fe714b5acae1a65f4ce06c9e6b449ba9d5da3989cca

Observation 57db35f6-fe3e-49e7-a9fd-989fbf4a5cde · outbound

This paper cites TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T04:30:15.167292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.371579Z digest=sha256:27b7eaacd4670c790d7a972211fa5fdfdca3e7c5af9e34262d6561338aad6472

Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.375574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.375574Z digest=sha256:e46a5fb88f18d93548be6d726006e1d87eb846d0b44abe3b12b86763185833dc

Observation 72812fb3-2681-4c41-b044-677deb2ad573 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.379313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.379313Z digest=sha256:38b96502a3dd4afe0472d98995efbf5dd7c5ea509acaf342bb598ea26f28046d

Observation b7c065a8-5ccc-4a65-a1a6-4ab6c37f6b8f · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.382832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.382832Z digest=sha256:36b1160c18473d88c92a867c4abbb42c3392710304dc6c5aa427f74d8ab2de91

Observation 19a85498-d981-42e3-8a41-06ba36b4a12c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.386155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.386155Z digest=sha256:8f3488458ac50f96ff80cee8baae4c0156b4e0865f9f19550774cb5417d92d2e

Observation 868f1a82-2656-424b-b0a5-40c2d9dde5d3 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.856172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.390028Z digest=sha256:a3f08647b15af88404f620b2e162a5737a1d2f046ae38ed481729f12d9f50bab

Observation 709f01b0-9b09-40de-bf1b-d9a8590fbc55 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.393401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.393401Z digest=sha256:b4fdc9da372927c0439b671c2645dc4b66ff3b4cd6e7cc75956920e73a2d143b

Observation 444a931e-74ea-4114-b2aa-e45f19f5938f · outbound

This paper cites ICML , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICML , pages=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.845287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.396351Z digest=sha256:8cfd97f288c61aa81b34fbbb5d786ed05be435f2a6f174b1a3b3d55bf7ffe873

Observation d3efd223-49a3-4e6d-b047-506a4b0bddac · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.399358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.399358Z digest=sha256:782bfeeb381262bac97b7813e72bdf0b83384f5a6b182d4aa78221eba21ba697

Observation 035b1fdb-dc96-40b6-aa2e-e4ddf1c81ea2 · outbound

This paper cites ACM Transactions on Database Systems (TODS) , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ACM Transactions on Database Systems (TODS) , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.834253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.403709Z digest=sha256:39ecaf5ea81292dd1922f982a1a90633a17e661e75a7c0e8d83d87dce4c25ec6

Observation 6adf513c-42ee-4652-8ca4-90b192613326 · outbound

This paper cites GPT-4o System Card.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin GPT-4o System Card

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.406724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.406724Z digest=sha256:124bde51ff9f5c681b41fff839955786f5f2b1498c554a6f108d5547cddf39ab

Observation b8db08b1-f786-44c5-981b-5861db570550 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.409858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.409858Z digest=sha256:85f5faf85d8e33a100202c5662779cdb78f31a16c4bf7d93377b222b93672fa4

Observation 2d548e17-b98a-459e-b572-53109317b777 · outbound

This paper cites Proceedings of the 32nd ACM International Conference on Multimedia , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 32nd ACM International Conference on Multimedia , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.414203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.414203Z digest=sha256:fd8c8b24dd8f3cadff4abfda13c879e1fb2d227982782268f937c46aa16b3f4c

Observation 4717ca4d-4ce6-4747-b553-94e6f052838b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.417640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.417640Z digest=sha256:17880e8211e3715c071c5590edf12cc57dab6ff14d221580c6263b0e6ac6f178

Observation 7d630164-4544-449e-8ca2-cd2eb7387854 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.421428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.421428Z digest=sha256:33d428eb1f76d942af444c0c0826495994b643ddeab1f570c327d89b3b338e7d

Observation 2bc0c6ba-b6c3-4fb4-a600-100a0fafb2da · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.424914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.424914Z digest=sha256:f49ba938ffdfff66de89e1c50ec9b68e2ef45d954f72376ad43ebdf4d599655e

Observation db8c182a-0f84-47e2-98b8-c2c81904cb62 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.800164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.428501Z digest=sha256:a5a92269e17996dccd110311954793791e8953878093bc02f16c78febdb3feaf

Observation 457f16e9-e01a-42bd-a39e-d79d004b12fc · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.432345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.432345Z digest=sha256:a3d4df8de66e936fbe4057cfb700d31f842075195e8884b1171bd2e7b5c357a9

Observation e148b1e9-b912-4325-bbd0-0e0adafa7149 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.436166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.436166Z digest=sha256:c93d46ca0ece4cf4fb65f5acab22ff61ad59b1ee480a35714ef754d3f6930761

Observation 4f4ca747-1fac-483b-b8cc-40afb9b37933 · outbound

This paper cites ICCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ICCV , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.788679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.440011Z digest=sha256:d8f0bf89861ae3bc6bc17954bbfefdc6b12f2437e97b74bdfb9cdc79e7e11d09

Observation 289f2c93-275c-42ca-99dd-849180301cb3 · outbound

This paper cites Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.443529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.443529Z digest=sha256:ef6417597cff6923c5f0d51da011c8bc277d505092e622f888a296abd8f385df

Observation e7414c56-1099-4ff0-8fe4-97423e1612eb · outbound

This paper cites Qwen3-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3-VL Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.447236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.447236Z digest=sha256:f7c4088348168eeb829e116bd49f694e4a231945deac89e68cf1fff26404f657

Observation 34ee7d7b-0e4d-426e-92b7-98060254abd8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaVA-OneVision: Easy Visual Task Transfer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.450759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.450759Z digest=sha256:15c093b8a3913aa5330757e0a4df8a3e3b8ad2e7b3dab694b5fc0d31f9272a0d

Observation 04c5c2e7-9d73-4d01-9880-688927c997e7 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.775868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.454179Z digest=sha256:2b8064c01222cf3bf204b0d6e1516a39fdb4143913b1d6ee963d3bed9edf1c63

Observation b7787d07-583c-4a18-81f7-f54d3e940750 · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.765249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.457768Z digest=sha256:805c5f30fcb11e4f132d5154ae1e98469bea048e23744b6333445de19de3be90

Observation 3ec38fa5-dacd-434a-b02d-69faf847e68e · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.754472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.461441Z digest=sha256:7588d4e4efd52a6ddd440cb756122421717a8ecf8ba25a3f0af7375d133c3273

Observation 40a305b0-f104-48b5-866b-d91e4a27bc00 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.743459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.464845Z digest=sha256:002ddaba2679b98c7dc639fa6196341f9d0ce4a12c33f98173371e9ee757800b

Observation f66abced-ddcb-4954-8520-3e15de978fef · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.731968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.468161Z digest=sha256:b2a34f7933dfd24f18ea1f069a9965fd8f932088c4d393480f68f41d779fafd5

Observation caa523ce-41c7-476e-a371-1ba415be5e77 · outbound

This paper cites Seed1.5-VL Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Seed1.5-VL Technical Report

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.471570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.471570Z digest=sha256:209205429fd456fc89d3d4c1b057501c7d6ac37557696720ccb26d549febd9be

Observation 3e278458-c443-4484-aac3-85b23c675074 · outbound

This paper cites InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin InternVL-X: Advancing and Accelerating InternVL Series with Efficient Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.475211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.475211Z digest=sha256:55c4afde8504f7b9c97511854febfc8e17597abeb31eec5d7fa6f48b26576283

Observation f3da428c-2a7a-49f5-ad92-2ef1b917e3e8 · outbound

This paper cites NeurIPS , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin NeurIPS , year=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.478903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.478903Z digest=sha256:7a0a2a47b61975baa9c9c14de119e909d68fa13b164791fc422bf8ab8286b3f8

Observation 33aaf4d2-def7-4a77-8995-12df540f570b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.482592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.482592Z digest=sha256:3023a215356dc7f0161a35b4717606ff71fe6371e1c74b6ce526137b64e11cb8

Observation 019cdcc6-3c79-4eb6-aa13-02a89bf8fe65 · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin DeepSeek-V3 Technical Report

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.486359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.486359Z digest=sha256:060b382b99fa0fdb03393806bcc0ca441f485e0093a1ef3a96f3b4986282fdb4

Observation a9fe5aea-29c6-4dcd-8398-07cc5861b1cb · outbound

This paper cites Qwen3 Technical Report.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Qwen3 Technical Report

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.490105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.490105Z digest=sha256:8a665820a01b560c4fd70c248bcd26270c1518be639e362a67919959f66ee00b

Observation 4d21e384-dc0c-4128-b777-d704c90f14bb · outbound

This paper cites AAAI , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , year=

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.712087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.493811Z digest=sha256:81ac77598ba66381c01348b77bcfc428df2c98068bd03430936bf7658f91b70f

Observation 044e74ac-68f1-45b6-b553-57d2505b46b7 · outbound

This paper cites ECCV , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin ECCV , year=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.497187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.497187Z digest=sha256:400cb1073d7b320d0a1fbabded95f3761092f5529b92aa5c5441c48d9cd8dea0

Observation f2c659c3-613d-410a-a1eb-09ec16f00ad9 · outbound

This paper cites CVPR , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin CVPR , year=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.500265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.500265Z digest=sha256:a44da5b87a032517fdaa600c41c52881bcf16c4bd0635f5c1bb1f5a167f65902

Observation c403b17c-df75-49e6-a8f4-2bb4ad3f6767 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.503529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.503529Z digest=sha256:c72c5a02e67ddcf04dc04ef9c65d6baa8c1c2f6906664eeb7791ac791b6f1ca5

Observation 8ec82576-f1d4-4563-a1ce-99ee577f7eab · outbound

This paper cites SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.506806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.506806Z digest=sha256:85b9fb4cbeb01bd0d47841c8c637ff67fb6e145039ae6367e8098a54973bf3a8

Observation 204158c3-e71a-4351-8031-9963114faf95 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.510251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.510251Z digest=sha256:8999b577a0f567570bcd13cba6e2046e4b311f422264d74626cb4b999139719f

Observation 41a5184a-a68f-454e-a9cd-4422249b0d01 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.686593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.513961Z digest=sha256:6023cf676189f55dfd04e85c88973c1aefe64d2f5524b3be450be71bcf8c9eb8

Observation 09aec212-b394-4a0a-ab5a-ec21b1b564b9 · outbound

This paper cites OneThinker: All-in-one Reasoning Model for Image and Video.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin OneThinker: All-in-one Reasoning Model for Image and Video

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.517712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.517712Z digest=sha256:07b3dbccf5de59ea0af3edacbdfd6a0c5c8f034cc7010a99945aa45573729d8a

Observation b8dbd646-f81e-45ae-9aed-dae7da1668b8 · outbound

This paper cites LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.521293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.521293Z digest=sha256:9e03b941b36cde7588ff851ebc13da48264c2a2f6eedfee049883b980addd5d7

Observation f5d1790a-cddf-444e-b41d-4f6533f24bba · outbound

This paper cites arXiv preprint arXiv:2601.22674 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2601.22674 , year=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.525086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.525086Z digest=sha256:d721fa272d642fe665ef6833b45d33ebab45aa96c9fac5355c49efe6724aa25d

Observation e7b0d891-b32c-481b-b396-e3e10cc6872e · outbound

This paper cites AAAI , pages=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin AAAI , pages=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.675107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.528582Z digest=sha256:9c5cbb237983c47a943996836b6f5dcff751ca5353ba6fa6b5db5ce7ab5ceae9

Observation 13f1b919-3589-4789-9d9a-62e60ffdf5f7 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Findings of the Association for Computational Linguistics: NAACL 2025 , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:30:15.662753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T04:30:14.531924Z digest=sha256:ee25e59d012f13cf5cd2e6dbf20e98dd4a3e1aa26dadb72b932c936e98938240

Observation f250d59f-7c12-45d5-ac2b-c74a27ad96b2 · outbound

This paper cites arXiv preprint arXiv:2602.23699 , year=.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin arXiv preprint arXiv:2602.23699 , year=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.535753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.535753Z digest=sha256:8d905930f66d7728bbb7b336c19971584c96373e35898d155d193a62f0703c23

Pith citing papers

No inbound Pith citation observations are available.