Pith. sign in

Paper Citation Record · LEDGER

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

As of 10 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 3 inbound Pith citation observations for arXiv:2501.19036.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.19036 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:39:21.861654Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:01.186306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:48:11.380911Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 222ea20d-5bed-4f91-ac7a-2818be6e5d5c · outbound

This paper cites online" 'onlinestring :=.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.594738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.594738Z digest=sha256:383843aa99b2b110a0d4d406fcdc9eaedd9bce2b3045142006ceec4484aa87dd

Observation 3c35d43a-5902-4d3c-b109-f298906454de · outbound

This paper cites write newline.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.600420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.600420Z digest=sha256:d845fdc17626b786a7a23912cd4dff3021d796d6665f8246297384e7f2d76389

Observation a3d53ece-7c89-4d3f-ad24-5edcfff3ae1a · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.888534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.605830Z digest=sha256:3fec98ee615a95637fe7060e828b2a6438e0e241ab7ab7f471dd23a62ce068c8

Observation dc10cd2e-a6d5-43d5-b44d-345cf927049f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.610540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.610540Z digest=sha256:9dd0a56d823adb2eeee4df2f5bdb3ce950c76ef6e0443fd83729685dd245d3a3

Observation fc84dbf1-6106-411e-8a4a-02d046500bbd · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.874634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.615711Z digest=sha256:94eaa4c1de7e8c5cf23493e6e148d4e0bffad2924bb4b449bd348ba600a5ca46

Observation ca3ca5c8-7e53-46b0-be24-42e7e55a3e61 · outbound

This paper cites Honeybee: Locality-enhanced Projector for Multimodal LLM.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Honeybee: Locality-enhanced Projector for Multimodal LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.620447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.620447Z digest=sha256:7c7c1edab4b6d35ce8a826c4fed57f7c4971874b3ed060f1db117c938ff1f2a5

Observation 077d4ead-ccc3-41fa-9f43-9f3eeb15daba · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.625551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.625551Z digest=sha256:8f7efc6b626f92aee687788b6a91aa2a2c335d6e7bf04c5f62fecebe9b8fc277

Observation bd44a04c-86c7-44f6-a822-f14fb55347d5 · outbound

This paper cites EVLM: An Efficient Vision-Language Model for Visual Understanding.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs EVLM: An Efficient Vision-Language Model for Visual Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.630533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.630533Z digest=sha256:bb75142d07e7e9b057217f392e7936ddecdb5e5b4d12dd46df524cc95531e266

Observation 215cdbb7-72a1-49d4-9395-b0523370df17 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.859880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.636011Z digest=sha256:b3e268a563519deab3bc38605a2f79de03aa256ce994983e6d38564a85fdc515

Observation 8766b507-a8a1-4cfc-bdb9-841312557341 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.640784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.640784Z digest=sha256:ed28a12bbc2b28b342dc9cbfa691ed3f176583ddfb0ca66471ea75d39185e269

Observation 3316b658-e3d9-4c2f-ad8b-e26eed8ef1e5 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.645567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.645567Z digest=sha256:02b1846d011e8bef1b117e2289f89c458d77e8019be4b5568a488d1e24567dd2

Observation a6332ece-49de-468c-911d-650a0a05bce6 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.844867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.650396Z digest=sha256:dfbacf6979d02b958f5252505c7f5a212ccf0e298d120ab23dd27e969b83e314

Observation 2741abf0-0153-4bb6-b11d-fb990c6f6022 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs NVLM: Open Frontier-Class Multimodal LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.655000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.655000Z digest=sha256:696381daa745a0a02963c12dc3458fcb0e2e718f8b94f073c56376309538ce0a

Observation 2174a1bc-c5ba-4299-aaab-6013ab08f594 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.831353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.659804Z digest=sha256:efa5466ba14ca72dfb193a77e81fbf268f8057e63ffeb9d1e1c6a2f81734b6d0

Observation 7de5c3ec-827c-4522-a171-be12207d8278 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.664380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.664380Z digest=sha256:ea78c315a4a2cdb512d98fa7faa8df866315cc22031e539dabdf8457acfee756

Observation 887fbce0-fff0-4376-b56c-d62e7a8b2f03 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.817100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.669334Z digest=sha256:420019fc6515a920696b4c69c954829fadc3a5f403fcbe32833b832fd00bc7f6

Observation 6b3084b6-0e01-438c-9225-8b44ceab7b43 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.673873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.673873Z digest=sha256:ce3c58632fc92403b25d84f860b423f3a4892506406990a8732204fa95801ef5

Observation cfc2897f-05b8-4284-be91-6c7857f74d71 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.802353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.679248Z digest=sha256:3f9123e7ced105c85c01105d0892dd9721a17e46a3bf5aab75ab91dbe5dab4a2

Observation 8cb16a2f-43e5-403f-ae05-be68372db9e9 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.788815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.683986Z digest=sha256:fb07f5e4b8f799c010e3f53f04d6d2e8eb3f4467c65861b8e4ad1df31a605ff7

Observation a9fa66db-bd1a-44e2-9ead-887ff5a80c2e · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.774328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.688518Z digest=sha256:4c039ec3b6f4b417bcdc9624fcb144bc7cf2544d6642205bb774e666ff274206

Observation 237ec53f-4b41-4e22-a08c-21165a974fdb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.692800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.692800Z digest=sha256:ee167622be448ac805744b9b7e27a713a277c24b8d5de3c77e0b252a65d872a2

Observation 41854b49-82df-4c20-af01-61df17409811 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs CogVLM2: Visual Language Models for Image and Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.697649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.697649Z digest=sha256:6665c32745a579ef2eff89e00a8463dc20baaf474fbcf9a9edff14b64f33c1f2

Observation 87caebc6-9a09-42e7-85d4-f1a68256f4f8 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.702360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.702360Z digest=sha256:f91e8a3414599ba99c0d94ac9e155ce0265c00ebcfd6ce90f229a5572b21e599

Observation 7828f2ba-fc4f-4316-ac5e-3fb15c75e4c4 · outbound

This paper cites Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.707003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.707003Z digest=sha256:1f6a45d0a2c2e51c21f003c64b11543347697141c9e5e5be1bc9cb6181ecbda1

Observation a6dbdbe5-26c2-4a7e-8986-ed95ddb263f6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.712002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.712002Z digest=sha256:0c35e0aa9e5cdf9fd8a5a59d260419e4aac3efae9b6a245d511d071fa0445ada

Observation 7fe2c3aa-28cc-4d73-89b8-53e7f6111f89 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.760743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.717944Z digest=sha256:eee318e0e4c4b1f910150918f599655a207e9aa021c90f6c6add4956b14384c2

Observation 2fc817c6-8eaa-4b45-b329-e0d3eb99c862 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.747349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.723054Z digest=sha256:5136b1e89f034dd7169ab830a21f9c1543a5cea57aafed104d8e069305232f56

Observation 8904b5bd-f1e7-4be9-8402-952d135fd0e0 · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.727495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.727495Z digest=sha256:354161be74128528794aaf6715219425dc968f25cc585cc6f28bc13015122670

Observation c2915a8a-18f5-4106-8305-89651466dcec · outbound

This paper cites Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-09T21:39:22.353394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.732229Z digest=sha256:ccb7df3aa755d461d07805a22727b0f2042b2d22a45b25bf31d53d57b1c59fca

Observation 26fbfe0b-bf26-41e8-b36c-53bbc4b8f4b2 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.733912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.736777Z digest=sha256:e28f2a221f07db07e86455d6a80f36d92562178f6b36541e469d013d47a295c4

Observation 0028a4ce-6e4a-473f-8965-2d503642b468 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.741090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.741090Z digest=sha256:44cdcf5cfbaedb55b7422239cce2ae9ee52f408fbf99f0f9e04d3138a9b0f8d7

Observation eb1f799d-778e-4011-aa34-48b8caaebd8d · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.745753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.745753Z digest=sha256:a6d36fe801ad9ce78b8b7839238c5821705afe105e0d79a58a08388c6cd905f3

Observation eee53010-8633-4f0b-b8b3-3a0d6f51fcd6 · outbound

This paper cites The Llama 3 Herd of Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs The Llama 3 Herd of Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.750252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.750252Z digest=sha256:21756915836a6741475723a8227fb2767cb51dfbe657ec1231421403d858ae01

Observation 8520d8b6-57ba-41ef-92d4-c0964048dcbc · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.755070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.755070Z digest=sha256:df513518c7736747fa056a109607d8bc2cdd102a9df4de14cb68cfacec60204d

Observation fbe7a7a7-d44f-448a-9640-e04258538e1d · outbound

This paper cites EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.759828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.759828Z digest=sha256:5788d40e5086a077e752ed9efedaa7216acff568fbcc9c1e26b059f7aeb177bb

Observation 92050a1f-59cb-43cf-b926-c22eadae543a · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.720298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.764670Z digest=sha256:b4b5aacf6df86a362726e8b446d0c7b42550281655affb332f496f72859839b0

Observation 3bd30714-1a7d-4e8a-94db-203ec33f5f63 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.705955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.769059Z digest=sha256:fb68fe847c9c069d000356b59a3caa1a9a4bf23fd3e75f5e7476cb0dfd3f8435

Observation b89ca2fe-2e41-4e92-8b88-b3c5691e8c8f · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.691835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.773380Z digest=sha256:2a64b4690da110cebee842dcf4bac81aa3f1c37af8a4face3317cfa4830d5799

Observation facfb774-f4a6-4f47-bfdb-fa1bd6624521 · outbound

This paper cites GPT-4 Technical Report.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.777707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.777707Z digest=sha256:d737c86d6a17c96abb2d30154279eec276cf396ac872afd43e73dc50c8dbcd99

Observation 5604f3e7-d0b1-4d27-b5e3-47ba0a11af65 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.782512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.782512Z digest=sha256:f4bf655acdc018d465a178912ff6af418480117ff910837d548c4ee9f897aa99

Observation ef9f0241-1058-4be1-ad63-9e59597e1cd2 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.676596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.786916Z digest=sha256:253e9df6338feb3c796cd18aaeeb58bce419578a4f212840bbde9cb8da2609f5

Observation c7d9d465-43c7-4a91-b118-0b289190c335 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.661158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.791790Z digest=sha256:8267563dfc913e4c9758cb2351b28c0f51b8008ae883f103315d7679f592efc7

Observation d6de5c31-b80c-4914-8aec-5a89ad417600 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.796142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.796142Z digest=sha256:6e80b7f9b711eb7b740eaaa9152fa2ecdbbe5145832b9ad8a386dbc343df2272

Observation af78091b-cb08-44a2-9841-a6f6ed2f9306 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.646846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.800625Z digest=sha256:6577f376043607fb924a4f52cfcf6701527d1facac0231777ad0ebb6e4987549

Observation a8947d7f-026d-484f-8822-50cce7a77ca7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.804905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.804905Z digest=sha256:fd82beda88db6f62e906d3cda80eafed4b6183378d2f9743574b714c6bbc7c00

Observation ae49b89e-fb72-4cfa-9e34-b119b4cf6da8 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs CogVLM: Visual Expert for Pretrained Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.810068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.810068Z digest=sha256:93f8ada13209cf8143981135e0dbd31a43616c0c3b09c0557bec527c3efde149

Observation ffdc1457-f9ee-4721-9ab7-167d3ba27c18 · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.631355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.814948Z digest=sha256:775fd292b39da0579bf82f1a03a4f955484f6ec7bb823c935bf13070698ec60a

Observation bb3c3ddd-942d-4ca5-98b3-0f55824435f5 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.819432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.819432Z digest=sha256:39b7a4725cf8878b2a4bd5bcf957332e6ed638e58f39951749bb55c991b9153e

Observation b1a9de48-f850-4970-9e30-ba437f924d71 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.824073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.824073Z digest=sha256:3947c55de09d9caae4a1cec841b020f76818f265e2d65404632974170f959034

Observation 2c82955c-fac4-43e1-a7d3-409672aa811a · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.616383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.828897Z digest=sha256:0024f5beb04ed0b920e86b3cab7e0984859916b473b0cf6103dffe5cd51feacc

Observation 66755296-f36e-4202-8e1f-c7c826274ca8 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.833240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.833240Z digest=sha256:a9fe33d0eea2396ab8f578b183d87da019ec15b55c44609b7f3ba86b01a876f2

Observation 6f04beb5-771b-412e-b007-cf7135626d65 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.838073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.838073Z digest=sha256:b0f3c14ba498918813302e7c899b360521659773a8a3db8a555ea2faf803ed70

Observation 5f4d8a90-9c17-4568-987c-dd50b5a74b1f · outbound

This paper cites TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.842719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.842719Z digest=sha256:255d927e044affde6df11e4745b4f99c977851a4cf0f7c8696caccfa21bb3775

Observation 3600d8d3-1a12-40b0-ba97-0d77132c043b · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:39:22.599556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T21:39:21.847501Z digest=sha256:6c37121ca6d1757cde40d3e64eb3eeae0a50c27ffbe15e7f1de553de3e2bfb93

Observation e4ff13ee-acd1-48ef-b3e9-333580ea32ed · outbound

This paper cites DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.852342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.852342Z digest=sha256:048de1d6681cbb86a2bdf1a7c786e381ab0cc4d7059c5f618a3f7c3fd8c3cebf

Observation 471f1036-d7b9-4a2e-bf3a-37cc8e5b47fc · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs OPT: Open Pre-trained Transformer Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.856983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.856983Z digest=sha256:9ad5721c377697fd9628ebd08b25a50289a7722f9d8b4a6dce3b87feaba2270d

Observation 8fb32034-b8fd-4251-82eb-09d432f8f22b · outbound

This paper cites an unresolved cited work.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.861654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.861654Z digest=sha256:c1d8dbeb91721ddd3d958a8706ab7d6f563917f143372dc1603902c27517f3af

Pith citing papers

Observation 58936284-d05d-4d50-9943-35cc6971894e · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:01.186306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:01.186306Z digest=sha256:b85465d56a752ac1ad56608965a713c19f68ee8c0fbbea36eb52c9e2b6ad1241

Observation a1defa97-b381-48e2-abbd-47a45d5d425d · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:02.510363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:02.510363Z digest=sha256:caa049c693c9467e9f7e058ba589536dc1a65f0ff11c087f66ff1b0b82a41125

Observation 0a3f0be8-4038-4198-babd-117a6477a653 · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.382448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:dfe38758675f0de78e5c54a46eee6ab73555e0576e24fc145c627adeef5e8bbf