Pith. sign in

Paper Citation Record · LEDGER

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models

As of 23 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2507.20842.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20842 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:18:06.616164Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact1
  • verified fuzzy51
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9bf00e6c-7602-4633-96e4-f9a9d413c4b0 · outbound

This paper cites HiRED: Attention-guided token dropping for efficient inference of high-resolution vision-language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models HiRED: Attention-guided token dropping for efficient inference of high-resolution vision-language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:00.011411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:00.011411Z digest=sha256:8bac4f5df65572901ef95b92b69a6080d2678137eaea148c426f643b4ee47430

Observation c1485921-bb5c-4034-9de7-180bcad8880c · outbound

This paper cites Qwen Technical Report.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:00.053585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:00.053585Z digest=sha256:710e3b1879dce475b814cd0e9dad54cf91c4764d7e70fce67500b5a79e7195bf

Observation 87ea521d-1dbc-492e-9cda-e099ef20f5d9 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:00.134131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:00.134131Z digest=sha256:69e192cd3961cc12511933329de6a3d0269c03589ee3eaa9588dd049c188b0a0

Observation ead1432d-477a-4e49-9bcb-f4c9afa33128 · outbound

This paper cites Language models are few-shot learners.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:00.225142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:00.225142Z digest=sha256:4e1d6c8dacd08326b4912f7fe2f5da278298219a0ebe1b3bc7a0f424cef7f61f

Observation 999b5cec-07e7-4562-9256-9980b3894810 · outbound

This paper cites Matryoshka multimodal models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Matryoshka multimodal models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:00.280685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:00.280685Z digest=sha256:2c3a9edf81deb454fddd8b7cc413083f5cb036d759b6627b701bfe77b4d2ba76

Observation a357e339-40d8-4090-b109-d7355d21abd4 · outbound

This paper cites Open-LLaV A-NeXT: An open- source implementation of LLaV A-NeXT series for facilitat- ing the large multi-modal model community.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Open-LLaV A-NeXT: An open- source implementation of LLaV A-NeXT series for facilitat- ing the large multi-modal model community

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:00.380288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:00.380288Z digest=sha256:a1021387ecdbefa2d0244cebcde8a6a5d6cc059ce215e885bfff7e8cf48ed8be

Observation 31abbaa4-7649-4f69-9f56-0f1ac3899f23 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceler- ation for large vision-language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceler- ation for large vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.864395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.452458Z digest=sha256:9fcf0a4c23b2f2292fdaa88237ec791143c9cc84e64437341b4b72641afb1845

Observation b63ad6fe-d42f-4d92-8267-c952b4650498 · outbound

This paper cites How far are we to GPT- 4V? Closing the gap to commercial multimodal models with open-source suites.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models How far are we to GPT- 4V? Closing the gap to commercial multimodal models with open-source suites

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.853260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.540214Z digest=sha256:2a655728b0c46802e4390d4bd4f6711b13537f5e03e18d99efb846795534b29d

Observation f1a7e27f-3eba-4793-9aaa-97f459c7bee0 · outbound

This paper cites Intern VL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Intern VL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.841852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.606313Z digest=sha256:397ffb2246aae8302691e079839b7b919029ca030b68f7d9ddf571043a1ca56e

Observation 5f33187f-e21e-4350-9a3b-7088eca57be7 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Gonzalez, Ion Stoica, and Eric P

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.829581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.736571Z digest=sha256:d1277944efecbc6069d93f7d42d7ca5b2aa97b85553d1b41e07868bae6498431

Observation 37646758-eeef-4934-a2ef-29c79e748467 · outbound

This paper cites InstructBLIP: Towards general-purpose vision- language models with instruction tuning.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models InstructBLIP: Towards general-purpose vision- language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.817783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.845969Z digest=sha256:a5f4c59972e9a9cdf1230d9b63a6517a903b41100a2dba0a786b410b16b231f8

Observation 0c6b303e-068b-47bf-b7c6-3d6146ff82f1 · outbound

This paper cites Transformer-XL: Attentive language models beyond a fixed-length context.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Transformer-XL: Attentive language models beyond a fixed-length context

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.806432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.931596Z digest=sha256:690f8314406fd7fb996ade4651ccc58c55bdac56d2d2920e442cee02b5bdf7b3

Observation 5a9f1c54-b5cd-455c-9b18-fbe011776efb · outbound

This paper cites Prune spatio-temporal to- kens by semantic-aware temporal accumulation.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Prune spatio-temporal to- kens by semantic-aware temporal accumulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.794124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:00.996225Z digest=sha256:3acc3724f6bafd8fac80c7f422214e4c1b13680fd940e035fac21c9b1bf6c5d7

Observation 7fac3077-35de-432e-899d-6aaaa01bad56 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MouSi: Poly-Visual-Expert Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.081840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.081840Z digest=sha256:8daa83dbebfb7c07b8a383ff0234493b9416a70a068374d1e751faa2fab486d7

Observation 67d42725-c1e2-4d6e-bd75-eaa9a483f1cd · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.159137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.159137Z digest=sha256:403b0cc3048d0208a548607545d3353e13cb9fc82965cca6fa5d739fa572f141

Observation 88619815-830c-4a39-a3b3-44a8887b3f14 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Llava-uhd: an lmm perceiving any aspect ratio and high- resolution images

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.783034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:01.245744Z digest=sha256:ccbc1f6711d1f489331618cd567b3499fa81d62bd52021798cebb3e8497706f1

Observation e7f6c151-3316-4943-9fc6-eb171a65499d · outbound

This paper cites Re- thinking token reduction in MLLMs: Towards a unified paradigm for training-free acceleration.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Re- thinking token reduction in MLLMs: Towards a unified paradigm for training-free acceleration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.321659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.321659Z digest=sha256:38bfc8de801df386ef07778f80ad9e355c533b1c4b3f483e34a55517dc4df9b7

Observation b21e7469-2038-42d5-85b6-e5790f74686a · outbound

This paper cites Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.418539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.418539Z digest=sha256:62f8474bd0621981d65c32d83e56530e28b1ae449a3af5e6c774b6088941c220

Observation 26d5b24b-0dec-42fb-960f-03fc79cf751c · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.488615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.488615Z digest=sha256:26a8af05635ec22a87c27269e02a0e521c0aed9e2ec1cc686adea827119ae02a

Observation 1bc995b1-b506-4ba3-8c1e-9a3d6fe5ede7 · outbound

This paper cites iLLaV A: An image is worth fewer than 1/3 input tokens in large multimodal models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models iLLaV A: An image is worth fewer than 1/3 input tokens in large multimodal models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.597860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.597860Z digest=sha256:da932ba62d3e367769d7b2cd957de1b2493964440b9d20a8936b62ba4a3c9341

Observation 3898d8c7-9120-4921-a83b-c31e29a5a7e6 · outbound

This paper cites Matryoshka query Trans- former for large vision-language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Matryoshka query Trans- former for large vision-language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.772607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:01.664723Z digest=sha256:b6db21cb166e711dfebf402973bc026ee381e5967bc4c1f7add2585d61c4b827

Observation 5e583f37-876a-4183-bdcd-00b3f57fe84d · outbound

This paper cites IVTP: Instruction-guided visual token pruning for large vision-language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models IVTP: Instruction-guided visual token pruning for large vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.761920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:01.730301Z digest=sha256:855b9b903d94792d7df1a4d75b54b92d9c4f845adc06e0509ab82992ea46182e

Observation 4842f778-bffd-4c00-becb-7104d57a6a65 · outbound

This paper cites GQA: A new dataset for real-world visual reasoning and compositional question answering.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models GQA: A new dataset for real-world visual reasoning and compositional question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.751293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:01.830501Z digest=sha256:a70cd54c01f0a1c28d89d506d82d8340704e6114439d48979307cf88433517c8

Observation 31d9f30d-3254-4f1c-8878-c08df639758a · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.907280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.907280Z digest=sha256:d5083e21b36f8e98ba4f8e769811fbe8a1d1ad67a8c56b3d37e82b404d092048

Observation 41ffb6cd-4e62-4975-bb51-ac4ff2dec3bb · outbound

This paper cites FoPru: Focal Pruning for Efficient Large Vision-Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models FoPru: Focal Pruning for Efficient Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.973663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.973663Z digest=sha256:42a59a6d6ac4986b92c461cd2bd6cae61dcdb97523cafae822dff428a0f9b8aa

Observation 76f4f518-fae1-4b0b-b269-393b8a11e5a4 · outbound

This paper cites Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.740474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.040705Z digest=sha256:cf83c886930bdabe19cb179391357ce34162ac79dd3cfa02bbbc082b98a34bbf

Observation a4427d5e-302e-4b23-ae12-c54ea29f5961 · outbound

This paper cites BRA VE: Broadening the visual encoding of vision-language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models BRA VE: Broadening the visual encoding of vision-language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.729482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.136557Z digest=sha256:df499d6622fbccb7bc119b1a623e922b9feba632723e2f1443dd84e5d22ae163

Observation 8dd021e5-ca85-4eb9-b5f1-346cd0c4358e · outbound

This paper cites A diagram is worth a dozen images.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models A diagram is worth a dozen images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.719364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.232705Z digest=sha256:cdba71f74ea6ff097ba7f5c774321ccbf403b6a2a652bb9226c45da6b4080049

Observation cfcabdef-1f29-4320-ae73-b6344af3bff9 · outbound

This paper cites MoAI: Mixture of all intelligence for large language and vision models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MoAI: Mixture of all intelligence for large language and vision models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.709154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.308242Z digest=sha256:53c79849d4a1dfe5e5ae596b40b4e40a70ccdacdbd866b457458658599be431b

Observation 55b7590d-b294-472b-a4b8-5a07d4faaf0f · outbound

This paper cites Pix2Struct: Screenshot parsing as pretraining for visual lan- guage understanding.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Pix2Struct: Screenshot parsing as pretraining for visual lan- guage understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.698366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.410969Z digest=sha256:8689b6083a4d58bd7a7f54f98c10beb248039faa0af8b2c373c61e1127215065

Observation bd6385b9-6310-4559-b82c-97769486e6ed · outbound

This paper cites SEED-Bench: Benchmarking multi- modal large language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models SEED-Bench: Benchmarking multi- modal large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.686313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.461864Z digest=sha256:7cf8b5167e90a7b3fd39faf432bdbbd95d0fd8bfa1502dc55a2790d3dba8747d

Observation a1defa97-b381-48e2-abbd-47a45d5d425d · outbound

This paper cites RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:02.510363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:02.510363Z digest=sha256:1084fda20bbb025c2890f0245e11b0779ca1e4adee7ef16f6f4ffcbbb2e986a1

Observation 040ac8eb-4830-45eb-bd0d-516ac390e385 · outbound

This paper cites TokenPacker: Efficient visual projector for multimodal LLM.International Journal of Computer Vision, pages 1–19, 2025.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models TokenPacker: Efficient visual projector for multimodal LLM.International Journal of Computer Vision, pages 1–19, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.675355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.584449Z digest=sha256:d7f4f13e9f925692ef215305bb9a82e6d78335d3085dda7a8395d19e7033c5e6

Observation 41ad1dc6-0ca1-4e89-b11a-e48559200bef · outbound

This paper cites Evaluating object hallucination in large vision-language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Evaluating object hallucination in large vision-language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.664717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.667810Z digest=sha256:395e24802860f9f032c621abede9677621e4b49cbe7c6ba326a17d3c57bd9899

Observation cdad8285-6e6f-4c2a-9e95-6cca04aac393 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:02.711071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:02.711071Z digest=sha256:ac5dbcb17b5a6c8f197115dfdc06e1180d346de858c9f5e48121532a6bb86fb6

Observation 3b6694fd-4a77-40fe-9369-c16178c0e4f6 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.654102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.795267Z digest=sha256:2d600b188ec521494c7bd6d82f30620b1344f91000da24e110a8189b0a942e99

Observation 62166a73-9f04-4031-9caf-5f17fd56f6d5 · outbound

This paper cites HRank: Filter pruning using high-rank feature map.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models HRank: Filter pruning using high-rank feature map

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.643503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.859823Z digest=sha256:40cffa03dd65d8d6f201febb572106fd1c3bafe6f59ca370fd6cd9f5dbd1eb01

Observation 6d34a01d-7111-4e9d-97c4-c004e9b98230 · outbound

This paper cites SPHINX: A mixer of weights, visual em- beddings and image scales for multi-modal large language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models SPHINX: A mixer of weights, visual em- beddings and image scales for multi-modal large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.633793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:02.925986Z digest=sha256:527f5215b2a793dacdc0acf80fa1f9f63f339fc77f8255266fdcd78f40b22cef

Observation c7edf160-0946-44c9-beb1-fab7aefc1333 · outbound

This paper cites Boosting multimodal large language models with visual to- kens withdrawal for rapid inference.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Boosting multimodal large language models with visual to- kens withdrawal for rapid inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.623493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.013426Z digest=sha256:343c6a363d229d1439cbb99fd13bbc5256359efa0bf1a7add11b6d53f1a4ae69

Observation 8ee48e26-ab84-46ec-8f53-0fbd3e941446 · outbound

This paper cites Visual instruction tuning.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Visual instruction tuning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.612633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.129577Z digest=sha256:d51c74e364ba19fd707103bef869fe5528c6e7e3983a8f60a15f468e1a107f0c

Observation ca253b1e-7681-4ec9-9b34-a803633737dd · outbound

This paper cites Improved baselines with visual instruction tuning.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Improved baselines with visual instruction tuning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.601485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.211066Z digest=sha256:f24ac7a39afba6c83c7894ce476cc7cdca1479ccb8b3e40a3e86d95c8227d04a

Observation abbbc795-999e-4f68-8531-01db123ee236 · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models LLaV A-NeXT: Im- proved reasoning, OCR, and world knowledge

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.591311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.273922Z digest=sha256:e9abdafb031363a17bd2765d2befb80e88a371ad9750a0d2e2c508096ec333c0

Observation b5ddd117-2279-42ae-a3d2-fd28eafd8d09 · outbound

This paper cites Prismer: A vision-language model with multi-task experts.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Prismer: A vision-language model with multi-task experts

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.581568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.360116Z digest=sha256:9951b3d292037d62a41dbfd2811235e8364592d6f8f0150d558cd21f9f09232b

Observation d0eca418-d976-4faa-adc6-75a5568b58d2 · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:03.452277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:03.452277Z digest=sha256:c4130913a6673720c20b6b6fbc3568b92b9c640ad6bc7c7a596aeb3cacc11a4f

Observation bce14f18-8693-40b5-a6be-ebd619080353 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:03.510688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:03.510688Z digest=sha256:53a2ec7d190c41d1ab2339c82aeeca55020df9edb235e674092bfda4e67ac73a

Observation 480f02da-cf24-4459-a1fe-aa3824c592b7 · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? In Proceedings of the 18th European Conference on Computer Vision (ECCV), pages 216–233, 2024.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MMBench: Is your multi-modal model an all-around player? In Proceedings of the 18th European Conference on Computer Vision (ECCV), pages 216–233, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.570794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.599690Z digest=sha256:2a7011622b6f94174453ab7a14d4df34a0add51641089e58ba10dd8239265f50

Observation cee4870b-154d-4b25-95d9-b48a0270594a · outbound

This paper cites A ConvNet for the 2020s.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models A ConvNet for the 2020s

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.559619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.700886Z digest=sha256:ea5d6d3d633131c0973d13d103862131e8c22dd16b13ef9b7450b05d0f12a703

Observation ace58567-937f-457f-bebb-715a9d7d50e2 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:03.786752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:03.786752Z digest=sha256:68d53d02db4137beb4b693336e6d64eab50488fb4ef3dcf95a1139b609e7002b

Observation ae65ec97-940f-4ab8-b051-589240f5ab3f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.549553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.851817Z digest=sha256:50ec5384787873a3aecf4cf67a5832cfc876222230811da01f120c21fc085a24

Observation 2c83b4ec-4b18-47d4-b38c-59f88bd7146e · outbound

This paper cites Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Feast your eyes: Mixture- of-resolution adaptation for multimodal large language mod- els

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.539139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.912455Z digest=sha256:71f15f47f386ebc36fc18e2ac78f2b59207f56280bac2cf3f11fd27f9fed52bb

Observation 70a081ae-ac1b-43fa-8d46-07caadfed91d · outbound

This paper cites OK-VQA: A visual question answer- ing benchmark requiring external knowledge.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models OK-VQA: A visual question answer- ing benchmark requiring external knowledge

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.527940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:03.967417Z digest=sha256:28aab0079b64a434c43121b52747e49c6cd51ac0bff8682ca0207c734422949e

Observation 0937d56c-abf6-49c9-85fb-0718abffb9c1 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.515372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.032069Z digest=sha256:358a49a8f73de309617680d5eef2fde926fe08fdcb9017df8cac5366af703be5

Observation 7ab13193-4dec-45a0-a24b-f1aebadadae2 · outbound

This paper cites DocVQA: A dataset for VQA on document images.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models DocVQA: A dataset for VQA on document images

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.504388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.097524Z digest=sha256:8022d473eedbfe3125b0fe51975085fbb6152b290cf55f4378857feff7f5d6b7

Observation 5066e26c-01aa-49ca-a91f-0f801504a090 · outbound

This paper cites MM1: methods, analysis and insights from multimodal LLM pre- training.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MM1: methods, analysis and insights from multimodal LLM pre- training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.493930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.196959Z digest=sha256:aae6b6b3f21aaa96452d1d877cef11b7a37770f94cff2a2a7d6aff2d65fa4091

Observation 6da9200c-4623-4c24-bd9c-bdd68b00b26d · outbound

This paper cites DeepStack: Deeply stacking visual tokens is surprisingly simple and ef- fective for LMMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models DeepStack: Deeply stacking visual tokens is surprisingly simple and ef- fective for LMMs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.482988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.251891Z digest=sha256:6ad939e4d635285465b4c8c04b1801f8f1be27182c6bba7aad1ccc41fc363e8d

Observation 1181e175-b624-4b44-8bda-bc0c2878e5cf · outbound

This paper cites Learning transferable visual models from natural language supervision.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Learning transferable visual models from natural language supervision

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.471349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.341414Z digest=sha256:adccc619db8ceff10a32eabbee757ffae739b5e7c6aa70b29f59f01b8ac73770

Observation 6ae5eacd-d64e-4b0a-8573-1e8e45d967d9 · outbound

This paper cites LLaV A-PruMerge: Adaptive token reduc- tion for efficient large multimodal models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models LLaV A-PruMerge: Adaptive token reduc- tion for efficient large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:04.431699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:04.431699Z digest=sha256:04a0b1170f145fe015e5f3987c6080f4ba1730a76ccdb1a9efba61a773d386d7

Observation 3e60f03f-2a83-4af4-a291-5bbf33ee083f · outbound

This paper cites Eagle: Exploring the design space for multimodal LLMs with mixture of encoders.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Eagle: Exploring the design space for multimodal LLMs with mixture of encoders

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.460870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.492453Z digest=sha256:3d350638add616d3d25b5041d534ef8e7aaa6d98a4453635df2a9e6fe1cff995

Observation e706f53b-662e-4af2-abf9-bae93973a430 · outbound

This paper cites Towards VQA models that can read.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Towards VQA models that can read

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.450379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.591655Z digest=sha256:02c022fd7d55e3c8ac0860c1d8095dc0f6e2da2d318fb1cd1f115896e418fab6

Observation b0249041-ae47-478a-8c72-4f287eb73c4a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:04.652268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:04.652268Z digest=sha256:0fd0d4a673f559b6e865c5d3d3dd16dd9356249a8047edb51dbd885c7167a93d

Observation e013054e-98ce-486e-8c52-63808d777c61 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.438994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.720209Z digest=sha256:58c14ccbe13556bd9383516b68a009abb05301315f77606230dc7e7d887d26c2

Observation fb96b329-82b9-4b0c-965c-452cf3350ee7 · outbound

This paper cites Eyes wide shut? exploring the vi- sual shortcomings of multimodal LLMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Eyes wide shut? exploring the vi- sual shortcomings of multimodal LLMs

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.427427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:04.803313Z digest=sha256:91f1dd04615513e24c513af86d797e8ab0572a48aa38251d18485de60d3de10a

Observation 282b5c52-48b4-4e8c-881e-9ce32466a1d7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:04.867176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:04.867176Z digest=sha256:10646872f07a6954340c9d8180e4557e04148c45f2cb8508f4599a79ee48756a

Observation e9cfad62-f673-4087-bdbf-bcec45303e3c · outbound

This paper cites [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:04.956498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:04.956498Z digest=sha256:b76a02f726f23ce322cae8e83fd3cdd3202580d97d98610be82c895218df7e2e

Observation d91cfad3-cda2-4078-8ccc-637e3f969956 · outbound

This paper cites FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models FOLDER: Accelerating Multi-modal Large Language Models with Enhanced Performance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.021161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.021161Z digest=sha256:8e3b95ec3e28136610856eec443cb45906a8159c72f7a787c1231c5c69276672

Observation 9a0cabfe-691d-4cfa-8656-f7ba10c3efb5 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision-language model.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Vary: Scaling up the vision vocabulary for large vision-language model

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.416656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.079387Z digest=sha256:374cf9a9dfa8fe221caa0ad90f358711bdc5208af5a23f781f5b7a036aa4d604

Observation 8e692868-cfec-4a7c-88a0-1d30e254f151 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.173548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.173548Z digest=sha256:8ba3fd6576d8ed2b3e51c6c1f69ffc8c8d2f39f47339adb831dfaf037e602d81

Observation 5d4cf09f-6e98-4990-a85f-6589f2c41054 · outbound

This paper cites Mitigat- ing hallucination in large vision-language models via mod- ular attribution and intervention.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Mitigat- ing hallucination in large vision-language models via mod- ular attribution and intervention

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.406074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.264656Z digest=sha256:b6f25cb8fad884ebf773799b09522728be12e88f02cff321751c25d5539f36f7

Observation 77bc4401-8c03-49ac-8110-7207b35e9308 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.305866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.305866Z digest=sha256:818568a18d98ddf8de90a7139f62584abfd300813dcf0450da46833d4c8b8b73

Observation d637267a-2472-4098-a01c-724fea3ddf31 · outbound

This paper cites mPLUG- Owl2: Revolutionizing multi-modal large language model with modality collaboration.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models mPLUG- Owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.395012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.361829Z digest=sha256:21d678a386e64e67d6c4dbb7afc7bcfbbfdd548def87bd94a8b4a66bb3015877

Observation 4331c16f-9315-4166-b6d1-784332056de9 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi- modal large language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Fit and prune: Fast and training-free visual token pruning for multi- modal large language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.383458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.459651Z digest=sha256:b3a30642404c5d6dcd678aae1a8ac4c36f10adcd924c59d0a6f3ea16405faeba

Observation adce2a7b-2be5-4439-a3e7-a01e4d4d0829 · outbound

This paper cites ATP-LLaV A: Adaptive token pruning for large vision language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models ATP-LLaV A: Adaptive token pruning for large vision language models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.372227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.553886Z digest=sha256:feab4e3eae4fd5fa684586260d572ff85d117fcf1dbfae44dfc0db560e248dcc

Observation 37c35d01-b84b-4c1d-a36b-98c03346ace0 · outbound

This paper cites V oCo-LLaMA: Towards vision compression with large language models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models V oCo-LLaMA: Towards vision compression with large language models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.342800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.613592Z digest=sha256:aafdeaedaf72387f7ced92f2017256cada31a1935745785e63ecdc0d8b8be4a7

Observation 841174f8-18ae-47fb-be48-393b559fb4a1 · outbound

This paper cites Lifting the veil on visual information flow in MLLMs: Unlocking pathways to faster inference.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Lifting the veil on visual information flow in MLLMs: Unlocking pathways to faster inference

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:08.184101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.707052Z digest=sha256:8467c0efe4b47758ce2b9c402f6c479dcfa49099c6f119435faa75dc70d155fd

Observation 18328f71-a4db-42db-a401-0e22267c7c09 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.802359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.802359Z digest=sha256:355f4b833d80e7fe933548071bcf174fc1e41920fbb666d8a2e0b1dd3eeef74e

Observation 20c3e7d9-921a-4bb0-8a53-73fcbf1558ff · outbound

This paper cites LLaV A-Mini: Efficient image and video large multimodal models with one vision token.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models LLaV A-Mini: Efficient image and video large multimodal models with one vision token

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:07.952000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:05.877868Z digest=sha256:6209a9fad187ff66e5d0a4cb4c4bdc83d8253a3115fa8e238130c0dc8477e476

Observation bacc4ae4-e3ee-494c-9aba-71a237fc9cf9 · outbound

This paper cites Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:05.927482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:05.927482Z digest=sha256:9f563f0f29a7da1182ca6d063ad45692345d2b6556e0b449420637adaf5d350e

Observation 6c1df209-b7b8-4371-b01c-df36f4f61cc0 · outbound

This paper cites SparseVLM: Visual token sparsification for efficient vision- language model inference.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models SparseVLM: Visual token sparsification for efficient vision- language model inference

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:07.783244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:06.027505Z digest=sha256:f1108d6a5c83ea8c950cbc93228f8a3b1377b5035c23d7685518f302ca6bbef3

Observation 342ec6a2-1905-450e-b7f2-2146ea7721b6 · outbound

This paper cites Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.117105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.117105Z digest=sha256:b26922f3e6043a8356d32797108ece1350925d4691ae13a6f85a26460b1a71f3

Observation 3ceeb469-71fe-446b-85ce-77e1fbf36a68 · outbound

This paper cites Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:18:06.881928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:06.236925Z digest=sha256:57ec573c088ba163d797ed14b74610ed54e54503a554fb72b506f9a2b5eaf43d

Observation 7d65a803-0200-4e6b-aa91-044e0e65175b · outbound

This paper cites AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.370053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.370053Z digest=sha256:594c4eaf1939ea646730c356f1435654dbdfb58e297b811bb31fa17a6c5c761c

Observation b3bbd1c0-7f4e-49bb-aaaa-2dae09d258cf · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.501322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.501322Z digest=sha256:f3e84de4245186692b2dde55ece4fb72cdc0112d4aa3db94f3512bba813476b6

Observation 01988274-4b3b-42ef-bd1b-1caf36fbd444 · outbound

This paper cites FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.528428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.528428Z digest=sha256:7c0b0a87d3b5e4ed212dd4009abbf2800361d62ee97f65937688eb0fa68e2907

Observation f050d5af-14cf-4662-9620-6bcb4201c868 · outbound

This paper cites MoV A: Adapting mixture of vision experts to multimodal context.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MoV A: Adapting mixture of vision experts to multimodal context

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:18:07.614354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T13:18:06.616164Z digest=sha256:879026cbf596a2b9d1417af61601423c6d4acfbe06ed5e0a51044c3f7b38c4f6

Pith citing papers

No inbound Pith citation observations are available.