Pith. sign in

Paper Citation Record · LEDGER

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

As of 14 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 28 inbound Pith citation observations for arXiv:2506.10967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10967 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:20:33.522423Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:49:29.556959Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:59:57.457886Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9580dc0-9e63-41c9-860c-842923ee40f6 · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.500266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.500266Z digest=sha256:58a3212ce11db3328d7244a5db794d89f9e738f8c0a563fc535e097a9efed3eb

Observation e25074c2-8e0b-46be-a3d5-4b0e7c0947ca · outbound

This paper cites Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:35.249313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:20:30.843591Z digest=sha256:eee4b4e4d9c2aef1b4e44fdaa6bc1cbe464105612daac6ba451580b91d254255

Observation bda22afa-70e1-4b24-b976-6ab55fef2e88 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.095679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.095679Z digest=sha256:afab2e033e780ed302806fcb814b0effe084913394d796379a75e666f745eb7a

Observation 75faae6a-b724-4b8b-abd7-d58a1dbf889b · outbound

This paper cites Mistral 7B.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.290255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.290255Z digest=sha256:32794ddc3aa1094d1a2ac5a4507d18e036407656e0bdaab87cde126bb0e7117b

Observation cb6f6bb3-cb43-4549-877f-ce54049d6d93 · outbound

This paper cites A diagram is worth a dozen images.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs A diagram is worth a dozen images

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.400104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.400104Z digest=sha256:863e1a9ee37858e1bc61f7f1bf6589d4633da668d0d7f03ccd2da0baca650b78

Observation 0b5a60a5-a5be-480d-9837-88bf333bfa83 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.581672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.581672Z digest=sha256:2fb401c245dae38ee7c1c1cce32851c171dcc771a2d964679e99639f0e649d07

Observation 880c77ef-a514-40ae-8dac-e2ee90fdcceb · outbound

This paper cites Microsoft coco: Common objects in context.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Microsoft coco: Common objects in context

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:34.975202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:20:31.664718Z digest=sha256:b2aa1f1c8d177b62500a315ce267d28ff9f165c85119429d1422b82e2c3d5e4d

Observation 731567cb-7de2-446a-b222-3b662f87a2da · outbound

This paper cites Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.753984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.753984Z digest=sha256:be86f4c0c657ebb5676f28f8140bd6b298d1485f1f61686491efde702d07efb2

Observation f46ba898-242d-4dcc-a67e-660f1cd1d0a9 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.837013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.837013Z digest=sha256:0eaf478a836bb57b381410fecdc3923bf710dbf63f0b5fbf85730756e0eb11d2

Observation c8783900-4263-4e24-8bf3-dfdd7c03543e · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.068161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.068161Z digest=sha256:4719c8ca5a0e64fcc1fb084d89437df7db53269a57e14fe6cf6de16b6d2725c4

Observation 5d13ddfd-0112-46ce-bfb0-8a3a9baa8099 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.245624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.245624Z digest=sha256:b6a8d4a168d5dafac7d5b5f0754dae8716ad2f3b049d62f96b8344dd9e44d876

Observation b416dba4-3258-4a0e-ae97-07e441622bd0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.325718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.325718Z digest=sha256:d289279805e9ef7f3eed1a30391208f581ab3bf7fd4df9257be2fd8551689c0e

Observation e2fd1dc9-808b-4d11-be3a-44643dea61ac · outbound

This paper cites Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.526592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.526592Z digest=sha256:b7718bd9f8eb114c8c4801e1a29e887318efd807376d898cd51b4b0fcb6e9a41

Observation 5a32adc7-b150-4ae5-b10a-6b9661ed0910 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.613673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.613673Z digest=sha256:4c03595ad30d85a064989a52b36b179fe0f69b1a9e3e2ce4c8c4e44e13f29d8d

Observation c735d4a5-be86-430e-9d94-702ab8595c14 · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.701129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.701129Z digest=sha256:7f43f2701ae19656bf84defd7b12e39c1a15c6ca6ec42228ec8133132c251c1b

Observation be676b47-8fa3-4528-acab-1b7fa509bc28 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.927138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.927138Z digest=sha256:ce0f4ef9c8d9c19ecb01b2863146ec22a5897a35f52941ea1329f4532db86102

Observation fc8edb99-3740-495f-8468-1737e9cc18aa · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.004390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.004390Z digest=sha256:c3f509261f1c73ffbf15f2eca1f205b35b318bf0c0a3cf0f553f2b01b7b7250a

Observation 26b7f326-3691-4767-a429-eddfeaa91aef · outbound

This paper cites Long Context Transfer from Language to Vision.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Long Context Transfer from Language to Vision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.079691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.079691Z digest=sha256:7a22bdc91ea8d973e50bfb52199c9bdea09e8215f8b516d6514ddc583b7c6e14

Observation fe9d77f7-1109-4755-be87-518e9ba954a8 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.148022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.148022Z digest=sha256:64fbcec9be1080ed505b973f48ddf1de63badcba702659418d2c0274545d9c22

Observation a8344eeb-4409-476e-8514-35f8472891fe · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:33.229755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:33.229755Z digest=sha256:7985638173ab5773d5a3f8f9ab61bba7f5d06db34346f5c71e72ac65c845757e

Observation c2123e3a-1bd5-40af-9385-258fcf6886da · outbound

This paper cites Appendix B provides some details of the experimental setup, including information about model architectures, evaluation benchmarks, comparison methods and implemen- tation.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Appendix B provides some details of the experimental setup, including information about model architectures, evaluation benchmarks, comparison methods and implemen- tation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:34.643823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:20:33.317956Z digest=sha256:1c8a02bf0cb48eee88db4a7e2b2fbf49cd0f42838437bd0fe337014ff95b3467

Observation ac6f7c56-d168-4c80-b242-6bc6aaf2d9a3 · outbound

This paper cites Therefore, the greedy algorithm runs in O(nm2) time.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Therefore, the greedy algorithm runs in O(nm2) time

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:20:34.442063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:20:33.433526Z digest=sha256:75ea5f08762e846f52cebead3e8f4cec13fe9cfcd9af7a0521fad89245f4a65d

Observation 82d897ab-868b-4802-9c48-46acc1b2dfde · outbound

This paper cites an unresolved cited work.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:20:34.277604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:20:33.522423Z digest=sha256:729e51237a6910d2fd7f5555068cc27cf4786340f6611a7acce584595f22aa2a

Observation 80118ccd-ae30-4541-b8f9-194c876cf8b5 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 1975

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.906806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.906806Z digest=sha256:165643f24e1c5941da659b361ee48902ed7f907e0fc7fcf1d18c877a5e94ff56

Observation 6e0c639a-dbd6-4e51-b060-7be707b167b8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.510710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.510710Z digest=sha256:64467a28eec5fb597e4df244ef95204553bbb53ed75d98ebf390818a4dbe4cdc

Observation 9b2a4f8a-0bb3-4df5-b1a9-9275074cf3d4 · outbound

This paper cites Qwen Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen Technical Report

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.562632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.562632Z digest=sha256:a1766471bb4cd8a50c0ef33940272abb1ae74eef2ea481f70eb84161ff727ed2

Observation 9fccbff3-c235-473c-827c-647ccd7a1c5a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.413526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.413526Z digest=sha256:95c81093b3fb7a31ee8621f0f8b5401c3a7ec090a3818c38dd2429fcfeb54889

Observation 38da4126-715b-4717-a507-f8506b4447bf · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.765313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.765313Z digest=sha256:2b809f2e3887821cf201a73dd93f23fce6878b1e4dde197a7f834584a247b261

Observation 8e4ee5fe-b3d4-4fb5-864d-64b5088904fe · outbound

This paper cites Similarity-Aware Token Pruning: Your VLM but Faster.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Similarity-Aware Token Pruning: Your VLM but Faster

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.175344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.175344Z digest=sha256:9fab2efdf63e2d36634140f97ce5a59a77debc1c533fa2739a1d2986612f77e2

Observation 2895cc3f-6db3-47f7-999b-8ece88d84bc0 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:31.982439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:31.982439Z digest=sha256:dbcfab777c17b7d492f072b1680dfc89a1db9f5cf3b8c255116635bdea7fc23b

Observation b7ffd2ae-9a9a-4e15-bd0c-5f1fa2a75316 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.949286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.949286Z digest=sha256:7f922b427703af22994bd61c5daf9a3a28de2b3dd8b747d9d9fe0e3cc0d953bb

Observation 480d693e-9f13-40e3-96a6-4f0a64979a92 · outbound

This paper cites Qwen2.5-VL Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Qwen2.5-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.635629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.635629Z digest=sha256:af08fc28a38ebdd08f7d6a06d7bb47e42eea896944ef2b21536108ca5f2fef46

Observation 8c99589e-8e7b-40ab-8eb9-0c9b25a35eea · outbound

This paper cites Mdp3: A training-free approach for list-wise frame selection in video-llms.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Mdp3: A training-free approach for list-wise frame selection in video-llms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:32.151078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:32.151078Z digest=sha256:d4860f9963b34fe910158bae359599f098130535727578f29501b323dd921721

Observation fd123cd7-c6c9-4391-93fb-7314899df7ac · outbound

This paper cites InternLM2 Technical Report.

Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs InternLM2 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:30.698008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:30.698008Z digest=sha256:71caced57d48763fa38e58b4e43c6bd290d567d3eb886b92e285cd10282e6cde

Pith citing papers

Observation 00c809f9-844e-4f80-8722-f6bc21ce394c · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:45.462835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:45.462835Z digest=sha256:ba01a67dc8ade1155eb9fe123923e32d89f36775a1e769ab26198a27f0a9970c

Observation d6e1b3bf-d94e-4c4b-959d-e59484535d8d · inbound

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models cites this paper.

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.450623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:37:35.450623Z digest=sha256:6f93431f46cb87dc05af7d7a00d6ca516093f976f93b364d5f429fc694e7d105

Observation fb5cc1b2-a8b6-494d-ace8-c05441757435 · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:34.271065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:34.271065Z digest=sha256:e365215d0d39083049d06c54d537f3f459b0308e7e8b8a1dafcf753893836908

Observation 2867c7e5-692b-4aee-9976-eee272cbf749 · inbound

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models cites this paper.

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:05.867186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:20:19.653806Z digest=sha256:97e87dd77fdd83d5ed9724f3e97d77a89b80fed94bb86afcbf1ae9b79116790c

Observation b952a8a0-3d83-4f70-b36c-5b1ea2f0bffe · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.417048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:ced37eac1d61dd062fd813c5120f92949d50ba8feaf0c9d1e9a960253b105286

Observation 5e597c98-ab8b-4547-8774-cd44bd08c998 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:08.455532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-09T19:49:28.591419Z digest=sha256:0c72c2871d2aab6a5c3e1cf90b4f1c56348376071082c91b48883574a78244cd

Observation 23d1383c-d06e-4817-b8b6-88c678f3ea48 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:19.376132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:16:31.918225Z digest=sha256:2e02f0b668040d9a1bf0ba2707328512497183ef75a2af2a0fb6e7deec2cbe62

Observation 3490a9f9-7ee5-4a16-8cbe-0a9c697dee77 · inbound

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference cites this paper.

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:31:25.772099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T10:28:15.293381Z digest=sha256:d133cb0d905b28a47cb3dcd0f9acf7f65c4e7f6343b42aadd7972f6c46b451db

Observation c50d254f-c09e-4cd1-872d-b44ea1beec21 · inbound

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models cites this paper.

Focus-then-Context: Subject-Centric Progressive Visual Token Reduction for Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:23:58.587932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T05:20:55.448430Z digest=sha256:26343dded128694e83d6d4049b2ebec1bbd7ee2673d89ba0657b737160916707

Observation 70eb3121-3269-4247-8ead-97d6e6805ba0 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:43:14.998270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T08:40:10.152344Z digest=sha256:04b38a9c58be125fabf6ee9f6a6e59afcb30e8240cdeea0f7fafa1b5b7c3015b

Observation bb008883-066e-4558-af78-5629603e4af6 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:43.228543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:43.228543Z digest=sha256:83b516b86418cdbd1d5449584f6f3356505fcc356d1851a2747ffefd18bf1d6c

Observation fda3f518-4f09-46cc-95a7-8aa849d3982e · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.963802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T15:06:22.102725Z digest=sha256:a8e51b1712665a5ebce95413a58e728be44e27f7ee1e6abae1e3ea9f056208fb

Observation edbc6b09-c160-4a09-9a4a-b6b3f11b6fc7 · inbound

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding cites this paper.

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:44:36.407579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T10:42:37.401221Z digest=sha256:98419fecbb9065eddcbab21f78056b4cadfe1bc2f5c5dfd2d18c6e1c7e46425c

Observation 722a29f6-c5a9-4c48-847a-f1da057ebcbd · inbound

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models cites this paper.

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.715894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T11:04:54.171291Z digest=sha256:c1546d883e6770e94b59418284a52595b671ceda01143d95e26c69e62af4dbe8

Observation e6a61e6e-7b9c-428d-a4cc-5f379bf31ec5 · inbound

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models cites this paper.

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:59:57.459369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:04:20.802531Z digest=sha256:02a9bdcd440f8e9b30a27fa0994ec8b1fa56d75becbd1c4956ee80c9d350d14f

Observation 3d777fc8-3027-41a4-9912-24a5f4ce079e · inbound

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference cites this paper.

TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:09:53.798585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T04:24:38.917137Z digest=sha256:f54ac51a7a86c4b0aa6dcff2b797f07d0677217b497b01a82b1a1e0345de283e

Observation e7538e8f-8d90-4758-9d8c-e24379874ebc · inbound

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning cites this paper.

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.406972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T14:47:35.377391Z digest=sha256:2a8814c98fa0bd277056e1d431f560374c2a90d3700c57a9f79d0e527e5dc17a

Observation af6ec779-02cc-4cd9-9823-eff642eca4cb · inbound

RADIO1D: Elastic Representations for Condensed Vision Modeling cites this paper.

RADIO1D: Elastic Representations for Condensed Vision Modeling Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:29d2e922104d712d1cd606ae5c7c534476365d8e9157e095509824293f39aed5

Observation f15b7aa2-54f0-49dc-a303-7d5f893e19aa · inbound

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models cites this paper.

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T22:32:25.110039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:32:25.110039Z digest=sha256:2c0eb018a64713faa7affe87e02090a7e2a5f1261d72e9d7d3c13b0ac37c6756

Observation ab72c83f-49fb-468a-906a-74a8aca1be14 · inbound

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models cites this paper.

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T04:21:09.449182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:21:09.449182Z digest=sha256:3c408639074579f13cabe0e131461c837728a3a6634e61add92a90ac7750d467

Observation 1d712359-c0ab-48d8-8081-4f12ab56b5ef · inbound

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs cites this paper.

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:50:32.506198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:50:32.506198Z digest=sha256:d49aa6c50c6c21f2f45736e8f3685e69a1b1734f8d08f5a4b97dd3c32fbf1134

Observation 787a3619-c726-44eb-b279-9170356085bd · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:18.339215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:18.339215Z digest=sha256:975623eadd60d6039e70718e1382d61a9b2c1ddfffe447a81dfa8a133099a9a0

Observation c940e4a4-e27a-4e34-b139-3498d8d14cc3 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.622068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.622068Z digest=sha256:90ec8d42cad21fd3443924b1c1bfedea03cd9c953e7631375c161ec6015bb76e

Observation d99230ee-52b9-41a9-911f-f0300818985a · inbound

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models cites this paper.

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T01:26:50.833164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:26:50.833164Z digest=sha256:0ad1f9b3f973fa554ca3961103264f8b6ad0aa1e122009649239071cb95016f9

Observation 814f8c48-6d41-4d5d-8183-f3ad966214e2 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:57.677655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:57.677655Z digest=sha256:5d928ebffeeef7d4f2256285c037820ba7195de6ca85e1ae3c71bf96e8f60102

Observation b32330f0-3ef6-43b0-a933-40627f1db851 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:49.179772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:49.179772Z digest=sha256:76a536d10fa700441b9b33652b39e0a24f544680d5c38f02a9b83a33bcdf96b2

Observation 9355748a-d097-48c7-a5bc-667d093bcbcd · inbound

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding cites this paper.

Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T18:36:14.731592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:36:14.731592Z digest=sha256:d887a5098c21bdfb35940e6af11eb4c21e15ef1ca6461edda9c1999b6fed53a1

Observation 0c2c288f-8d7b-460d-98b2-1fa8c38830dd · inbound

Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers cites this paper.

Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:49:29.556959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:49:29.556959Z digest=sha256:24124d73c7011ec37ab82fcba7bbf5bcac8b28f4107f47c769ceb03a4295ecc8