Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:14.766061Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 30 inbound Pith citation observations for arXiv:2411.10803.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:14.766061Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:31:12.932387Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T14:09:53.785860Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 90864fc1-f830-4d83-8382-7eb430667e2e · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba16e14-8183-487c-8ee0-98d720c1287d · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf07494-df73-4d0f-9a8e-af7dab5165dd · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec4711df-b311-4043-81cf-720c20357eac · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Vizwiz: nearly real-time answers to visual questions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59ce88ca-8c8d-4533-a298-17232e8455b6 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Token merging: Your ViT but faster
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9f75e7-9399-4572-8c87-08ccdb358a52 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Lan- guage models are few-shot learners
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 63edf87f-65b1-43e9-849d-4134efc7fae6 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74fb69d7-221d-451a-9d70-0359b041a938 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Qlora: Efficient finetuning of quantized llms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0fb3af-f239-4543-98bb-7e7a38d06a3c · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7e8d820-3121-4b8c-b1c1-d96c6da4c693 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7632f18-9eeb-44a2-ae4f-f5cc592fb010 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7900d443-6024-407c-a331-a9f51ed33260 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Cogagent: A visual language model for gui agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d2aa05c7-75c4-4c84-b425-2ed6babad4e0 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Lita: Language instructed temporal-localization assistant
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 21fe686c-4fec-48f2-ba7a-fae3718c4987 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model GQA: A new dataset for real-world visual reasoning and compositional question answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f008554-ad28-4410-be93-66d9c3929c91 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 546f1cdc-6d81-4930-9172-add9e7692d8f · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Turbo: Informativity-driven acceleration plug-in for vision-language large models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5048f24a-5b5d-4765-afc9-a1772d323ab7 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e524ec4-cee8-46f3-9228-d64dad492558 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model VideoChat: Chat-Centric Video Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3acf5675-711b-4a42-8e66-45b338422fa1 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Univs: Unified and universal video segmentation with prompts as queries
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b1b4e3f4-7f34-42dc-ab74-a7cfcff93cca · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Not all patches are what you need: Expediting vision transformers via token reorganizations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b8202f7-6a04-4e51-bdac-b5119199b81a · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82f5b42-a0ba-4bb4-a49b-1ebea5a60523 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Improved baselines with visual instruction tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d920a1f9-03ac-4574-9bb7-3b8b9f077175 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25aa0eac-4f08-4ec1-8bb3-d2c2c31cfad5 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Visual instruction tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120106a2-2cff-4f79-9ad0-fa0feedcb873 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model World model on million-length video and language with ringattention, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ffe86533-16a5-4952-ad35-51c99c3969bd · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Sparse-Tuning: Adapting vision transformers with efficient fine-tuning and inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2784e95-37e2-41c9-9435-c82d1d70b763 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Mmbench: Is your multi-modal model an all-around player? In European Conference on Computer Vision, pages 216–233
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5926270f-a6de-4566-90a6-1453b9c0c142 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109af63e-a835-47c7-a941-64aef8be0380 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f795fdce-73ee-491b-8a09-57add7987980 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Language models are unsu- pervised multitask learners
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c4afb1-a030-4e76-9c51-a5a55d6471c7 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Learning transferable visual models from natural language supervision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f523732-1f7d-4d7c-87ad-1d0807245fc1 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Direct preference optimization: Your language model is secretly a reward model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 218ee6ca-d0cd-406d-afa8-ab4659e9c1db · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a515ffd-d675-4c9d-8312-df2957c7c3db · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Towards VQA models that can read
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d69b5c5-551f-417e-bba5-93bfb65da4d0 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Gemini: A Family of Highly Capable Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad812a0d-efe5-44e1-9170-5edaed08b250 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8687ff-1aa5-46c4-8d04-46ad9834cf6b · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Attention is all you need
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8f65aa-c970-41fe-a198-dcaa65a8dfb4 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model CogVLM: Visual Expert for Pretrained Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd05ada-3e01-4ee0-8061-a1c98917708d · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Can i trust your answer? visually grounded video question answering
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03ac8755-664e-4e3c-add4-7f33b7f4c244 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8cbbe51a-842f-47a5-888f-e5c9936cb268 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Zero-shot video question answering via frozen bidirectional language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c966751d-677d-4b39-983b-438a783cffc4 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30fd9868-b180-4977-b031-44e152b3eaa3 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Tree of thoughts: Deliberate problem solving with large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68427e43-9baa-4705-ba2c-1513d9ef9ee4 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Ferret-ui: Grounded mobile ui understanding with mul- timodal llms
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09742380-dbcc-4ccb-9a39-312369a06e4f · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Magic tokens: Select diverse tokens for multi-modal object re-identification
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3ac0fd5d-705d-4259-9f12-a93c0077f143 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model Llama-adapter: Efficient fine-tuning of large language models with zero- initialized attention
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a467bd66-ca9b-4b1a-99ec-a406103992d5 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c738d7b-faee-466a-92d7-1fe9f50bc749 · outbound
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f142b27-cac9-4666-a592-a5452ab23db9 · inbound
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 83320cf7-427f-47a2-9850-377bd67ebdbb · inbound
Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 274
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af17887-9523-4eef-99ba-52009acf4f1a · inbound
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba2346a-6f98-45c9-a41c-80c1c81e0659 · inbound
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b97be4f-8fef-4016-847b-682b060c615b · inbound
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 731567cb-7de2-446a-b222-3b662f87a2da · inbound
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762b43ab-b672-4b1b-b4ac-cbd32c5250f7 · inbound
GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c4a900-5f2d-47c0-821f-41283f54c042 · inbound
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eca418-d976-4faa-adc6-75a5568b58d2 · inbound
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ddc02e1-1c64-435e-a3f3-29e50eaa25fb · inbound
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcff0bad-9c40-47ae-9458-c26b0d523365 · inbound
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a264c7a1-7d9c-4008-99ed-4e807bc9dd5a · inbound
Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04c2190-6cf4-4248-a223-772acb2fbd56 · inbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9d97cbca-ed4a-4607-b337-de37b2c26d8e · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e536b025-dadf-432c-896c-4eeb897aa608 · inbound
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e35a6860-e50e-4911-b3a9-e07247782165 · inbound
FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2adc0b12-fb7d-4065-999e-67083037bda7 · inbound
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5eaa3acb-9858-4fa4-8e60-b1bc9710040d · inbound
EarlyTom: Early Token Compression Completes Fast Video Understanding Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b92e2ef-7f1f-4dde-83e9-3d2bcd497210 · inbound
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c9a07ee-aa4e-4607-95e9-d6b945e612eb · inbound
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 78b96a77-ebe3-4938-adc2-b267a4e911e4 · inbound
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c222cc31-d96b-440e-80bd-170880e2e2bd · inbound
MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 13493d25-1882-4ac2-9c22-5251d778b29b · inbound
Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cac17e1-2c7a-41fb-9ea6-fe26463ff5d8 · inbound
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d05624f-a8bc-480b-9d30-1829bd76b99e · inbound
SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ac6f38-9978-46a6-aff6-7293d6157dac · inbound
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2aee47a-d6c7-43da-bbc9-29666f38fb20 · inbound
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8bd15c3-4ac0-4d71-9160-9ad85923ffd8 · inbound
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b888b749-faa7-4be3-b08e-3024b0e967d0 · inbound
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f395aee9-5bdd-4512-9f45-d0f050a073f6 · inbound
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.