Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:42:35.606021Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 10 inbound Pith citation observations for arXiv:2412.03248.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:42:35.606021Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.890317Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:41:03.999598Z
100 of 102 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0029a0e-2d85-4f64-bf21-3b59605dedd3 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a517098-19be-47ea-967d-9a3c809b86a4 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Conditional Computation in Neural Networks for faster models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66289086-3f93-4906-8a59-72f723f4ab3c · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Token merging: Your ViT but faster
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd99e16d-42e9-4d6b-8619-f97814eb7df5 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f07c04b-41a2-471e-a262-8a5c6b562efb · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning An image is worth 1/2 tokens after layer 2: Plug-and-play inference accelera- tion for large vision-language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c3365e-a694-4982-956c-8e52125de695 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Gonzalez, Ion Stoica, and Eric P
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6522194-ca2d-4132-b55a-6df2d3c786b4 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Generating Long Sequences with Sparse Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc93ba2-a3bf-4978-8744-ed063811f16d · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72f6d86b-a88b-4be4-a8fb-fc4c955b7f61 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36d8d08-4439-4c6c-8880-2dfe9a25e92c · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bca0f33-9232-44b1-8c62-abfb515c7a90 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning HeatViT: Hardware-Efficient Adap- tive Token Pruning for Vision Transformers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b394b5fd-0e4c-410a-83ad-899aa9c32586 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Glam: Efficient scaling of language models with mixture-of-experts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2ceb51-bacf-459b-9560-690a844a44ab · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Hsu, and Shang-Hong Lai
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4a36c6-2fe0-4eeb-b36b-2734fda1b327 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Adaptive Token Sampling for Efficient Vision Transformers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b902260-180f-45a8-9d41-51803814a93a · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Spatially adaptive computation time for residual networks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37316ee7-8a7b-4b70-8305-4fb2a3363bd7 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e89941-f4ea-47ea-abe5-7ebd81815d09 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e51cbe8-edf1-435c-96f1-073d904d239d · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning PoWER-BERT: Accelerating BERT Inference via Progressive Word-Vector Elimination
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e73f7a-154d-4448-bfa4-a893656e4d6b · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Making the v in vqa matter: El- evating the role of image understanding in visual question answering, 2017
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38655c9-030c-4e82-a64c-71a71c8d5e58 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Speedboost: Anytime pre- diction with uniform near-optimality
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8f9fa3-87c4-4a3b-96ef-f32f91eac3f2 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mamba: Linear-time sequence modeling with selective state spaces, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e87409-3aef-4da1-93ee-bcf6254b8109 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Dynamic neural networks: A sur- vey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0600ee-3fd0-4408-9720-d7210b29bac3 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learning anytime predictions in neural net- works via adaptive loss balancing
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14db3d63-a0d8-47ed-932d-dde1e2027987 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Longrecipe: Recipe for efficient long context generalization in large language models, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e093b3-e80b-421d-8892-b1679769d893 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Language is not all you need: Aligning perception with language mod- els
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362ebd48-17f1-4b06-8dc4-1fd2086584d4 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14874ec9-b7bb-47fb-8efa-37e38cf72d37 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Anytime recognition with routing convolutional networks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf2311f-1c4e-4b38-a586-8aefa9dec776 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Anytime recognition of objects and scenes
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c55cd4-6d31-4ac9-8cb4-4f394139b434 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8500cadb-5d0b-45a9-b01d-ff4130395d82 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learned Token Pruning for Transformers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b730ae6a-7ada-49ba-8c56-6820884d089d · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning SPViT: Enabling Faster Vision Transform- ers Latency-Aware Soft Token Pruning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d118ceed-a1ec-48ea-9f7c-790a8a5b2c11 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Text-conditioned resampler for long form video understanding, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 49b89833-6f25-43fd-88c9-fb5def88c31e · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LLaVA-OneVision: Easy Visual Task Transfer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5837c2c-d2bc-42ad-b801-f47ec19ea465 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning 2d or not 2d? adaptive 3d convolution selec- tion for efficient video recognition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 240bd5a9-c36f-4a23-9933-a287b6d98e0f · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450e9e6a-6069-4479-a81b-3a6d5b58cd4c · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mvbench: A comprehensive multi- modal video understanding benchmark, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 880428ec-814b-47b5-978f-618641a884b9 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mvbench: A comprehensive multi- modal video understanding benchmark, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa7bce63-e9cf-4fab-bb1a-08ed492d2570 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Evaluating object hallucination in large vision-language models, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 03390d55-db6f-45ac-b385-58661ceaa3a4 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29085ce-8130-491a-9e3e-94c02e51fc3a · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Boosting multimodal large language models with visual to- kens withdrawal for rapid inference
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b146270-361b-468b-95f7-99f9db001ecc · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Improved baselines with visual instruction tuning, 2023
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 040a0f2b-e501-42b1-85f5-12aa6a7fb64c · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cd49640-e815-4fd6-9613-86b83b60705e · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mm- bench: Is your multi-modal model an all-around player?,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8258f6ad-9655-464a-9a08-a25085257104 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7005e5a2-b4fd-4291-996b-c573d4e492d6 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a0b85c-9408-4b25-b55b-2cf851f878f8 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Egoschema: A diagnostic benchmark for very long- form video language understanding, 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d8e3a40c-93ca-4ac9-8563-51d4dc79cede · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llama3.2
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 457fb89c-8351-4a77-b0ad-bbda2d801f59 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Adavit: Adaptive vision transformers for efficient image recogni- tion
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d18a5a18-d1d9-49d0-bd2d-9e07a68b1f16 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Ar-net: Adaptive frame resolution for effi- cient action recognition
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4ac4217a-9349-4990-8340-b75bfc0a5d3f · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Compositional chain of thought prompting for large multimodal models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5c6b7c86-c7ae-4b2c-8ef6-895367126778 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Encod- ing and controlling global semantics for long-form video question answering, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9a3428c4-d063-4309-b640-4bb0824975d0 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d97414c5-bacb-4d04-b009-efc026c23638 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Gpt-4 technical report, 2023
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff77bf7b-c85c-4bfa-a91f-64abccbfbcf1 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Training lan- guage models to follow instructions with human feedback
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd422b71-84c5-4a32-aaa8-3cf4bbcc2929 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning The pagerank citation ranking: Bringing order to the web
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e6327bc-7b09-4b33-8b10-550cec56181b · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Ia-red2: Interpretability-aware redundancy reduction for vision transformers
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f83faf88-c2f2-4738-ad7d-08674c4912e0 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Perception test: A diagnostic benchmark for multimodal video models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5accbd04-bbae-4721-9691-191243fa9763 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learning transferable visual models from natural language supervi- sion
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1a0956f3-402a-4f0d-a575-423db1e20a99 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c80718b-24b9-47fe-b54e-1dbce6f47b46 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Finding the sweet spot: Analysis and improve- ment of adaptive inference in low resource settings
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec62986d-32f0-4ecf-8b81-fb2b003c378f · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Interpolating video-llms: Toward longer- sequence lmms in a training-free manner, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1edc9274-79e0-4e84-911f-923a55456b2f · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6ede8a3c-8eca-4565-a815-9aaf4d00261e · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Towards vqa models that can read, 2019
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0b4c81a8-7082-49f3-b23d-8aed28fd2fb5 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Dynamic Token Pruning in Plain Vision Trans- formers for Semantic Segmentation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1d22d025-eb75-4db5-886f-97daac3cb8c9 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Dycoke: Dynamic compression of tokens for fast video large language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 038999bd-7622-49e7-b053-8cc33e73319c · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LLaMA: Open and Efficient Foundation Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0789ecf3-c8ce-494a-95d9-7643ad399a1d · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning SpAtten: Ef- ficient Sparse Attention Architecture with Cascade Token and Head Pruning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 47e67a27-d0a6-4fcb-a93f-35500c172fb7 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Zero- tprune: Zero-shot token pruning through leveraging of the attention graph in pre-trained transformers
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1227f3ed-7a15-4a42-9eee-8cf8e7d7434e · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4b5ed8-9e99-4eab-9beb-bbdf5ef19fd1 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Skipnet: Learning dynamic routing in convolutional networks
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e2bba78-430d-4540-965f-c24aea22fe1a · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture, 2024
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f5a96be5-5d41-47b8-9b50-2bcdbbba7819 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c638345b-7acb-4397-a2b1-5bfe2aea6aa0 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Videollamb: Long-context video understanding with recur- rent memory bridges, 2024
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a4c1ce93-90a4-4ac0-802e-7c8f7fd582cd · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Visual context win- dow extension: A new perspective for long video under- standing, 2024
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 23241af9-9f4e-43c9-9542-b371537fe413 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cfcc0320-87b6-4b8f-8565-bd8a1bdad02c · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Longvlm: Efficient long video under- standing via large language models, 2024
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 47846cea-a78b-4d35-8e4d-85e9b8e37c97 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Blockdrop: Dynamic inference paths in residual net- works
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4648a01f-e97b-44bc-b8d3-aa462314bc7a · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Next-qa:next phase of question-answering to explaining temporal actions, 2021
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 69b0ac4a-7df3-4171-a32e-707d0c6a62bf · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Conical visual concen- tration for efficient large vision-language models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e2cfcd85-85d0-4b5c-a3e2-d397f22d1590 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c679c919-d041-4e74-ad92-17cc044aa0bf · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Smartadapt: Multi-branch object detection framework for videos on mo- biles
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2e149b68-1f3c-4aba-b4ea-299c66d8b911 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning The greedy miser: learning under test-time budgets
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b42e201-f2c7-4aa2-a3ea-52167ea5b291 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learning to inference adaptively for multimodal large language models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 071fcb59-3f8e-4437-acd7-d53341c590dc · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a1cada-6559-4c95-85e3-a887118f5636 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Qwen2 Technical Report
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c82efb-326a-4499-83ab-a3caa6c57001 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7689da-4d9d-4c25-8847-cc9589d1b9b4 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning mplug-owl3: Towards long image-sequence understanding in multi-modal large language models, 2024
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d628cffe-e20e-45bf-ad9f-5f37ede00a6d · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Fit and prune: Fast and training-free visual token pruning for multi-modal large language models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b0bff02-7148-49f2-b7bc-6a5c1ea8452b · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f37e1d6-9c72-451e-9a23-fd816d9d0a07 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llm inference unveiled: Survey and roofline model in- sights, 2024
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9c8afaca-8878-42c6-9f66-921e155a77d1 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Sigmoid loss for language image pre-training
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ce5201-7e5d-4032-bfee-955a1cec7fbe · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Long Context Transfer from Language to Vision
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cc18f74-ef5e-440c-8bf0-04d87bcdaf89 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning OPT: Open Pre-trained Transformer Language Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf704827-ccb7-4c18-976d-ddc899f90bf2 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llava-next: A strong zero-shot video understanding model,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c2bfff-7ab0-49dd-a3b8-863289b432fc · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video instruction tuning with synthetic data, 2024
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e38301fb-a2b6-444d-83cd-7ae159be5830 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f041d933-1902-412a-9ac3-1acf2b510799 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Multimodal Chain-of-Thought Reasoning in Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e68b2d4-546a-4dd0-8c91-07175928aee8 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae94114-09fb-465c-b2fb-d5ca5bed3427 · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Beyond embeddings: The promise of visual table in vi- sual reasoning
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a9a29163-bbce-4992-842b-e601b02006ff · outbound
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mlvu: A comprehensive benchmark for multi- task long video understanding, 2024
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7cf09e0b-8c67-4d98-8be8-2fb881ea285f · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd6af17-6836-4c4e-bfc1-b15169f0b67c · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e2f337-f90d-448d-b986-cf87959df91b · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 174
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f61d0b3a-33cc-4d09-9490-858863c934ec · inbound
Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ed619fe-2182-414b-afff-bd1a1ca8dd69 · inbound
AdaTP: Attention-Debiased Token Pruning for Video Large Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b344fd13-b7d1-46a2-af72-80b526c10748 · inbound
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d65a803-0200-4e6b-aa91-044e0e65175b · inbound
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6074cee1-1481-471e-9de5-4e6c03422122 · inbound
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1228267b-abf2-44b3-99d4-d9c23b22df29 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6db44e5d-2150-469d-81d6-7bd821c2662a · inbound
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.