Pith. sign in

Paper Citation Record · LEDGER

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

As of 18 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 10 inbound Pith citation observations for arXiv:2412.03248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03248 v2

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:42:35.606021Z

measured 110 of 110 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.890317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:41:03.999598Z

Reference resolution

100 of 102 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0029a0e-2d85-4f64-bf21-3b59605dedd3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.441164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.441164Z digest=sha256:2a0daad2141b6d54932974ad501bac82c31fb71b588e18123c95447bdd9f37c7

Observation 4a517098-19be-47ea-967d-9a3c809b86a4 · outbound

This paper cites Conditional Computation in Neural Networks for faster models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Conditional Computation in Neural Networks for faster models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.450087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.450087Z digest=sha256:6c53cbfd9cc0b2202583a648ffc4e44d91055cc04a0003e8582b1d1f56776eeb

Observation 66289086-3f93-4906-8a59-72f723f4ab3c · outbound

This paper cites Token merging: Your ViT but faster.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Token merging: Your ViT but faster

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.457746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.457746Z digest=sha256:5791dc05a4e820e3c618128b955e5470660d7feb85b7b24d98967f46f0c701ee

Observation fd99e16d-42e9-4d6b-8619-f97814eb7df5 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.466118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.466118Z digest=sha256:77685ed98e4a7b90cac52b14fff77ecd5c2b4b95e5f6fcdafcb2035a1325728a

Observation 2f07c04b-41a2-471e-a262-8a5c6b562efb · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference accelera- tion for large vision-language models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning An image is worth 1/2 tokens after layer 2: Plug-and-play inference accelera- tion for large vision-language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.474254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.474254Z digest=sha256:ff9c61dcc6dd9bcd18e12aa1a4299c775c2e5460d4c0a9997bf4c146a1812da2

Observation 97c3365e-a694-4982-956c-8e52125de695 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Gonzalez, Ion Stoica, and Eric P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.492269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.492269Z digest=sha256:cef614f97e0da46a43145183fefd4545f4a674a44ed3758d4efb3ef7cd5bd565

Observation d6522194-ca2d-4132-b55a-6df2d3c786b4 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Generating Long Sequences with Sparse Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.508061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.508061Z digest=sha256:35983cc9714e2bb0a16ad15df18b8caf371a329d27f199d843ced8414290d665

Observation dbc93ba2-a3bf-4978-8744-ed063811f16d · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.522544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.522544Z digest=sha256:8cd944c9966feedfc33cd3d461231af56298e8ca54ed1a477c812eda0c1383b8

Observation 72f6d86b-a88b-4be4-a8fb-fc4c955b7f61 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.531047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.531047Z digest=sha256:aa645fbe95c3717cc9d7d99e1f048c3f9711c50e72df293945b18686157c1c8d

Observation e36d8d08-4439-4c6c-8880-2dfe9a25e92c · outbound

This paper cites an unresolved cited work.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.537975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.537975Z digest=sha256:34b905eef0b81b90bfeea6a7eb07ee9c8f923d5d5723ddb005cfb2c864e50525

Observation 4bca0f33-9232-44b1-8c62-abfb515c7a90 · outbound

This paper cites HeatViT: Hardware-Efficient Adap- tive Token Pruning for Vision Transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning HeatViT: Hardware-Efficient Adap- tive Token Pruning for Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.547798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.547798Z digest=sha256:e1d74c81282cc809b4cc38a46b9620cfe54aa8b5e0ee35cbb03f3df502bae740

Observation b394b5fd-0e4c-410a-83ad-899aa9c32586 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Glam: Efficient scaling of language models with mixture-of-experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.559597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.559597Z digest=sha256:d78f05fced1a6b263770791f360c9c40679aa327ae64e73b545bcff4929b8e6f

Observation be2ceb51-bacf-459b-9560-690a844a44ab · outbound

This paper cites Hsu, and Shang-Hong Lai.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Hsu, and Shang-Hong Lai

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.568230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.568230Z digest=sha256:96c07fea3aac965f31e1b2f3637f23ee8bb67401fe290065bddaf479c398f7ee

Observation 1c4a36c6-2fe0-4eeb-b36b-2734fda1b327 · outbound

This paper cites Adaptive Token Sampling for Efficient Vision Transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Adaptive Token Sampling for Efficient Vision Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.628752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.628752Z digest=sha256:3de33b8412be377110ee49b34e54293977b1915bf0d46e0b2c67d892bbdb14ad

Observation 9b902260-180f-45a8-9d41-51803814a93a · outbound

This paper cites Spatially adaptive computation time for residual networks.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Spatially adaptive computation time for residual networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.720208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.720208Z digest=sha256:c793a19a24857e05eb098aa31983a1835fba086d81d653eb6e134485836f9381

Observation 37316ee7-8a7b-4b70-8305-4fb2a3363bd7 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.813015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.813015Z digest=sha256:b6dfddf207da09b39b50ce5efabccdf22ebeea36a769f8fed940d205e122d8d1

Observation 94e89941-f4ea-47ea-abe5-7ebd81815d09 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:32.928121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:32.928121Z digest=sha256:24b0db7a6d7237f67ac8dc0c9daec5761dab8fa15759de207217788a777fa6e8

Observation 3e51cbe8-edf1-435c-96f1-073d904d239d · outbound

This paper cites PoWER-BERT: Accelerating BERT Inference via Progressive Word-Vector Elimination.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning PoWER-BERT: Accelerating BERT Inference via Progressive Word-Vector Elimination

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.027600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.027600Z digest=sha256:5c9a263c655f3bd0396b46ae0ca6ac6e0ac6c3e62967bcc71331e8e59c15b38c

Observation 35e73f7a-154d-4448-bfa4-a893656e4d6b · outbound

This paper cites Making the v in vqa matter: El- evating the role of image understanding in visual question answering, 2017.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Making the v in vqa matter: El- evating the role of image understanding in visual question answering, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.040285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.040285Z digest=sha256:c92233155ce3a8271b484dec4d3f70ec213e708a72837521482fb24c167918d3

Observation f38655c9-030c-4e82-a64c-71a71c8d5e58 · outbound

This paper cites Speedboost: Anytime pre- diction with uniform near-optimality.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Speedboost: Anytime pre- diction with uniform near-optimality

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.048908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.048908Z digest=sha256:bcbea27a2c3b1c34c527e7e368d5a2a00d96d22019519101d591ab050adc8c30

Observation be8f9fa3-87c4-4a3b-96ef-f32f91eac3f2 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mamba: Linear-time sequence modeling with selective state spaces, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.054914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.054914Z digest=sha256:b6a26f33248c6798f277f986962ef41fb06e023ff4bcaeea5be67ea0e0d6eb33

Observation 14e87409-3aef-4da1-93ee-bcf6254b8109 · outbound

This paper cites Dynamic neural networks: A sur- vey.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Dynamic neural networks: A sur- vey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.060957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.060957Z digest=sha256:45baa5b8c5a9757b5c09fbff927e17220b11b7fa8e8cf213460bb9293af7d4fc

Observation df0600ee-3fd0-4408-9720-d7210b29bac3 · outbound

This paper cites Learning anytime predictions in neural net- works via adaptive loss balancing.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learning anytime predictions in neural net- works via adaptive loss balancing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.067491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.067491Z digest=sha256:fb45b92bab82fddd37a4759f54d3f922d3d96f4b42df96536a51dac70bc5c8ed

Observation 14db3d63-a0d8-47ed-932d-dde1e2027987 · outbound

This paper cites Longrecipe: Recipe for efficient long context generalization in large language models, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Longrecipe: Recipe for efficient long context generalization in large language models, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.079353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.079353Z digest=sha256:782249a2188cda879e065d56105f0a368d22596ccf460838bd1912571ddf8a24

Observation 31e093b3-e80b-421d-8892-b1679769d893 · outbound

This paper cites Language is not all you need: Aligning perception with language mod- els.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Language is not all you need: Aligning perception with language mod- els

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.088729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.088729Z digest=sha256:c284682cee07078a722fee350c76b99917dad35a7fc7ee07e3b1d7d142bd94ec

Observation 362ebd48-17f1-4b06-8dc4-1fd2086584d4 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.096469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.096469Z digest=sha256:17ff5cc9b29323463a017e7429be4a144088b703afdea7dc590d16860dc14aee

Observation 14874ec9-b7bb-47fb-8efa-37e38cf72d37 · outbound

This paper cites Anytime recognition with routing convolutional networks.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Anytime recognition with routing convolutional networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.101574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.101574Z digest=sha256:8a5018b1bb2359919698a413afb5cbd94464c9f41bf6c82dd447930da7c369df

Observation 2bf2311f-1c4e-4b38-a586-8aefa9dec776 · outbound

This paper cites Anytime recognition of objects and scenes.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Anytime recognition of objects and scenes

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.107004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.107004Z digest=sha256:4230c431ea27119f0d1494c67da77488ce05d435aed4fb8686b5ef8d0481c20e

Observation a7c55cd4-6d31-4ac9-8cb4-4f394139b434 · outbound

This paper cites Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.113003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.113003Z digest=sha256:b06a12b43ea5ab9729b743c5e70c961128aec8ac0767a0893cafbe54fe9cea01

Observation 8500cadb-5d0b-45a9-b01d-ff4130395d82 · outbound

This paper cites Learned Token Pruning for Transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learned Token Pruning for Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.119961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.119961Z digest=sha256:b9334960f82f72556560825fc1ed16bb2720ab1ebc8dd5aaa04c7d4043e3803d

Observation b730ae6a-7ada-49ba-8c56-6820884d089d · outbound

This paper cites SPViT: Enabling Faster Vision Transform- ers Latency-Aware Soft Token Pruning.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning SPViT: Enabling Faster Vision Transform- ers Latency-Aware Soft Token Pruning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:40.139405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.126479Z digest=sha256:c464b4e46496d15a87af5fd1b372d3969be45eecba534e35fe46c16c573b03d6

Observation d118ceed-a1ec-48ea-9f7c-790a8a5b2c11 · outbound

This paper cites Text-conditioned resampler for long form video understanding, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Text-conditioned resampler for long form video understanding, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:40.097815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.137605Z digest=sha256:514cedcdadbfe132b8d1a5cf6d9d3732175c4b4b510a0a2bec72311d486f888d

Observation 49b89833-6f25-43fd-88c9-fb5def88c31e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LLaVA-OneVision: Easy Visual Task Transfer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.145410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.145410Z digest=sha256:5cffb4d1627d48a43ecbdd9bc6704818924de7a7dee74fdb87e1747f14c0577b

Observation b5837c2c-d2bc-42ad-b801-f47ec19ea465 · outbound

This paper cites 2d or not 2d? adaptive 3d convolution selec- tion for efficient video recognition.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning 2d or not 2d? adaptive 3d convolution selec- tion for efficient video recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:40.044840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.158147Z digest=sha256:0e76dcaaeb449800f9bf19d8bd272af1a26d1feed1e58d729049e8d626afa7a8

Observation 240bd5a9-c36f-4a23-9933-a287b6d98e0f · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.163460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.163460Z digest=sha256:563a3e32f3ebc593e7a9a80edb3c673592267d02c82a4b2cd0cbdb770d85d629

Observation 450e9e6a-6069-4479-a81b-3a6d5b58cd4c · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2023.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mvbench: A comprehensive multi- modal video understanding benchmark, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.992675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.227811Z digest=sha256:c0e4d85dbf96e2a4306e163c8daacb37b855b3d6beec517d82ee69f64571869a

Observation 880428ec-814b-47b5-978f-618641a884b9 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mvbench: A comprehensive multi- modal video understanding benchmark, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.915018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.350463Z digest=sha256:59417f337d262ba1ead906345cf815af8090c90630f4185864b4bc9822ad0feb

Observation fa7bce63-e9cf-4fab-bb1a-08ed492d2570 · outbound

This paper cites Evaluating object hallucination in large vision-language models, 2023.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Evaluating object hallucination in large vision-language models, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.830333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.486802Z digest=sha256:4d5ee92ae7efa18c8b2234b8a6de125c6521cf0197c38065246e8be0dadf9f68

Observation 03390d55-db6f-45ac-b385-58661ceaa3a4 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.499475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.499475Z digest=sha256:b58d305ce45af78182d118e204e0e913e116121c195c82292d85f0b0ec4ca7de

Observation c29085ce-8130-491a-9e3e-94c02e51fc3a · outbound

This paper cites Boosting multimodal large language models with visual to- kens withdrawal for rapid inference.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Boosting multimodal large language models with visual to- kens withdrawal for rapid inference

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.710130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.510751Z digest=sha256:59a8e1904145e409a4bee368425f0656eeb32d6c1ea648ad290079e0e3ac081e

Observation 6b146270-361b-468b-95f7-99f9db001ecc · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Improved baselines with visual instruction tuning, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.611242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.524516Z digest=sha256:113374be31b0a5c155795008f9933f8ee06d92d67ae5f702f1f05d8b84947a92

Observation 040a0f2b-e501-42b1-85f5-12aa6a7fb64c · outbound

This paper cites Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Adaptive Sparse ViT: Towards Learnable Adaptive Token Pruning by Fully Exploiting Self-Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.536591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.536591Z digest=sha256:77d4ddba81d8d696dd3a177c3dca19739b18ef387f8a144279700d0b41a16f9d

Observation 2cd49640-e815-4fd6-9613-86b83b60705e · outbound

This paper cites Mm- bench: Is your multi-modal model an all-around player?,.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mm- bench: Is your multi-modal model an all-around player?,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.524869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.547736Z digest=sha256:7a65f3127bf7bbcec2eca2c644fd18c683175cf9dd745726f1ad9e44ea130d35

Observation 8258f6ad-9655-464a-9a08-a25085257104 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.488095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.556286Z digest=sha256:a0a047a8e80c23ae709d7e2f2bc4bd4991e35235e468391cc0afac0445dd0e72

Observation 7005e5a2-b4fd-4291-996b-c573d4e492d6 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.565798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.565798Z digest=sha256:b3cd77effdc582377251638bdc6b5efaf703b8dc7dcc1400c6cf2fe2144e2833

Observation 86a0b85c-9408-4b25-b55b-2cf851f878f8 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding, 2023.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Egoschema: A diagnostic benchmark for very long- form video language understanding, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.384758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.575589Z digest=sha256:507b1a6895caaee6b257c84b3fea71b509b210bb867a22e855155b62d6db9878

Observation d8e3a40c-93ca-4ac9-8563-51d4dc79cede · outbound

This paper cites Llama3.2.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llama3.2

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.235114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.585461Z digest=sha256:f01637884bec5af12b73c885e7e72ee4e6535dff5b879aa72bc4890e58047e43

Observation 457fb89c-8351-4a77-b0ad-bbda2d801f59 · outbound

This paper cites Adavit: Adaptive vision transformers for efficient image recogni- tion.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Adavit: Adaptive vision transformers for efficient image recogni- tion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:39.061038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.598813Z digest=sha256:90cad48c8b3f5980d83c209d430782cf49a1fb819891f40ca1e93a512383055f

Observation d18a5a18-d1d9-49d0-bd2d-9e07a68b1f16 · outbound

This paper cites Ar-net: Adaptive frame resolution for effi- cient action recognition.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Ar-net: Adaptive frame resolution for effi- cient action recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.958217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.616022Z digest=sha256:5a087082efc6d29ac6eaf686bac87d8013755a93bbe50ae15bdabdf3143c7be7

Observation 4ac4217a-9349-4990-8340-b75bfc0a5d3f · outbound

This paper cites Compositional chain of thought prompting for large multimodal models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Compositional chain of thought prompting for large multimodal models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.887809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.624337Z digest=sha256:04931aee8c63921d45796eb8354067ff88d1ac6ecb68ab48537d87bab8ff4dea

Observation 5c6b7c86-c7ae-4b2c-8ef6-895367126778 · outbound

This paper cites Encod- ing and controlling global semantics for long-form video question answering, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Encod- ing and controlling global semantics for long-form video question answering, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.804751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.630775Z digest=sha256:33cd79b5892b51057ddf1251d8a11db2173d84280a61d115b9e7af48ad33323c

Observation 9a3428c4-d063-4309-b640-4bb0824975d0 · outbound

This paper cites an unresolved cited work.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:42:38.754751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.637576Z digest=sha256:efebe85a409a72bcf359be610e3aa25233770543977d1446f99da9491e45e8dc

Observation d97414c5-bacb-4d04-b009-efc026c23638 · outbound

This paper cites Gpt-4 technical report, 2023.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Gpt-4 technical report, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:33.655655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:33.655655Z digest=sha256:54b37787a01591b2575c073881dc2e0d9d2ef91768b33cb4714a6995a70ebb6c

Observation ff77bf7b-c85c-4bfa-a91f-64abccbfbcf1 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Training lan- guage models to follow instructions with human feedback

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.590187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.668797Z digest=sha256:50e6bc9189f5e3605a6f2e173473689d68d2f5e0b95ea04624dadcb3712ff5d6

Observation dd422b71-84c5-4a32-aaa8-3cf4bbcc2929 · outbound

This paper cites The pagerank citation ranking: Bringing order to the web.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning The pagerank citation ranking: Bringing order to the web

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.507077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.762151Z digest=sha256:7d1e837b23372a92937d688884cebe1d3700cb185e07a5cde29ef0aba0835495

Observation 0e6327bc-7b09-4b33-8b10-550cec56181b · outbound

This paper cites Ia-red2: Interpretability-aware redundancy reduction for vision transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Ia-red2: Interpretability-aware redundancy reduction for vision transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.454239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.849543Z digest=sha256:65b002dd09e4086340bdb069bbc569c8d3438472746d43d3e3d806ab2bb9102b

Observation f83faf88-c2f2-4738-ad7d-08674c4912e0 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Perception test: A diagnostic benchmark for multimodal video models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.428599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:33.975137Z digest=sha256:307787512bea7b59775643e1633f05955a51ed3ee87c2382866b2f7aff127451

Observation 5accbd04-bbae-4721-9691-191243fa9763 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learning transferable visual models from natural language supervi- sion

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.380154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.063231Z digest=sha256:c5fc20cd97fab163e941e296e1e87f90773035efd348da52cb29de67d0d91770

Observation 1a0956f3-402a-4f0d-a575-423db1e20a99 · outbound

This paper cites DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.333199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.158807Z digest=sha256:56f35c1be6f2a743f834d8b4120c6712ec357a327776a20207ce5422da181db0

Observation 7c80718b-24b9-47fe-b54e-1dbce6f47b46 · outbound

This paper cites Finding the sweet spot: Analysis and improve- ment of adaptive inference in low resource settings.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Finding the sweet spot: Analysis and improve- ment of adaptive inference in low resource settings

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.309827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.168279Z digest=sha256:3e8630a0f0fe176769934ceab236a1fc6005266dcd900d9c3cc01453279ee173

Observation ec62986d-32f0-4ecf-8b81-fb2b003c378f · outbound

This paper cites Interpolating video-llms: Toward longer- sequence lmms in a training-free manner, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Interpolating video-llms: Toward longer- sequence lmms in a training-free manner, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.284187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.174000Z digest=sha256:788521e87ed4eab9018fa03bcd3eea531a2611e3f748ba57d80a30a16726816c

Observation 1edc9274-79e0-4e84-911f-923a55456b2f · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.201277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.180714Z digest=sha256:56688b19861081c8c2c1bd3c3d4f06ca6656bb2f0d0ef5f9b7eb9c87e3cbc801

Observation 6ede8a3c-8eca-4565-a815-9aaf4d00261e · outbound

This paper cites Towards vqa models that can read, 2019.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Towards vqa models that can read, 2019

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.115737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.194341Z digest=sha256:8dcf9809e7542574d7207a8fc0548e31d63f270dff1320f6d8c7318c7b82b89d

Observation 0b4c81a8-7082-49f3-b23d-8aed28fd2fb5 · outbound

This paper cites Dynamic Token Pruning in Plain Vision Trans- formers for Semantic Segmentation.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Dynamic Token Pruning in Plain Vision Trans- formers for Semantic Segmentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:38.073037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.214748Z digest=sha256:75cce27752c8765e0bb57ef85c58cf464a9d7e53b5e2df8756f4b7594210c60f

Observation 1d22d025-eb75-4db5-886f-97daac3cb8c9 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Dycoke: Dynamic compression of tokens for fast video large language models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.973209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.225942Z digest=sha256:07d35174fae3ba458fdf43d10108ab6992ea3cfb291c1276fc3098ab37fef15c

Observation 038999bd-7622-49e7-b053-8cc33e73319c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.245840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.245840Z digest=sha256:108d60a35e604ca701af795d17e217de1271fe6c6f3cf0b220a2102c01034103

Observation 0789ecf3-c8ce-494a-95d9-7643ad399a1d · outbound

This paper cites SpAtten: Ef- ficient Sparse Attention Architecture with Cascade Token and Head Pruning.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning SpAtten: Ef- ficient Sparse Attention Architecture with Cascade Token and Head Pruning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.945733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.254738Z digest=sha256:fdd189450a56769fe67ec622e345bc7a713eb1be152946aad8335771c0b55343

Observation 47e67a27-d0a6-4fcb-a93f-35500c172fb7 · outbound

This paper cites Zero- tprune: Zero-shot token pruning through leveraging of the attention graph in pre-trained transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Zero- tprune: Zero-shot token pruning through leveraging of the attention graph in pre-trained transformers

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.918130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.294829Z digest=sha256:8256be2bd49dc221867fc059464e51f474079a3171fcc89da9c67249e61c2b59

Observation 1227f3ed-7a15-4a42-9eee-8cf8e7d7434e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.320401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.320401Z digest=sha256:27c8b876ea7da7edbec1c175d82c4362b6a48a7220baa8ad92440ecf7e1880a7

Observation 6e4b5ed8-9e99-4eab-9beb-bbdf5ef19fd1 · outbound

This paper cites Skipnet: Learning dynamic routing in convolutional networks.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Skipnet: Learning dynamic routing in convolutional networks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.882126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.400690Z digest=sha256:17bbb20fb4a5d1a7c991b61a19853b25ccfdf6d9fcc5958bdb386239601fd65c

Observation 0e2bba78-430d-4540-965f-c24aea22fe1a · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Longllava: Scaling multi-modal llms to 1000 images efficiently via a hybrid architecture, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.836613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.534876Z digest=sha256:98b201ecafa48144870ea4773823f05ec491a7c918e50b8beb5ac2898cd918e8

Observation f5a96be5-5d41-47b8-9b50-2bcdbbba7819 · outbound

This paper cites Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.792671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.655401Z digest=sha256:b9b79efd6860dab6190bb067de192ecdbab2c89fe053fd601204fd71942944b9

Observation c638345b-7acb-4397-a2b1-5bfe2aea6aa0 · outbound

This paper cites Videollamb: Long-context video understanding with recur- rent memory bridges, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Videollamb: Long-context video understanding with recur- rent memory bridges, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.768762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.718307Z digest=sha256:b4e6a0d23c883182101f412d08ea295046f5d3e5cba5318ed7a0f7c3c47843a8

Observation a4c1ce93-90a4-4ac0-802e-7c8f7fd582cd · outbound

This paper cites Visual context win- dow extension: A new perspective for long video under- standing, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Visual context win- dow extension: A new perspective for long video under- standing, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.744662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.731408Z digest=sha256:bf96e983fe5fbdcad72b072d9be803f8b422d5243dcf66d13aff0e04c9fcafe5

Observation 23241af9-9f4e-43c9-9542-b371537fe413 · outbound

This paper cites Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.679928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.740948Z digest=sha256:bd2584a55090232b19aff1fe1fa2691a7c6f2e0c4c955a950ff87a795d4394e7

Observation cfcc0320-87b6-4b8f-8565-bd8a1bdad02c · outbound

This paper cites Longvlm: Efficient long video under- standing via large language models, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Longvlm: Efficient long video under- standing via large language models, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.656283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.751404Z digest=sha256:2f8f9f18ec12b5ebc3fa6d26b3a29cc51b7b77862d9f2c66d04d52b51ddf69e4

Observation 47846cea-a78b-4d35-8e4d-85e9b8e37c97 · outbound

This paper cites Blockdrop: Dynamic inference paths in residual net- works.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Blockdrop: Dynamic inference paths in residual net- works

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.624531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.759486Z digest=sha256:c3afe35fd9be1ca53ff00fdcf039c976dafc7b0710f6482af5112358757661c1

Observation 4648a01f-e97b-44bc-b8d3-aa462314bc7a · outbound

This paper cites Next-qa:next phase of question-answering to explaining temporal actions, 2021.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Next-qa:next phase of question-answering to explaining temporal actions, 2021

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.570100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.770470Z digest=sha256:0d2dd9ea62f92dfcb5295a9f5dde9084bd031f8687168b3c173e92234a5807c9

Observation 69b0ac4a-7df3-4171-a32e-707d0c6a62bf · outbound

This paper cites Conical visual concen- tration for efficient large vision-language models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Conical visual concen- tration for efficient large vision-language models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.531799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.783913Z digest=sha256:7633e17b7d2349c0b37a8634acf1f064e0bf46739f9a56146ae6abc1af576786

Observation e2cfcd85-85d0-4b5c-a3e2-d397f22d1590 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.808922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.808922Z digest=sha256:b0d38c6a6d7443c90ea98c41c979dcb98bbaa1e598a6ea422a315692b4caec5d

Observation c679c919-d041-4e74-ad92-17cc044aa0bf · outbound

This paper cites Smartadapt: Multi-branch object detection framework for videos on mo- biles.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Smartadapt: Multi-branch object detection framework for videos on mo- biles

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.454749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.822061Z digest=sha256:67243e5e68cb60efe744e0c83a057ee7c47d94d46c4b99d4af69dd3f260ba794

Observation 2e149b68-1f3c-4aba-b4ea-299c66d8b911 · outbound

This paper cites The greedy miser: learning under test-time budgets.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning The greedy miser: learning under test-time budgets

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.386268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.832500Z digest=sha256:b5638e1f01f50c42b4562b930ac24db4bddfc5220994270a9d4b38b1689d1423

Observation 2b42e201-f2c7-4aa2-a3ea-52167ea5b291 · outbound

This paper cites Learning to inference adaptively for multimodal large language models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Learning to inference adaptively for multimodal large language models

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.337698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.842779Z digest=sha256:d604ac9992532b85ba4d123fae2fc481f538be90adda7dfd2eb4181c27ab321c

Observation 071fcb59-3f8e-4437-acd7-d53341c590dc · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.850695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.850695Z digest=sha256:62d04b7c782c7d85f218ac7aea4132f61edab382698ac9b023cddf1c1c4f8528

Observation d2a1cada-6559-4c95-85e3-a887118f5636 · outbound

This paper cites Qwen2 Technical Report.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Qwen2 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.856995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.856995Z digest=sha256:ff00bb8c7564b3d8b9a0fb36e2c7fd3d83d0eb85217fa4f3a595bec7fb784899

Observation 33c82efb-326a-4499-83ab-a3caa6c57001 · outbound

This paper cites TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:34.872462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:34.872462Z digest=sha256:62cb793d9447cf6f8c1fb1285d4f32c95d6b6a8568dc015279adfda485f38040

Observation 2a7689da-4d9d-4c25-8847-cc9589d1b9b4 · outbound

This paper cites mplug-owl3: Towards long image-sequence understanding in multi-modal large language models, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning mplug-owl3: Towards long image-sequence understanding in multi-modal large language models, 2024

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.274750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:34.954224Z digest=sha256:7026eaa91c4d832e437787e76c52cc71ec892a59a183e54784901575bd1e6940

Observation d628cffe-e20e-45bf-ad9f-5f37ede00a6d · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi-modal large language models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Fit and prune: Fast and training-free visual token pruning for multi-modal large language models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.208625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.098347Z digest=sha256:ae9bf697cf8355791bbf30236274daf638fcc290b2404ece5cfe0d94a4a6af7d

Observation 2b0bff02-7148-49f2-b7bc-6a5c1ea8452b · outbound

This paper cites Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.163076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.204596Z digest=sha256:5eaeff05250addebe60589aba789766089f62df44b23ee84cef4144e7315a183

Observation 0f37e1d6-9c72-451e-9a23-fd816d9d0a07 · outbound

This paper cites Llm inference unveiled: Survey and roofline model in- sights, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llm inference unveiled: Survey and roofline model in- sights, 2024

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.130770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.309467Z digest=sha256:000a70470336a619126b7d1a6e3bc4d413e34dddc7fefc54f7618580558219e3

Observation 9c8afaca-8878-42c6-9f66-921e155a77d1 · outbound

This paper cites Sigmoid loss for language image pre-training.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Sigmoid loss for language image pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.411927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.411927Z digest=sha256:472c1a6ff741fdd8a42128ba8e2e8585cfcecbe5d1b3533268ee1988ea9cb2ed

Observation 30ce5201-7e5d-4032-bfee-955a1cec7fbe · outbound

This paper cites Long Context Transfer from Language to Vision.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Long Context Transfer from Language to Vision

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.454216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.454216Z digest=sha256:27e0c9104f217c2a7de96bbf7896fa16cd8e57b873fc3b8a5a6154e24928c5cf

Observation 3cc18f74-ef5e-440c-8bf0-04d87bcdaf89 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning OPT: Open Pre-trained Transformer Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.470309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.470309Z digest=sha256:670cf35a3889d9ae91e4ac73a1792cdd276682623b99d3edb2614c9599b1d504

Observation bf704827-ccb7-4c18-976d-ddc899f90bf2 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model,.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Llava-next: A strong zero-shot video understanding model,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.480309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.480309Z digest=sha256:f4f4e1c7139aaf28c0793b988bcc589664cc099c0c568577cca8afb206b9fa63

Observation e8c2bfff-7ab0-49dd-a3b8-863289b432fc · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Video instruction tuning with synthetic data, 2024

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:37.049515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.493317Z digest=sha256:2b8bbf3b82af4732416c17d0f82f3bc47be32efb94abdba6b67b9464bc62fa71

Observation e38301fb-a2b6-444d-83cd-7ae159be5830 · outbound

This paper cites Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:36.961412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.501667Z digest=sha256:946b9dc26bd9982fdbdc764a1ba9416b2bb54b204a20197c6b17978997ef648c

Observation f041d933-1902-412a-9ac3-1acf2b510799 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.510876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.510876Z digest=sha256:8fd996e0ce6af88652da6dad69abf906ec9eda9964d74d7e0fd47ec312638527

Observation 0e68b2d4-546a-4dd0-8c91-07175928aee8 · outbound

This paper cites Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:35.570387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:35.570387Z digest=sha256:995517c0d9218e9b6028b1dbbd28ed6bfb4ce36fbc920ce44f43bc6cab70f85b

Observation 8ae94114-09fb-465c-b2fb-d5ca5bed3427 · outbound

This paper cites Beyond embeddings: The promise of visual table in vi- sual reasoning.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Beyond embeddings: The promise of visual table in vi- sual reasoning

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:36.926701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.595613Z digest=sha256:a1141309470ac9fca254567dfc21bbd79aa5fe1edf377c4474f66f1df0bab529

Observation a9a29163-bbce-4992-842b-e601b02006ff · outbound

This paper cites Mlvu: A comprehensive benchmark for multi- task long video understanding, 2024.

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning Mlvu: A comprehensive benchmark for multi- task long video understanding, 2024

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:42:36.681732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T22:42:35.606021Z digest=sha256:e6054d151ad5a7b86bd2afb4ec30ccb74ae03bfcad136bd54c682c4cd6fd2093

Pith citing papers

Observation 7cf09e0b-8c67-4d98-8be8-2fb881ea285f · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:47.668642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:47.668642Z digest=sha256:08fd785b87a279f7ea158c08919eb09cc4c6dbaee2d62a923a146108892c4bf8

Observation ebd6af17-6836-4c4e-bfc1-b15169f0b67c · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.185790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.185790Z digest=sha256:e60c815eac36234370b32c7b4a4a9e1dba4e4aa4caf29fa8482e76fa1bf8a9a8

Observation 77e2f337-f90d-448d-b986-cf87959df91b · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.890317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.890317Z digest=sha256:92ce3c2dae90a187750b5e6ae4ce2fa7705af3b76a43c2389f6c0ae5753a8378

Observation f61d0b3a-33cc-4d09-9490-858863c934ec · inbound

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM cites this paper.

Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:12.434364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:12.434364Z digest=sha256:1f5b4c95e12e5f483da57a3cafc91675a4894fb4ef9bcdc614288d4d9f010089

Observation 4ed619fe-2182-414b-afff-bd1a1ca8dd69 · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:10.915558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:10.915558Z digest=sha256:9a7d990ba4d9a2f71591c1a76e26eb6ddae36392ff9d06ccc16b807220eb8050

Observation b344fd13-b7d1-46a2-af72-80b526c10748 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.555597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.555597Z digest=sha256:74e6767423834847f78917c8d724b1988675e6b8019f79f6e13fcb6a7f7fe157

Observation 7d65a803-0200-4e6b-aa91-044e0e65175b · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.370053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.370053Z digest=sha256:1f4b6dd30f78f658aed0ae2eb5307fac5a1206da9c76695a774f5ae62c66591c

Observation 6074cee1-1481-471e-9de5-4e6c03422122 · inbound

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects cites this paper.

Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:50.915833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:54:04.104227Z digest=sha256:85af8b60c5afdc914ecf7c7a3c6b1c12febb9b7378faef975d1dc4c0fb1ae0d8

Observation 1228267b-abf2-44b3-99d4-d9c23b22df29 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.005982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:777461951ba2e0a38a34700ce74e967ee1bda775b9d4457bd3feba8c528d1778

Observation 6db44e5d-2150-469d-81d6-7bd821c2662a · inbound

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs cites this paper.

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:42:44.443670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:42:44.443670Z digest=sha256:2e4417fcfd77f225f9f537c363e97f54bb33a00f88c1b179c1cbfd72b7f21679