Pith. sign in

Paper Citation Record · LEDGER

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

As of 14 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 10 inbound Pith citation observations for arXiv:2412.05819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05819 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:23:18.738520Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:57:32.530971Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T19:48:11.386460Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af9cca93-691e-416b-aece-26acfe1cacb9 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.356250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.356250Z digest=sha256:1f9c4cfe5ddfc205a896e35a56406161af330a53fd2773f6a757513ab1c98781

Observation 9cc3c0b1-7af7-40d9-9745-df66af626c35 · outbound

This paper cites Multimodal large language models in health care: Applications, challenges, and future outlook.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Multimodal large language models in health care: Applications, challenges, and future outlook

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.704931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.360947Z digest=sha256:f2824a108bcf43bf9b2ab62f0fbb6206064d7acef2f11d4a4c8cd6a8cea594d6

Observation 8b1831cd-5dcf-436b-9ffe-2064b2d15e2b · outbound

This paper cites Qwen Technical Report.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.365056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.365056Z digest=sha256:33bdcd825ed1bff40a5804c4a0ff17654e355ac19dbd10207588e006807370f5

Observation 977e4197-4e10-4757-80a4-3048724a4db1 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.369183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.369183Z digest=sha256:ea9c9d0f4d5927b6bb1434a054d1ef067877d7d0422f951ccf951c398638d393

Observation 755f1968-0778-46cb-a56c-2023e2213c83 · outbound

This paper cites Introducing our multimodal models, 2023.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Introducing our multimodal models, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.694317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.373356Z digest=sha256:8708e1eb149d856469bc7c60d40e9e8420859b62f8eff28f829aa0db2f39cdd4

Observation aaca4f51-1d4d-4002-8de8-9af90ae780e4 · outbound

This paper cites Token Merging: Your ViT But Faster.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Token Merging: Your ViT But Faster

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.376596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.376596Z digest=sha256:a751c121c4fbf56431db41ade442927d80868f587aa8169d67ef28d43864134f

Observation d2279f39-cdca-4968-bd6a-6f88994610fa · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.380176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.380176Z digest=sha256:de24e9bef2331c75f1cf337d533633c9ec26a9e260d2abd925f9e2fc6f4fa26a

Observation fd24611f-5d71-47f2-b737-b7ba2b7a7ac1 · outbound

This paper cites Show, observe and tell: Attribute-driven attention model for image captioning.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Show, observe and tell: Attribute-driven attention model for image captioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.682555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.384011Z digest=sha256:401164034b5025d55dad09de9cbbf674864b3ecb8b3a342ef5a69a87f7bdc2f3

Observation d50f9d2c-8fc1-4c9b-bad4-d1a90f425e0c · outbound

This paper cites Imram: Iterative matching with recur- rent attention memory for cross-modal image-text retrieval.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Imram: Iterative matching with recur- rent attention memory for cross-modal image-text retrieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.670506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.387158Z digest=sha256:5024c3b8880c46357059d239c77deada3878b784ad00317de5745377ddc823d9

Observation a1f51f80-4cd2-4477-b5aa-eaf8a52d8ad6 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.491097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.491097Z digest=sha256:f9c7b1af38f5f8abad1ed5428fb4d4ebc2887847faa225e07b65a113f4e85cd4

Observation db7cf982-5136-4bdd-8a21-7b8c849cdcf1 · outbound

This paper cites MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.494894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.494894Z digest=sha256:5cb162d1250e9b20265be937f7dc87088c045132e0ef518e9645bfd7fd159330

Observation d066c14d-e3fa-479b-9329-c87cb9a403b4 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.498892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.498892Z digest=sha256:28f068ac75b693f908cdb3de870c73100f7c7d7d623dd4564f7752aca0328d43

Observation 181bc5c6-cede-4180-a16c-967a7f8e7cf5 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.502572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.502572Z digest=sha256:2165189d7094a8291d1784d89ee0930c5488bdeffc31f04243826652cc2a23df

Observation f97ebf50-acbf-4769-baee-01799f66bb91 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.506754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.506754Z digest=sha256:1b1b6230c5e533fdeb1bd47125b832fe1a7540dba8857a6e0a5107f20b226306

Observation 3c115939-9158-4eea-8c30-386cea8a6f76 · outbound

This paper cites Pearson correlation coefficient.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Pearson correlation coefficient

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.649869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.510908Z digest=sha256:7d3fa6701179c3296bab86d6ee1fdd04b7d9370fa1206f10a34769dda5b56096

Observation 628f96cd-d9e8-4ba3-8396-e15707789759 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs A survey on multimodal large lan- guage models for autonomous driving

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.638745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.514729Z digest=sha256:145a690f6f2b5170170c903ce5250c1670c232a27b4ad6eb0d4008c24c56de33

Observation ee1115bc-874c-401e-891a-d828a6b126d5 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.518220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.518220Z digest=sha256:a4589f339a20c405636fa5c3c26aa27dc7b2a567df65a41a59b43328a5aada97

Observation 082a0047-1d8e-4ef8-babc-301cd1ce799f · outbound

This paper cites HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.522486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.522486Z digest=sha256:a3524c46e4c1aa6bfb29bd3336621b53e7f47903aef254f99bc4e7de8fabd676

Observation 0234ad18-c79d-43ed-8c49-05159676214b · outbound

This paper cites Explor- ing structured semantic prior for multi label recognition with incomplete labels.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Explor- ing structured semantic prior for multi label recognition with incomplete labels

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.626557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.526566Z digest=sha256:ca3c83b2ec8347b3c60baafbe7b87643a4d7ea0011bf864e3a63d39d017bb656

Observation 4685981f-1e81-4143-861e-703334cb700e · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.530205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.530205Z digest=sha256:c69199dec6c9a19b87b9d7b03a2e4fbd2d41806809c4800a78cc42fe2c76a46e

Observation dd472030-80f4-47e9-9dda-2415f5a30f8b · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs PaLM-E: An Embodied Multimodal Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.533914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.533914Z digest=sha256:b0f74f9a611002dcdabdfd4f2bc735a1fc5f88a64971821aa77750ce71e81b59

Observation ac042213-1efa-4ed4-b1d5-dcae4239769a · outbound

This paper cites The Llama 3 Herd of Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.537742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.537742Z digest=sha256:109919117012dffb485891df4a6d4b45ca09624baa895de601fc5086059c98ba

Observation c3aa61ae-4231-4d93-a69f-8d6a900c5b4d · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.541725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.541725Z digest=sha256:e175e777ce7ffc0a1220eadb48b8a7104552907863e2434e16dd1126aae41b14

Observation 8a70487e-312a-47c5-bb4c-e18ee6319381 · outbound

This paper cites Anomalygpt: Detecting in- dustrial anomalies using large vision-language models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Anomalygpt: Detecting in- dustrial anomalies using large vision-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.613138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.546828Z digest=sha256:7a915beeff2243c88115a67e6cb80f5e3db3753e72a2b5143e65411462afe653

Observation 5ca43d82-4d3f-4ee4-a754-c7936e043c96 · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.550329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.550329Z digest=sha256:5ed2d3f7aa264462ff78d118329ae2f5a8ac1980cfae01fc49b12c1515a0b687

Observation 589894f9-6786-4385-a1c4-3d6dfe505f86 · outbound

This paper cites MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.554361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.554361Z digest=sha256:0851c423c87f34330f6bd3c458453beb2c787d849f96091e5559d96b8bc9bf26

Observation d1c4ac31-27be-42db-a67b-6f948b24ce9a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.600304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.557628Z digest=sha256:b15a189c3f7d2bab60971f9a0485011551c640b90045c99697faa1e82b31f1b0

Observation 1061a524-6356-4223-ac99-80a8a9316454 · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.587856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.560615Z digest=sha256:4a863764f781dd04bcc09b20b0d6f5f9ac34533fe84c5aef82dc82946baaa6ef

Observation 46b589e5-bec0-4a3d-bf5f-75d3dd71c1b6 · outbound

This paper cites Phi-2: The surprising power of small language models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Phi-2: The surprising power of small language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.576132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.563973Z digest=sha256:7f1801a3b5bf86d49862ccd4ec3160186d595b781a5af1d350b0daf2f699707d

Observation e259e76a-c326-4baf-8d8c-cce36f5f320c · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Seed-bench: Bench- marking multimodal large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.564150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.567268Z digest=sha256:7add5813b5f9b96c76c66741b0b8eddaecb63aab374747b99a51d9906fa06e09

Observation 970a86f2-3855-4ca7-af90-422022a341a3 · outbound

This paper cites Llava-med: Training a large language- and-vision assistant for biomedicine in one day.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Llava-med: Training a large language- and-vision assistant for biomedicine in one day

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.551757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.570381Z digest=sha256:fd2bf00bbc9216905ffe809a7224785697ba9835565e0d499926065efc66040b

Observation 31aff3c2-f15a-489a-b608-26a52167a965 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.573814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.573814Z digest=sha256:73d0fa02b39a81be56a39ab4cdac8ada7e4a541c4fc4c314d27b5b6100063ade

Observation ffa2c15f-b5a3-4d35-9ecb-a1318814a07b · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.577569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.577569Z digest=sha256:059ad3833796431bf256285ff51ec5103d0a754df0055d1874b7ba9ffdd5590e

Observation 28e5d65b-674f-469c-b890-4ebdf1513676 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.581794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.581794Z digest=sha256:d45752a0cf6a3e869ff1f96c30cc9ba3a8adfe88e72087440e55becd7daf6d1e

Observation 506f0483-8b85-402b-8cbf-aa427a9dc19e · outbound

This paper cites Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Not All Patches are What You Need: Expediting Vision Transformers via Token Reorganizations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.586120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.586120Z digest=sha256:931f98826c86ebbd521b165d0b00eff4f89b392f994bc1f4875cc430971d4721

Observation 5916b2a0-6577-4a96-bbd8-10eca8e57e71 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.590375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.590375Z digest=sha256:e424535f12e47b8b9c4b28133f4c460ab7660d6b3abcadf0c1febae7e75ada75

Observation 866ee1f5-fce6-4aad-8900-51efb24b6527 · outbound

This paper cites Microsoft coco: Common objects in context.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Microsoft coco: Common objects in context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.594414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.594414Z digest=sha256:38651173d7132ecb260bbc336cff1782126fbaef87bc4574cfaac2bc428a1401

Observation 106e35d1-3e28-49cb-a613-9b7eeb4fb0bc · outbound

This paper cites Improved baselines with visual instruction tuning.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Improved baselines with visual instruction tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.532914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.598631Z digest=sha256:e9f0ad89acb4cccd4077cf159f428708c1c07a7ab567ea3d701a5a46a263ec0c

Observation 63c86595-7beb-41fc-8714-8c5862dd76cc · outbound

This paper cites Visual instruction tuning.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.521307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.602465Z digest=sha256:07a75fcec6c73835c7df903547b480924b648da68e41d747d3bc7434cade3d2b

Observation 46d60ad2-f810-4a4f-8652-873b21524d2e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.606338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.606338Z digest=sha256:f42540cabb1002f7a01f565c3dec03e8ae96c4f6d4be0803175adb5c0c77027a

Observation e792b842-d7cc-47c0-8aee-2433f36c1769 · outbound

This paper cites Least squares quantization in pcm.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Least squares quantization in pcm

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.510328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.610362Z digest=sha256:33048a8e596b87aac20a8bfe9eda463dbdc7b7de78c3234092530bb7b8eabd0c

Observation b8e8b71c-995f-4f41-82e0-0258167937c2 · outbound

This paper cites Robollm: Robotic vision tasks grounded on multimodal large language models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Robollm: Robotic vision tasks grounded on multimodal large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.498906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.614428Z digest=sha256:21bff31e614bf59b1794198d84dbd75a5bb5e514f175d6016e7e11d91663fbbf

Observation b697e2cc-3337-41b9-9c0f-d012ec4cbd5d · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.618303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.618303Z digest=sha256:b9302c685d271c56f9bf3a4ce51695b58f6167975ad227afd6a66a30c374e200

Observation e0b70180-7039-44b3-8fb8-d7bcb64d7eec · outbound

This paper cites Some methods for classification and anal- ysis of multivariate observations.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Some methods for classification and anal- ysis of multivariate observations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.479962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.622316Z digest=sha256:dfc023ab541d99b6199e066063800278bb382ec1c303b40dfa318e7f0e965e1c

Observation b123146d-63d4-4d92-8c87-a48a13b683dc · outbound

This paper cites DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.626572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.626572Z digest=sha256:97f94930c0dcce1d6a805295a19453bd66f5afd37f36a01c9ff3eee19fd66f06

Observation cdec5db0-8672-4016-9cd2-ebdc0d60e3f5 · outbound

This paper cites The impact of multimodal large language models on health care’s future.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs The impact of multimodal large language models on health care’s future

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.469117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.630546Z digest=sha256:905702b3d6345a97fa1b4f4019291f22858636007c01ebb0754c0869f405d04e

Observation 0637749e-4422-4456-9a3f-32347f4afa38 · outbound

This paper cites LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.634324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.634324Z digest=sha256:64c2d92cc5df48a93dab14ad933ed746644efb370f23b128055edfcd395ad155

Observation 535ad5b2-2954-46c8-93b2-702b2ad801d0 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Learning transferable visual models from natural language supervi- sion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.457193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.638395Z digest=sha256:c07647bf0dac9087b5aaa2ccc967a42310f146169d805ce25cc18805fef0d365

Observation 28a5b9f5-0173-4799-aa3e-cfad16624250 · outbound

This paper cites Designing network design spaces.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Designing network design spaces

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.642232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.642232Z digest=sha256:45da2647f036495268572be6127c461c57c90555c91be077d0428147958eb849

Observation 51d6f2cc-a7f1-4e0b-87cc-1a0dcf127825 · outbound

This paper cites Spearman’s rank correlation coefficient.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Spearman’s rank correlation coefficient

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.438857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.646711Z digest=sha256:a014acdac85aa66e54123299c3828ecce51f7de72e5b458bd4d8537eabbe9b26

Observation c150cfcc-ba0c-428a-bbdb-603e35be8f39 · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.650470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.650470Z digest=sha256:695b5bb587830c615c54a921370ab5c4eea983abdbc97bf7bdabe89382510d9a

Observation 38d302b1-57fa-4005-9473-e754cfc6471d · outbound

This paper cites Towards vqa models that can read.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Towards vqa models that can read

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.428280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.654635Z digest=sha256:44c5f7385cc0d5257834e4445e37476679761d9407aea988112ec3eb270858db

Observation 2aca5a95-6775-4eda-af5d-71e564924332 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.658305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.658305Z digest=sha256:69f83db471ca664c003d74ba1cbed4234ca6d1f439a5f469977053e742c1e895

Observation 1fd20365-a70f-4ce4-8f81-1ccd2ec40144 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.662035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.662035Z digest=sha256:898324f6596d16fc7285c261a6df25c1cdc2832193084fa44ddd44302cf87286

Observation f963ca89-69a0-452d-a544-64d57abc17ad · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.665325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.665325Z digest=sha256:d1ef502f720706667184690a43b0cf50131e3543626bb0b703a01d05d4583f42

Observation add6ff33-db20-4c88-8b2e-dc0fdba1a4c0 · outbound

This paper cites Teaching matters: Investigating the role of supervision in vision transformers.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Teaching matters: Investigating the role of supervision in vision transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.418620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.668778Z digest=sha256:ff3177a0309eb0e2f68a3b94792abcae57d56f127fab0822d65349f51566f649

Observation 975e5536-81a4-4778-a559-4a3bfa2652df · outbound

This paper cites Hierarchi- cal prompt learning using clip for multi-label classification with single positive labels.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Hierarchi- cal prompt learning using clip for multi-label classification with single positive labels

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.407469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.672156Z digest=sha256:bd3d5b7f040eaa72fa5a82ebaa0ced8be1a6cda1ecc2c1b585959fc3b1131727

Observation 3cf076f9-39f4-465f-8f30-de9f6feb38f3 · outbound

This paper cites Cait: Triple-win compression towards high accuracy, fast inference, and favorable transferability for vits.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Cait: Triple-win compression towards high accuracy, fast inference, and favorable transferability for vits

Reference 58

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:23:19.019714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.675939Z digest=sha256:2189f47198e2ae4276eef840da3ca9125210a4dbdfc4f0b130fee41975ff92db

Observation 9f9d1d74-e30a-4b10-a53c-3f52d4d93f73 · outbound

This paper cites Repvit: Revisiting mobile cnn from vit perspective.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Repvit: Revisiting mobile cnn from vit perspective

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.679260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.679260Z digest=sha256:d1340629366c7f94159a368d0fe42eb5fd9f1d4eeb7d21e40db56c94b8cad22f

Observation 9174ab73-291a-4aae-9be2-b27dc785481c · outbound

This paper cites YOLOv10: Real-Time End-to-End Object Detection.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs YOLOv10: Real-Time End-to-End Object Detection

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.683446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.683446Z digest=sha256:b047efa5b82642639f5c20c241fd952bab68069eb2b123d58a29e928e9f8d9eb

Observation 1c27f763-85ac-4383-aa07-0b32be338146 · outbound

This paper cites Large Language Models for Robotics: Opportunities, Challenges, and Perspectives.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Large Language Models for Robotics: Opportunities, Challenges, and Perspectives

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.687251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.687251Z digest=sha256:255249eb75df79cac6af55a6ef139d61c3388bf2e79fc4b2b9b2a016317ab102

Observation b17676b0-df9e-4568-9076-ca336662f59a · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs CogVLM: Visual Expert for Pretrained Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.691345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.691345Z digest=sha256:95c60d31d77b77f426c797c87f5bd90849269e96aaf1ed859f851df0462a86b4

Observation ab0c678a-7cf2-491e-b9b1-8dc5e47dff05 · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.695485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.695485Z digest=sha256:28b53e3c8bc606440d9a5b5b25a8edcb9a7905297ab86ca33b2dd31420654272

Observation 30b716bc-0e36-48b4-9771-2b3c199b3e7b · outbound

This paper cites A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.699407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.699407Z digest=sha256:c56dc48b5f7d643d2c076b6082e81aa4a7afed59c1839f78ab2b02e96454c2e4

Observation 39613f6b-db26-414f-8f4a-bf6c37f746f7 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Drivegpt4: Interpretable end-to-end autonomous driving via large language model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.705098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.705098Z digest=sha256:a899a48ec27dbf166823391c0327dc445b28d12f7de0d1c394b43a0369827799

Observation 984286a4-b366-405e-97c6-53abeeab1c45 · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.709285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.709285Z digest=sha256:c24ec343af00c7f9f41ae45cbfa1a7357717ca797c3fb6ce8eee20c695da4d1a

Observation 6bcebf37-bfbd-47d8-be5d-00105a08efd5 · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.713337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.713337Z digest=sha256:c7540b7d2551f750adb97a66dff250d0eafa5e28a638a966e9a7465c7bbe0f62

Observation 94685f31-4229-4530-9ccc-b34e367696c1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.717641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.717641Z digest=sha256:5c35da51a3e7ac449dcc195f647471594093bc3a3981a3dca6d4dcc846f6193d

Observation 97b20cd8-00d5-4aaf-92ac-9380db2105b9 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.721571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.721571Z digest=sha256:0f27093e42c711a99fcac91a4e5adee17d294384b86a82e2f863a7fd7cf18a62

Observation 306ad005-f887-40c6-9df7-590f6111247a · outbound

This paper cites Object recognition as next token prediction.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Object recognition as next token prediction

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.383263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.725720Z digest=sha256:555752ee2d4c2dc0bd711b386fed3ec0c1b44c5509cf1e65ed20780d632970bc

Observation b245ff0a-2cbe-4362-9aa4-30e236ab5991 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.729654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.729654Z digest=sha256:7d028332d372562bb9e5e254ef2e53a0576a69fa868a1aeb325d9688665dba2e

Observation 6b9e68a6-f9f5-4d1c-b5c0-98c06c15dbdb · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:23:18.733345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:23:18.733345Z digest=sha256:60587225e545903e37d2dee441dbb5ed30cf13f051347e4eac085fa2421a2dd3

Observation 93971e6d-2a07-44de-a0b5-a99af9012878 · outbound

This paper cites Llava-phi: Efficient multi-modal as- sistant with small language model, 2024.

[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs Llava-phi: Efficient multi-modal as- sistant with small language model, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:23:19.365120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T20:23:18.738520Z digest=sha256:063eab5630fd9b6c976deb5a4cea8849bf75335590a7a543eb5c234d451758c8

Pith citing papers

Observation cad0e01f-f301-4383-a6e4-48966203a249 · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.530971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.530971Z digest=sha256:198a172acc867b348f856573bbc2a610ce6598029bbb96387587b882ce23b0f5

Observation 927d1f5a-eb39-4d25-a0d6-5dd9d7e7c996 · inbound

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens cites this paper.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.330686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.330686Z digest=sha256:d0c5e1be524b978b2107dfae1b49b0e0c98437b667016b81bb08a1cc421c6b71

Observation 0684c49d-fb94-46a0-b763-5b078f39c65b · inbound

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models cites this paper.

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:09.626063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:09.626063Z digest=sha256:f1faac528a674ae94c4012710ea65add946a6c1306a075fcb8afcf59dd3d04db

Observation 34f56709-6ad3-47d4-80cb-9719785e5851 · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:32.847260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:32.847260Z digest=sha256:5121cfaa3619a827b3a6076155bb4d82731ad1275e69dc3600a4eda4217be723

Observation bf08dcb9-4566-433e-b52b-f72d9f749cab · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.503424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.503424Z digest=sha256:1315b577b557b4a8966d5be8491836be357086023ac58c909654bfa3f7ece041

Observation 0453dc7a-29bb-4b04-bbfe-a7ce4230f910 · inbound

Training-free Token Reduction for Vision Mamba cites this paper.

Training-free Token Reduction for Vision Mamba [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:16:49.428490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:16:49.428490Z digest=sha256:2b3a20ca71d08d22780d3ff647cd97de2a512c872da751df348b2e26397f20d7

Observation e9cfad62-f673-4087-bdbf-bcec45303e3c · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:04.956498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:04.956498Z digest=sha256:b1d261b76359f594f7da691503d645462962fb04f7e7e99927ecb8b98978246b

Observation e058a539-f6ac-4009-9e94-a32b58190d49 · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.388264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:6db4f60ef64c66ae7a858b2c13980e08abb2510bf7d69214ea562df157d97fd8

Observation d5a48be8-1ce3-429a-aea9-64a04bc5560f · inbound

Geometry-Guided 3D Visual Token Pruning for Video-Language Models cites this paper.

Geometry-Guided 3D Visual Token Pruning for Video-Language Models [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:09.962866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T05:49:38.346274Z digest=sha256:450a3d806912d31ae6cecc68c28c156247cce041cbb6a796ec754bc6d3360bb0

Observation e0112050-4ed6-437b-a8bc-307ebd8fa627 · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.407469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:d67d70659feefcd3f15f285fe3a732c98001e8684102a9e2fcf9dfac165672e8