Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.837338Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 5 inbound Pith citation observations for arXiv:2411.14228.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.837338Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:38:56.772419Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T07:56:35.564771Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 85981306-6ba3-47e3-962a-d3eab24d426f · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be6b8822-90b2-4eb1-910c-1e857f7702bb · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44541421-4f6f-428c-b44a-f559bce6cead · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 84fcb041-986c-48ad-bba5-6b218014217a · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Introducing our multimodal models, 2023
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f18d397b-de95-4bc2-ae3b-3ab8387e1d59 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Matryoshka Multimodal Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d6801c9-0478-49c7-9347-393dabd23a54 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Madtp: Multi- modal alignment-guided dynamic token pruning for accel- erating vision-language transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e88b20a6-2d67-49ee-8bee-3bed8fda4fbc · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97077036-845c-4526-8562-84f99b63e1bc · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Allava: Harness- ing gpt4v-synthesized data for a lite vision-language model,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9521270f-9bc6-4c7b-8053-beeedba0edeb · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Geoqa: A geomet- ric question answering benchmark towards multimodal nu- merical reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6bdbb30c-4332-4629-becb-aa9a8773fa17 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7212cb04-4fd8-42ff-91c5-b22068d28a7d · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b86bb992-3cc1-44cd-b325-9dec4103f217 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2204a2cf-eeb2-4123-b637-d03e7f884779 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15604c6-b00f-4649-8317-86262a8f22bb · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d717807-1f12-4c24-bb31-d9a95e95ffa3 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression An image is worth 16x16 words: Transformers for image recognition at scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c1f498-9289-42b5-99ba-9d5d34673e90 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8cde052b-b0ad-451d-9de7-31cae2645616 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c4dcd8-496a-4d67-bf48-0cea65accaed · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Matryoshka Query Transformer for Large Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f679ab-04fa-4717-bd8a-835a9f7d7778 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Bliva: A simple multimodal llm for better handling of text-rich visual questions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa5d41bb-cc54-416c-bcc9-229d8d97e611 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Hudson and Christopher D
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 835b0f59-ef62-4db7-89d3-326a2d250049 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Dvqa: Understanding data visualizations via ques- tion answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c65a56df-4e85-4a4d-83f2-ccd473697ed5 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression A diagram is worth a dozen images
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4da35707-260f-41f2-b308-7ea4c9c3294b · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Ocr-free document understanding transformer
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07c3197e-4bb8-41db-a3ee-5068bb653d2e · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression OtterHD: A High-Resolution Multi-modality Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 496fa244-7501-45eb-97aa-517fc9a64638 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b5fd66df-daab-4f6a-9bca-a0aaa5b13eec · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Evaluating object hallucination in large vision-language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 10f5cb9f-cab3-4929-912b-d7b3d1096199 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7bb0d8d-92da-486f-83a8-aea9522449da · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 949bee26-8fbd-4dbb-808c-66e5fff87b8e · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a27c8e-2a75-4482-b212-2e70bf6e2fe5 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b41d26f8-b5bd-4fe4-8482-4facd44c33fb · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2d0748-67b0-40b2-8e29-7a8d8f00469a · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c24928d-5701-4377-a67b-653a8e538ef6 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Visual instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d4301e7b-0e88-45d2-bf19-4778c07ab290 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Improved baselines with visual instruction tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fc93e0bd-653e-4f05-a1ff-96bfb8ef4c90 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f907f837-daa6-4e16-a89b-b5a628eb484e · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression MMBench: Is Your Multi-modal Model an All-around Player?
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2086919b-ee04-48d5-a617-3364ef15fbc9 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e31518c-808a-4dbb-be17-301aa00dfb90 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 266936f7-9d45-42fd-bf87-d890353fd29a · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30d2dee-8f06-461a-af90-96c6184f7bfa · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Joty, and Enamul Hoque
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8dca30-61ce-4018-a297-bb4ecd828cc2 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d5548868-4c7d-49a8-a8f9-b94fbeb9a2a0 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Learning transferable visual models from natural language supervision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac31051f-4420-4236-b7dd-9be4eef286ba · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b18de5-609a-4227-9cdf-fd753c2f34c9 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Towards VQA models that can read
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9663235a-df3b-437f-95f5-446efddfc313 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0735f891-3805-45f3-9512-989adea90ff9 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68937910-ccfd-4f15-831d-2fa37fd7d194 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf3e0c0-7b68-4965-945b-5221e8622735 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68cb61a9-185a-4ec7-b68b-010ddd4ca265 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ce1aae-48ec-4cef-8db7-c1a9c615a421 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Mm-vet: Evaluating large multimodal models for integrated capabilities
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4ce4de7-7a45-411c-a2a9-754906917c48 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac41a21-f0be-49ee-aa91-a5f5766ced3d · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1776f1c2-6979-4d83-a4ee-428e21b33d76 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c85c23ed-b540-42fa-b5ae-eb93a40efafd · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1009fd42-b04b-4819-bfc0-7c95510b4de7 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression This is used to illustrate that a learned metric rather than hand-crafted will solve the prob- lem of performance reduction
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 65663ce6-e6ef-43a3-8c39-0f7064f82c4e · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression 6 and Fig
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e2054ada-94f5-42f5-b2fd-1bb84204a22b · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 81b7ee5d-849d-44fa-bec4-24c0724dfa62 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression During our training process, under the constraint of balance loss, the model is required to select three different visual scales with as equal probabil- ity as possible
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0ba032dd-de4c-42e2-a7e7-fd035bfe3871 · outbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work
Reference 2279
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 81522011-0c70-4c50-b4e6-f46a7adcdf3e · inbound
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01988274-4b3b-42ef-bd1b-1caf36fbd444 · inbound
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa3f6f5-d21d-4709-aecc-aeed2acded77 · inbound
An Efficient Token Compression Framework for Visual Object Tracking FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 924ecc97-c796-401d-83cf-d8baec19dfba · inbound
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · inbound
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.