Pith. sign in

Paper Citation Record · LEDGER

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 5 inbound Pith citation observations for arXiv:2411.14228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14228 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.837338Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:38:56.772419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:56:35.564771Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85981306-6ba3-47e3-962a-d3eab24d426f · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.877324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:36.997725Z digest=sha256:88572621bf29fde87cdf3b15ec425a3d98fefb0c49462271c74cd55a76958d81

Observation be6b8822-90b2-4eb1-910c-1e857f7702bb · outbound

This paper cites HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.002477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.002477Z digest=sha256:62fce726289e4d99b6500cefc3ab97e0b6ef82de012fd5b5524793720d72efea

Observation 44541421-4f6f-428c-b44a-f559bce6cead · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.720353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.007597Z digest=sha256:12651d8244a75258f155d18f7376dc8cb05b70a01a3c6829ada61ff0fef69f57

Observation 84fcb041-986c-48ad-bba5-6b218014217a · outbound

This paper cites Introducing our multimodal models, 2023.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Introducing our multimodal models, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.597815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.011954Z digest=sha256:57d49e9b7be7003e78a1190b204e0ae06181b774b063feb8330189443c5df5c9

Observation f18d397b-de95-4bc2-ae3b-3ab8387e1d59 · outbound

This paper cites Matryoshka Multimodal Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Matryoshka Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.016869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.016869Z digest=sha256:e00321faa27fcf14db9e4015508a8727ab40a26bf8178444a0106a6de1523f0d

Observation 6d6801c9-0478-49c7-9347-393dabd23a54 · outbound

This paper cites Madtp: Multi- modal alignment-guided dynamic token pruning for accel- erating vision-language transformer.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Madtp: Multi- modal alignment-guided dynamic token pruning for accel- erating vision-language transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.021723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.021723Z digest=sha256:b5c863e5dcb70618f7b937b2fb4cbce193573570fa8a6ad58f3d8bdc23d7e198

Observation e88b20a6-2d67-49ee-8bee-3bed8fda4fbc · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.026326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.026326Z digest=sha256:992f16fa7ad5dae2e9c2e2b7415843382066d1fb02c5adb22807b1a61c889a10

Observation 97077036-845c-4526-8562-84f99b63e1bc · outbound

This paper cites Allava: Harness- ing gpt4v-synthesized data for a lite vision-language model,.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Allava: Harness- ing gpt4v-synthesized data for a lite vision-language model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.570415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.030718Z digest=sha256:cc20fd10169832f0451f4fe3754e644890f661751ec6a7233c11763a8fc9b83a

Observation 9521270f-9bc6-4c7b-8053-beeedba0edeb · outbound

This paper cites Geoqa: A geomet- ric question answering benchmark towards multimodal nu- merical reasoning.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Geoqa: A geomet- ric question answering benchmark towards multimodal nu- merical reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.553874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.035534Z digest=sha256:984a86776a2af073dc42da15a1b63cc631b9070ec0720f6fad45f0b399b13df9

Observation 6bdbb30c-4332-4629-becb-aa9a8773fa17 · outbound

This paper cites Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.040643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.040643Z digest=sha256:5ee5b55c542d8d09164263aaab0d32d3736cc31dda6146ccbd341ec368f530f7

Observation 7212cb04-4fd8-42ff-91c5-b22068d28a7d · outbound

This paper cites Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.537965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.045721Z digest=sha256:df4245dbfd436014fe2d1277be0100c59aea343a1064f50ecafe3f4c85e8a32c

Observation b86bb992-3cc1-44cd-b325-9dec4103f217 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.050476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.050476Z digest=sha256:0639a9336ea20c20aaf3df3c4ef4a135c07dd70355f72b6e5c15c44354346781

Observation 2204a2cf-eeb2-4123-b637-d03e7f884779 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.075718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.075718Z digest=sha256:ae2a103317b3ad2a98b549e8b1c43933b26a67ff1bd0a410eb1a4bece6683d48

Observation c15604c6-b00f-4649-8317-86262a8f22bb · outbound

This paper cites an unresolved cited work.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:28:39.520715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.099273Z digest=sha256:a59e6ff9144d5648b5a6292daebd86c1ce63d21c31932941fe7f5fa62b220e8f

Observation 6d717807-1f12-4c24-bb31-d9a95e95ffa3 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.156258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.156258Z digest=sha256:6d44c664d27783aae573db57cbd6370c7f147970cfcaf6ba6ad6703fa466d240

Observation d7c1f498-9289-42b5-99ba-9d5d34673e90 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.381261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.205683Z digest=sha256:30c6709ac456b71ad78fbc6275836ed9b36b5c8fec0879e0f924c6f3e0b9bfb2

Observation 8cde052b-b0ad-451d-9de7-31cae2645616 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.209797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.209797Z digest=sha256:7505e7bf447954943513c5128c2e33610fe8842e27f44738375b15ed6891aafd

Observation e8c4dcd8-496a-4d67-bf48-0cea65accaed · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Matryoshka Query Transformer for Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.214098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.214098Z digest=sha256:2f4c11f0565a6ff5dc0c5cf1e4e69f6dab58ae0c3f4848b3a4c401a85d0e325f

Observation f7f679ab-04fa-4717-bd8a-835a9f7d7778 · outbound

This paper cites Bliva: A simple multimodal llm for better handling of text-rich visual questions.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Bliva: A simple multimodal llm for better handling of text-rich visual questions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.218265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.218265Z digest=sha256:00fcf1627099ec2cc703ec41dd4ee6a2ed6267b90ffd1613f7791e094232872e

Observation aa5d41bb-cc54-416c-bcc9-229d8d97e611 · outbound

This paper cites Hudson and Christopher D.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Hudson and Christopher D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.326081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.223393Z digest=sha256:4e6ba733330c75ef210d1c5f497fbc1110f2cd8b1e27fab5977ee45db3340223

Observation 835b0f59-ef62-4db7-89d3-326a2d250049 · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Dvqa: Understanding data visualizations via ques- tion answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.311235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.228109Z digest=sha256:ddc09172393bd28e163ce98c172f620ba30b57bcc4a87ff8ba6a9c66372f0f39

Observation c65a56df-4e85-4a4d-83f2-ccd473697ed5 · outbound

This paper cites A diagram is worth a dozen images.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression A diagram is worth a dozen images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.296116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.233751Z digest=sha256:b274df026b3ba137839ab3ff9b817aebcfbf2fca6f4ed0433509570744f2e1cf

Observation 4da35707-260f-41f2-b308-7ea4c9c3294b · outbound

This paper cites Ocr-free document understanding transformer.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Ocr-free document understanding transformer

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.281655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.238183Z digest=sha256:9f3780cce710c343c89d3b45882a1b702c78e7cb4ccc00f46925439237448dc8

Observation 07c3197e-4bb8-41db-a3ee-5068bb653d2e · outbound

This paper cites OtterHD: A High-Resolution Multi-modality Model.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression OtterHD: A High-Resolution Multi-modality Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.243021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.243021Z digest=sha256:3af2407181338e95782b68e019a34be98005b46c468c21a269db8390cb1ba1e0

Observation 496fa244-7501-45eb-97aa-517fc9a64638 · outbound

This paper cites an unresolved cited work.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:28:39.198504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.247894Z digest=sha256:2ca9792c412d03ee43df478925f553a7888b12b74c539b2bd1c5c3cdcf3e4f73

Observation b5fd66df-daab-4f6a-9bca-a0aaa5b13eec · outbound

This paper cites Evaluating object hallucination in large vision-language models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Evaluating object hallucination in large vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.151575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.252097Z digest=sha256:382232a24ede67e80427bfbd66257a357c12bb29a90e2d2feb15d00b9f3d9b86

Observation 10f5cb9f-cab3-4929-912b-d7b3d1096199 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.258045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.258045Z digest=sha256:1b4605c887c15f0c4203696b52967acb564e19cc2f77ee3fe7560d332b13f873

Observation d7bb0d8d-92da-486f-83a8-aea9522449da · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.137689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.263842Z digest=sha256:31a75c28bdd23655a0562d58ade889ec24230f6bc2bc95de28e7156cc21aadad

Observation 949bee26-8fbd-4dbb-808c-66e5fff87b8e · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.269526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.269526Z digest=sha256:68e9966ac4cf44606e80457940c2020c40349039c5475d1415f8e2b96be27df1

Observation 83a27c8e-2a75-4482-b212-2e70bf6e2fe5 · outbound

This paper cites Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Open-llava-next: An open- source implementation of llava-next series for facilitating the large multi-modal model community

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.123104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.274925Z digest=sha256:8d2c917f9417880e44b868f45ffc915338658fc9b199755f586fb6a13841bd4f

Observation b41d26f8-b5bd-4fe4-8482-4facd44c33fb · outbound

This paper cites SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.279341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.279341Z digest=sha256:e35ec22de6c25d42195c67c512fb7b594747af0a3e875455804c1534004203be

Observation 3b2d0748-67b0-40b2-8e29-7a8d8f00469a · outbound

This paper cites Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.284569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.284569Z digest=sha256:18b8736c8d57291a40ded0a036cdb4ea08b600c848b26bb4c088f8420e8333da

Observation 7c24928d-5701-4377-a67b-653a8e538ef6 · outbound

This paper cites Visual instruction tuning.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:39.054195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.289871Z digest=sha256:898a601448dc5a9a0bcc716eb8a63649ee6532462fde35525ff53348dab87dc4

Observation d4301e7b-0e88-45d2-bf19-4778c07ab290 · outbound

This paper cites Improved baselines with visual instruction tuning.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Improved baselines with visual instruction tuning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.903964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.362307Z digest=sha256:9de5d6add1b689c4c4af0ce065b9a84a1ecc5e851c6bc28fc7d9ad79d3873272

Observation fc93e0bd-653e-4f05-a1ff-96bfb8ef4c90 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.400560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.400560Z digest=sha256:20aa9b2af9db4dbe27906c75a65797e09c0ba7202f3b8623027406e5d1b7fb68

Observation f907f837-daa6-4e16-a89b-b5a628eb484e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression MMBench: Is Your Multi-modal Model an All-around Player?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.424539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.424539Z digest=sha256:e05e0d20fda04fc920c96ea0e114b90639a34f06e63a4ff580837b7f52d70280

Observation 2086919b-ee04-48d5-a617-3364ef15fbc9 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.430152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.430152Z digest=sha256:5592115462fdf64f5ffc9e8e99e2d41add4f55e283afbc1a1573b626c6fe4250

Observation 3e31518c-808a-4dbb-be17-301aa00dfb90 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.876420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.434873Z digest=sha256:8cd2a8d5824db77d5f21b8bc36781b73462cc7f69020e4c3a7553bc86d14ad72

Observation 266936f7-9d45-42fd-bf87-d890353fd29a · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.439399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.439399Z digest=sha256:6297e0702f965b5db0878679b6d4c2e8ef296d5881d7cb10e47f7aded2fdcff3

Observation e30d2dee-8f06-461a-af90-96c6184f7bfa · outbound

This paper cites Joty, and Enamul Hoque.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Joty, and Enamul Hoque

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.444054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.444054Z digest=sha256:c1d17a07e75a8035a294f3360d6da11f29627db0f07b34ade293385b525c0a14

Observation 8e8dca30-61ce-4018-a297-bb4ecd828cc2 · outbound

This paper cites an unresolved cited work.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:28:38.732723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.452968Z digest=sha256:1266ca430023969cfbfc5903f0e44189067cb823fa4de01191cd226d5d7df475

Observation d5548868-4c7d-49a8-a8f9-b94fbeb9a2a0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Learning transferable visual models from natural language supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.457732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.457732Z digest=sha256:bfcd8b3b2a198ce049e93f8c2116c2958a06afc295a244c227c690e2c7e88834

Observation ac31051f-4420-4236-b7dd-9be4eef286ba · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient large multimodal models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Llava-prumerge: Adaptive token reduction for efficient large multimodal models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.462581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.462581Z digest=sha256:425f01e48c1a6674fe89403b535eb7f8cc10def5e544917b1d40954fd4c88d8d

Observation d3b18de5-609a-4227-9cdf-fd753c2f34c9 · outbound

This paper cites Towards VQA models that can read.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Towards VQA models that can read

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.621897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.467134Z digest=sha256:c8418b338c482c3d8b8967af839cff0dfdd6f1a8c062eb612aef2486a080d38b

Observation 9663235a-df3b-437f-95f5-446efddfc313 · outbound

This paper cites Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.471245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.471245Z digest=sha256:ca34a4f36a3a5a7585a300441afd20bf87608e0934d27bc1b284f83d65ef2dcf

Observation 0735f891-3805-45f3-9512-989adea90ff9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.476018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.476018Z digest=sha256:9972c4e2c5b0ee47ddd7475cb814fdf2aaa2e3b5d4a8fa80d0443039e340cb2a

Observation 68937910-ccfd-4f15-831d-2fa37fd7d194 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.480933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.480933Z digest=sha256:abe55993b938e1f09f3a2c509ff54591a49176a5dff9746f24065e509eba65c9

Observation adf3e0c0-7b68-4965-945b-5221e8622735 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.486326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.486326Z digest=sha256:92742b47c18f69b150d26248ca4da5e0da5c2db2abf3d74ccf2e04b635dc81d4

Observation 68cb61a9-185a-4ec7-b68b-010ddd4ca265 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.491494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.491494Z digest=sha256:5469866737d7007a13fd59dfce252cf084a2d580fdd05af41add13d8a8906e8d

Observation c6ce1aae-48ec-4cef-8db7-c1a9c615a421 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Mm-vet: Evaluating large multimodal models for integrated capabilities

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.573491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.496679Z digest=sha256:205ccdba4a055b77fe8d9b5e31c0a61aa89384b81d90cac581c5b1b32502b5fb

Observation c4ce4de7-7a45-411c-a2a9-754906917c48 · outbound

This paper cites TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.501180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.501180Z digest=sha256:b496d18f3aa47ce14ad2402a9c6f8d1f4cf8d080fdf3561e448b4160e04b708c

Observation fac41a21-f0be-49ee-aa91-a5f5766ced3d · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.529815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.529815Z digest=sha256:953a7a45ebf5ecd29f8441a4597976b923b0406793715d91d5e8ecfe0db0ed03

Observation 1776f1c2-6979-4d83-a4ee-428e21b33d76 · outbound

This paper cites Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.620931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.620931Z digest=sha256:a6d9cb6b8d05e84b04d408de789185b44db1e43f40dc93caf34af56b211e4743

Observation c85c23ed-b540-42fa-b5ae-eb93a40efafd · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.702245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.702245Z digest=sha256:539a5628da3fdca2b9dbab12b3beee93bc645a71108127305a7d4aeda8c599a7

Observation 1009fd42-b04b-4819-bfc0-7c95510b4de7 · outbound

This paper cites This is used to illustrate that a learned metric rather than hand-crafted will solve the prob- lem of performance reduction.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression This is used to illustrate that a learned metric rather than hand-crafted will solve the prob- lem of performance reduction

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.558449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.819204Z digest=sha256:dafb9086169c60b4a4b21270fba53e9c3f9eab06aeb0b124efb1e962b58c9746

Observation 65663ce6-e6ef-43a3-8c39-0f7064f82c4e · outbound

This paper cites 6 and Fig.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression 6 and Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.542583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.825560Z digest=sha256:f62a1f8fa8fc6d03811d27e60a5a045392cbcba6da342fcc164cb12915a5b27a

Observation e2054ada-94f5-42f5-b2fd-1bb84204a22b · outbound

This paper cites an unresolved cited work.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:28:38.524062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.830168Z digest=sha256:2b104f308032c74a61939188927a80f0d14a923dc210c8288628a6ee9a333128

Observation 81b7ee5d-849d-44fa-bec4-24c0724dfa62 · outbound

This paper cites During our training process, under the constraint of balance loss, the model is required to select three different visual scales with as equal probabil- ity as possible.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression During our training process, under the constraint of balance loss, the model is required to select three different visual scales with as equal probabil- ity as possible

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:28:38.426533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.837338Z digest=sha256:bee05fd8e133d7da09f471cfd832054451480e149813cb708e38f3e94877f9f5

Observation 0ba032dd-de4c-42e2-a7e7-fd035bfe3871 · outbound

This paper cites an unresolved cited work.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Unresolved cited work

Reference 2279

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:28:38.851506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T15:28:37.448526Z digest=sha256:152f0ceb9b99d53f69f4bdb9dcf4181d10c37805a0b4e6ab907734acde7d70d9

Pith citing papers

Observation 81522011-0c70-4c50-b4e6-f46a7adcdf3e · inbound

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information cites this paper.

LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:38:56.772419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:38:56.772419Z digest=sha256:3adeae63324b6c82e74a579d529bd65f4352ab5125eeb1a45b8d12ba3b6c20e9

Observation 01988274-4b3b-42ef-bd1b-1caf36fbd444 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:06.528428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:06.528428Z digest=sha256:a6f80fbcc46ef661b60633b969c82cfcadf6a17de8936c204a09ebcc82bc9521

Observation dfa3f6f5-d21d-4709-aecc-aeed2acded77 · inbound

An Efficient Token Compression Framework for Visual Object Tracking cites this paper.

An Efficient Token Compression Framework for Visual Object Tracking FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:35.600359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:27:18.535624Z digest=sha256:5834298db536d2aca779a16463f6583ea2eecaea362e2757407928d9dc94b701

Observation 924ecc97-c796-401d-83cf-d8baec19dfba · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.296228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.296228Z digest=sha256:a8ae624349bdf6fb44ac5094a7a80de346eb32a7bc05dca86f55d1591992ac51

Observation c4bfe39b-061b-4146-86c4-abc0b20949bd · inbound

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin cites this paper.

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T04:30:14.375574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:30:14.375574Z digest=sha256:451b87ccd09a43adaf12e71469c0524ec4f2b93d1dee13aaf1215f3b8f288c7e