Pith. sign in

Paper Citation Record · LEDGER

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

As of 11 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.10548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.10548 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:38.969249Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T00:15:58.111384Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f32fe38-422a-4bc1-9057-a9cef7fd7d38 · outbound

This paper cites GPT-4 Technical Report.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:31.969501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:31.969501Z digest=sha256:47ffe880153bed79703e2226ff99f51ad189388fd01943b9283359bb01bd67d0

Observation e5e7e21f-8336-429d-9e68-3334dfe51165 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.124176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.124176Z digest=sha256:e3244d8e84b3d3a8f7562895c3b913b6ea16c8dc2d5a590d2161998de134c044

Observation 06fd9171-731a-4be9-b22d-45c42578b975 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.253101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.253101Z digest=sha256:33b7637022d0c36ac08f5a4d02e6a993f4aa200c2a03b693a1e9cce43c1bab13

Observation 98031748-f8f0-4cb4-8c1a-d4615879ced4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.364961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.364961Z digest=sha256:56add1daf80dde3f8b6d9fe3d075377357e7dcd5793161270d935245bae3aac4

Observation 0f600248-06ad-45f3-8739-0fa3a07e4d5a · outbound

This paper cites Hallucination of multimodal large language models: A survey, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination of multimodal large language models: A survey, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.516659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.516659Z digest=sha256:aaf31ca12fdb13027b48c0d5e96e41e5035b58999b694fe8afbaeb09f9cc12f1

Observation cc41d65e-3770-4a33-987a-68cb1b75de11 · outbound

This paper cites Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.630525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.630525Z digest=sha256:baaa1d904c45c0b90162a7d6e31c255a1b79fa40e361ff6f7d13a355018408b9

Observation f19e051c-5c62-4334-8c69-0282bd17e0c3 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.756479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.756479Z digest=sha256:6c8df599157a843543f5c1dc0184de60fcf28576368e05bc8a4f60d147cfe3a6

Observation d6d70d8b-6a42-4d6c-aacb-1020432feb1f · outbound

This paper cites Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:32.914509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:32.914509Z digest=sha256:cfd816f73e859db8ba8cd9847f51e08405c61fcc575676586376bba71787bea4

Observation 415163b8-b131-49aa-b806-ec6dac56b0f9 · outbound

This paper cites Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.033914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.033914Z digest=sha256:0431ebc74eb506df9f801c7e243baadf78076c3c18c0a49ca4fffa316a3d0bcc

Observation 132056ad-50ea-4d42-9024-82dddab03ac2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.154299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.154299Z digest=sha256:f6ec942c1414fe9bab1f72dd70126d53b6f6a234920f8819eb4d820e14059a40

Observation 76eac7fa-b439-4f1e-adf3-e55ada39e82a · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision language models,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Spatial- rgpt: Grounded spatial reasoning in vision language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.283044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.283044Z digest=sha256:96782f3aa08d66326c5a99e566711a9b1f8550f30afa95e2e18ed8305daebdba

Observation a29e3834-45fb-4375-8824-f45b17b54b14 · outbound

This paper cites Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.398979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.398979Z digest=sha256:8c8730d7cdfc3f982b1131c37aadef3f76fd214273fa313c429b0a5a7b7702d4

Observation ed9dac77-1ad7-410a-848f-e63800e634cc · outbound

This paper cites Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.555242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.555242Z digest=sha256:effeacd76d50b6e9e76ee8e9cbfad78cada2016e211748b5bd574b84567a77e6

Observation 225a961d-d473-4f67-acc4-d3a208aaeed5 · outbound

This paper cites Acceler- ating the super-resolution convolutional neural network.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Acceler- ating the super-resolution convolutional neural network

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.660599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.660599Z digest=sha256:6b6c9da9c774e2749557155b0afe1aa9a809b4c57bbd18dbdee38dfcdfb1accd

Observation 6ac304ad-c02f-4cbc-a3a3-d9701695e2d6 · outbound

This paper cites Multi-modal hal- lucination control by visual information grounding, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Multi-modal hal- lucination control by visual information grounding, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.746918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.746918Z digest=sha256:bc32b7cd8e2664e3a3848e813485acc57a20e4dc73a2bc89b5caead71ea0a1b5

Observation c15a0393-fed7-4933-8ab4-c68253dc1815 · outbound

This paper cites DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:33.923627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:33.923627Z digest=sha256:e3fd0ae7d4c7d64a5fcdb96674bf2b01e2e4d7b09878543b1e2cf0345de5f653

Observation d04eaf8d-4385-4bde-aa91-7c5edb3840bc · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.073705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.073705Z digest=sha256:0c22a0cb7f47a09f8f1aa41be1aae7c017e471834ba50a88857fe826df219ca7

Observation 8ea7cf24-597e-4089-aa6b-7b1d31388ee5 · outbound

This paper cites Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.244703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.244703Z digest=sha256:813b736776aaf85827279b11fea1c6dc54d61cd1af373ee2d5d2c9f02f5f5b0a

Observation 26d508e5-4d87-448e-927d-efc006131837 · outbound

This paper cites Beamlora: Beam-constraint low-rank adap- tation.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Beamlora: Beam-constraint low-rank adap- tation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.426101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.426101Z digest=sha256:ef933a6fd966a7feb0260bb5c0ebf86aa4e579d2894f3d87aedecc27a7b84f75

Observation f18c14b3-1ee4-4bcc-88c7-13c64e9d650d · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.507754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.507754Z digest=sha256:860df5f4ac7f047ad5f71dea2fd7c0a00f6d5f3070127ee9929f436d17ca3164

Observation cd0fdf4e-338d-489c-9b9b-911e6d7c33d1 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.700795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.700795Z digest=sha256:39abd8fbedb6c288c898bbff4f1bdf89f709fc60d1c0bbe2cb8693c65d3e7681

Observation 57913d8a-3348-4335-b498-5364c074a75c · outbound

This paper cites Hallucination augmented contrastive learning for multimodal large language model, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination augmented contrastive learning for multimodal large language model, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:34.887592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:34.887592Z digest=sha256:62174009526dd629f135afbd41813ebbe698a8cde71ecd81b636f7bbe84fc8fc

Observation 3d7c207d-cb2d-42d1-a6d0-27320cf6055e · outbound

This paper cites Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.029874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.029874Z digest=sha256:8c865a64977263691e5ae044c9b3f0b23979e12918b7a31df847b7f5e0d4488d

Observation 8df33330-5259-4fed-bac0-3ba32b12bcdc · outbound

This paper cites Accurate image super-resolution using very deep convolutional net- works.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Accurate image super-resolution using very deep convolutional net- works

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.187960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.187960Z digest=sha256:00d6133d5a7f97f3a5d63365e11d2a31763136612007250573f00ddd1d9f34cc

Observation bb761497-fc53-491e-ac5c-fced096a0003 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.331706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.331706Z digest=sha256:ee271a2b80f3caf5ff5c90a9f427ae0d87b4a1d2106c1e22705f502867387366

Observation 6c64248e-1ef1-4283-96f7-4f6d05e4eb20 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.462066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.462066Z digest=sha256:c7ce63f32051ec2cbd60ceb699f772d19ded9a33f6f2a8c396ae3779654c9c99

Observation ba5dbffd-f91d-4d93-9d2c-da234d9df531 · outbound

This paper cites Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.645936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.645936Z digest=sha256:deb3f5b950bf8ffdc9abf6fa95aa7a57f4360212c9d8d2c0f3af54ade1c57782

Observation 8ae0dc74-3ea3-484f-a6f7-c87df8108fa0 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.790140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.790140Z digest=sha256:b223a65b8c8b83ae32c39f050f7515ae644dbfb8ad1c281c0e61fbd9916bdb08

Observation c4cd6922-8f0c-40f3-9f69-f5992716d92e · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:35.981648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:35.981648Z digest=sha256:3536daf38ed67e8693317d94a20871461d570370fbc1d0330e47337962b5e4f8

Observation b88d4203-8a5c-47a8-b28a-b9da0dce8b4d · outbound

This paper cites Microsoft coco: Common objects in context.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Microsoft coco: Common objects in context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.135377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.135377Z digest=sha256:464de5f556a734cbd032170eb973e8a6702841d651dfb028d02167ea89c09b4b

Observation 263bf065-c9b1-4704-87cc-d0fdd74c2cdb · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.260287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.260287Z digest=sha256:34973fcc99a53df4fc4933e42bd46c342a1f231e31bb1eb083ca8155b3905587

Observation bf27882a-4faf-48c9-91f4-f0d0a73e5fb0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Improved baselines with visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.441445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.441445Z digest=sha256:a0315f9a808e9c77022e28f8e04e85b2fe972ec523de30e7e1c5c9090166b797

Observation c24e4044-c131-42e5-8cf8-0d76a396a91b · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.559709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.559709Z digest=sha256:92f92f7fabae2269c98123a3ab087b8567aedb1b0c37e7698c26e5c5521c4ca1

Observation faba4aad-1c8f-4096-bcf6-b7274439f3bd · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.675881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.675881Z digest=sha256:9a8cd4df1e08ec9db547c8632f0235650523d36e70d0425b39b465cc951277e0

Observation 114edba4-7afe-4b80-883b-69254ee70302 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.738767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.738767Z digest=sha256:a645badde787275ab6ae8b3ea83c283d60e6c6ad469d839578c6a977ce75ebeb

Observation 1b184aa7-8a20-44d6-a001-f40884234dbb · outbound

This paper cites Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.859880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.859880Z digest=sha256:eb2b70a77a84094078a47008dc0f923f88344e8ec309fcd81bd01981830545e9

Observation 98f33d94-5808-4029-8137-268054a72f27 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Ocr-vqa: Visual question answering by reading text in images

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:36.974350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:36.974350Z digest=sha256:f173873bc4033fb48227e12886a30ed61435c82f89e5551c4cf31d43f965667e

Observation ce4b50ba-1971-473a-9dc8-5acc4e0474ef · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.107535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.107535Z digest=sha256:5a630a01b183845f3395c067241e8f24f5f6fec95d7f40495f2222f3549c146b

Observation ec8c6013-d94d-44a2-aadc-f06a26411e2b · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.236417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.236417Z digest=sha256:bd4961db0dafaec1e266e602934a2b2650247f4b8e781d023ea588a732baf25a

Observation ac8345c7-e171-4208-a687-284a42c9e5ae · outbound

This paper cites Towards vqa models that can read.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Towards vqa models that can read

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.399420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.399420Z digest=sha256:70e75c2ce69bca00c15ab8fea20da8cde37bf7dac6392d8cd2804d19ca475061

Observation d6c710be-23b8-4ffa-9a21-abb42453bfef · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.581401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.581401Z digest=sha256:dfc40f82e1c7f9986891fdc789bc46a1c919f3979466fd3a493152b04be1fdd8

Observation 11c66680-208d-422b-8114-03d3941c5080 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.702564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.702564Z digest=sha256:5f458952584c487f01d69926570c77c91f740706ed8184cf9de5822baea8d0e6

Observation 5640b0b4-4a97-406b-bbbe-6fce424be439 · outbound

This paper cites VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:37.841750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:37.841750Z digest=sha256:f05374ba71d5b7bf3684a6e524f8671ca3a95b05bad3eefe5332768db2240e1e

Observation 2ec4e44d-88e9-4bb4-b788-08e4eeea8147 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Transformers: State-of-the-art natural language processing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.025746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.025746Z digest=sha256:0cb90b82e2121941e3646539357a1aa9ee24cf87cc1686217e04b1f4b3fcb1aa

Observation 978dc477-32ba-4afb-8e4c-f600cbfa318b · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.165356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.165356Z digest=sha256:315b19e9ad55abd7fc83bc67769ac8eedee5071477b389fd287c7e74aee9f970

Observation 8c8e008f-e00e-404c-a8af-c2ad5d00a78a · outbound

This paper cites Grounded chain-of-thought for multimodal large language models,.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Grounded chain-of-thought for multimodal large language models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.306215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.306215Z digest=sha256:61e4675cd7161adb8fc1ded54657354612975ee9e8538b44e7b979d8fdf6a4b3

Observation 6635180d-9c07-4ea3-8288-e28d31e73a4d · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.421930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.421930Z digest=sha256:a5e627e9079eb3c8cd6ba577527e41fbb34b2db9f4ddb4c37d8b641ab99f108b

Observation 8331d809-d8a9-4c1d-87c5-82b7ec5e5952 · outbound

This paper cites Fit and prune: Fast and training-free visual token pruning for multi- modal large language models.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Fit and prune: Fast and training-free visual token pruning for multi- modal large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.563170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.563170Z digest=sha256:5fae309a16efe9d290c4eef5f76351b971de8b38d5167564a000dbb971f8ad13

Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.749150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.749150Z digest=sha256:7e89eea6bea96b1f4ed2f353ea4fd8355e3e5416b38d035cd063578834d2206e

Observation 1c6baefc-31dc-4d52-91ae-aa6d452ffafe · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.835314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.835314Z digest=sha256:a98eb47646340f49ed2416376c6ce4217f66d521d21c7320fa9a2309f9f970a1

Observation 52371a29-5dbb-4375-95d8-72bfd611be92 · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.897509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.897509Z digest=sha256:5e2f15d9d0a255b4b8152f8c827951bbccdcf2c265d4298ecafc9fc5bd54313b

Observation d0cae116-48b6-49e8-9be2-be8c25934b68 · outbound

This paper cites Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.969249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.969249Z digest=sha256:5aabdcdf7ebf28be4556e90898819fe22feb04ef0a90a7e0c9aefc05110b8925

Pith citing papers

Observation fa65d30c-7fb7-470b-95d6-9e823029c87d · inbound

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? cites this paper.

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T00:15:58.111384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:15:58.111384Z digest=sha256:87c125eb1bbf50e0682bd3e361d2dff48fe5326a5a036fe70046f093c6487f5e