Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:38.969249Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.10548.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:09:38.969249Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T00:15:58.111384Z
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5f32fe38-422a-4bc1-9057-a9cef7fd7d38 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e7e21f-8336-429d-9e68-3334dfe51165 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06fd9171-731a-4be9-b22d-45c42578b975 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98031748-f8f0-4cb4-8c1a-d4615879ced4 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f600248-06ad-45f3-8739-0fa3a07e4d5a · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination of multimodal large language models: A survey, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc41d65e-3770-4a33-987a-68cb1b75de11 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Geopqa: Bridging the visual perception gap in mllms for geometric reasoning,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19e051c-5c62-4334-8c69-0282bd17e0c3 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d70d8b-6a42-4d6c-aacb-1020432feb1f · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Nacl: A general and effective kv cache eviction framework for llms at inference time, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 415163b8-b131-49aa-b806-ec6dac56b0f9 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Inner thinking transformer: Lever- aging dynamic depth scaling to foster adaptive internal think- ing, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132056ad-50ea-4d42-9024-82dddab03ac2 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76eac7fa-b439-4f1e-adf3-e55ada39e82a · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Spatial- rgpt: Grounded spatial reasoning in vision language models,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29e3834-45fb-4375-8824-f45b17b54b14 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Atlas: Mapping attention’s location and size to probe five modes of serial and parallel search.Attention, Perception, & Psychophysics, 86(6):1938–1962, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9dac77-1ad7-410a-848f-e63800e634cc · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Image super-resolution using deep convolutional net- works.IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 225a961d-d473-4f67-acc4-d3a208aaeed5 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Acceler- ating the super-resolution convolutional neural network
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac304ad-c02f-4cbc-a3a3-d9701695e2d6 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Multi-modal hal- lucination control by visual information grounding, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15a0393-fed7-4933-8ab4-c68253dc1815 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04eaf8d-4385-4bde-aa91-7c5edb3840bc · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mme: A comprehensive evaluation benchmark for multimodal large language models, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea7cf24-597e-4089-aa6b-7b1d31388ee5 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Tracking the will to attend: Cortical activity indexes self-generated, voluntary shifts of attention.Attention, Perception, & Psy- chophysics, 78(7):2176–2184, 2016
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26d508e5-4d87-448e-927d-efc006131837 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Beamlora: Beam-constraint low-rank adap- tation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f18c14b3-1ee4-4bcc-88c7-13c64e9d650d · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0fdf4e-338d-489c-9b9b-911e6d7c33d1 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57913d8a-3348-4335-b498-5364c074a75c · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Hallucination augmented contrastive learning for multimodal large language model, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d7c207d-cb2d-42d1-a6d0-27320cf6055e · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Cortical mechanisms for shifting and hold- ing visuospatial attention.Cerebral cortex, 18(1):114–125,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8df33330-5259-4fed-bac0-3ba32b12bcdc · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Accurate image super-resolution using very deep convolutional net- works
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb761497-fc53-491e-ac5c-fced096a0003 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c64248e-1ef1-4283-96f7-4f6d05e4eb20 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba5dbffd-f91d-4d93-9d2c-da234d9df531 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae0dc74-3ea3-484f-a6f7-c87df8108fa0 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4cd6922-8f0c-40f3-9f69-f5992716d92e · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Evaluating Object Hallucination in Large Vision-Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88d4203-8a5c-47a8-b28a-b9da0dce8b4d · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Microsoft coco: Common objects in context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263bf065-c9b1-4704-87cc-d0fdd74c2cdb · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf27882a-4faf-48c9-91f4-f0d0a73e5fb0 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Improved baselines with visual instruction tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24e4044-c131-42e5-8cf8-0d76a396a91b · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faba4aad-1c8f-4096-bcf6-b7274439f3bd · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 114edba4-7afe-4b80-883b-69254ee70302 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b184aa7-8a20-44d6-a001-f40884234dbb · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Neuronal mechanisms of visual atten- tion.Annual review of vision science, 1(1):373–391, 2015
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f33d94-5808-4029-8137-268054a72f27 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Ocr-vqa: Visual question answering by reading text in images
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce4b50ba-1971-473a-9dc8-5acc4e0474ef · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Learning transferable visual models from natural language supervi- sion
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8c6013-d94d-44a2-aadc-f06a26411e2b · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8345c7-e171-4208-a687-284a42c9e5ae · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Towards vqa models that can read
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c710be-23b8-4ffa-9a21-abb42453bfef · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c66680-208d-422b-8114-03d3941c5080 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5640b0b4-4a97-406b-bbbe-6fce424be439 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec4e44d-88e9-4bb4-b788-08e4eeea8147 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Transformers: State-of-the-art natural language processing
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 978dc477-32ba-4afb-8e4c-f600cbfa318b · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8e008f-e00e-404c-a8af-c2ad5d00a78a · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Grounded chain-of-thought for multimodal large language models,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6635180d-9c07-4ea3-8288-e28d31e73a4d · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8331d809-d8a9-4c1d-87c5-82b7ec5e5952 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Fit and prune: Fast and training-free visual token pruning for multi- modal large language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a86ab97-3722-41b1-92e4-34065fbedba8 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Introducing Visual Perception Token into Multimodal Large Language Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6baefc-31dc-4d52-91ae-aa6d452ffafe · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52371a29-5dbb-4375-95d8-72bfd611be92 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0cae116-48b6-49e8-9be2-be8c25934b68 · outbound
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding Open eyes, then reason: Fine-grained visual mathematical understanding in mllms, 2025
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa65d30c-7fb7-470b-95d6-9e823029c87d · inbound
Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration? Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.