Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:35.027792Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2506.12776.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:35.027792Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T21:58:53.702009Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:37:14.448914Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e0e82b7c-81b1-41c4-a0b5-6c052a6f472b · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6369f684-748a-4e43-9414-9bc5a166ee8b · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ce4dc1-81d8-4543-8384-df02b0d133c3 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a605ce2-499d-4e75-a0b5-7159b53b774b · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocean-OCR: Towards General OCR Application via a Vision-Language Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ffee03-72d9-4c73-9c81-4513847c6dcf · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d0afd1-97d5-4d80-80f4-e03ef8c6561c · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b260ec6d-0bbb-4899-8e6c-51b9aff3a854 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c52b55-10e8-4431-bf04-c0040d6601aa · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e0b855-cade-44c4-8722-8e0b0d8f28aa · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Bert: Pre-training of deep bidi- rectional transformers for language understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc9fdcc-c093-4040-8fc8-8c54c2b820d5 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d6d148-e214-4a23-b785-302d79a40750 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Gpt-3: Its nature, scope, limits, and consequences
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e32e37fb-45e4-4006-ada8-2c8766dc3bcc · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bd0a28-93a8-4458-9fc0-5186a0340428 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2629bd-4eee-4440-90d2-4227f3240ef6 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Seed1.5-vl technical report, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 383e8978-26b3-4a4a-8418-e1043bfa123e · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7ff2c4b-6a7e-4a70-a0bc-1de1c4fcf400 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models A diagram is worth a dozen images
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757c8926-79bc-4b83-8eef-656cf9bd9112 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models BERT: A Review of Applications in Natural Language Processing and Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f45a04d0-49c7-4083-9a7b-458e621671fc · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecc99e65-9cfe-4be9-8062-c03588f67105 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4294a87b-37ca-4dd7-8684-3793d217ceef · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8e8a796-7b0c-4d36-bc25-67623a38812a · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4544fb5c-1eb9-4760-9932-a995bd17dc91 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d39cb784-ac20-49f6-858e-3926719fd61f · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Monkey: Image resolution and text label are important things for large multi-modal models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf36f4bd-1fae-47c5-85ba-3a755db47a5a · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316f67e6-7fba-4b5e-84ee-da9e18e278e2 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Improved baselines with visual instruction tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c5f576-1235-4f76-ae9b-4cc6d7c4d51b · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaV A-NeXT: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50640e39-879f-48dd-9f2f-013a9676da49 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef186e24-61e5-41cf-bf81-fda06bac4f5c · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f8843cc-3055-4ec6-9c48-159ad1fb97a6 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39b5d73-5718-41e6-9886-f2394b594b86 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a607c0fa-11a5-46ca-8516-3e96c1384d3a · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Decoupled Weight Decay Regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb20c4f-02ed-43fd-a417-9700eb54ee93 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56555408-5b6c-498d-a63f-ba3f12f10b05 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888b9194-5eee-423a-a6c0-3abbdc2e8561 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3622c02b-b745-4a0a-84e8-765ad3c35cd1 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Infographicvqa
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2035eb1-a169-47b4-a1c2-e9e7a65c513c · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Docvqa: A dataset for vqa on document images
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a20a55d-df00-4e28-a867-2b132bc3d205 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b0c52fe-66d5-49d3-8796-b11c09c98a05 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Ovo-bench: How far is your video-llms from real-world online video understanding? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 18902–18913, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd369666-9d8e-44d8-a40e-dd555445bd52 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Learning transferable visual models from natural language supervision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396cba8e-2143-49c2-a0a5-5cb1159e68b6 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models When do we not need larger vision models? In European Conference on Computer Vision, pages 444–462
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ebac53f-50e3-4040-a501-c1ec785f2b0b · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Towards vqa models that can read
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f22c7c6-e324-451c-bd75-660a10d00b58 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Roformer: Enhanced transformer with rotary position embedding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a6e65a-bf4a-4826-8a9c-fafbb8a22589 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d899ee3-e8d1-4120-8aae-aebb65305b14 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Kimi-VL Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e9f43c-e693-4959-852e-547fcdf21f5a · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b0045e-7682-452e-b341-2d0e1956b838 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc87ade-6243-4ef1-8afa-2e443bdfbefc · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f6e633-21d2-455d-a48f-02b01855bc70 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b412e13-b0c9-482f-9f09-a2dae75db265 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80137b3-24e5-4535-8511-8cd6d336417a · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8db740-51b7-4499-9a91-3e129af10f55 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Qwen2.5 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2db8e0-340b-4ab7-b439-f2c2c86854f8 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94d73980-c679-4485-84c2-0af603b7dc22 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd675666-0726-40a7-ada7-dcd14b2e0ccb · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727ce860-a44f-441f-8fed-e0382e7af1ba · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Sigmoid loss for language image pre-training, 2023
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cef6c13-6340-4263-8359-99ce2afbf1c6 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 309da78b-8d1e-4153-836e-aa64ee4ed65e · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d866024-da5b-4073-b695-8bf8b0b3a9ae · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80f41d3b-5c2a-4624-901b-1e90b1c58f28 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd9d7f2-d8c0-4b09-95a8-8012cb60fa00 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models Swift: a scalable lightweight infrastructure for fine-tuning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7e63517-4296-4df2-99a0-37102c2a6997 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce9c156-8fc7-47df-8f60-e76731845313 · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a882fd8-dd3a-4dfc-ba9a-baa3e1f05b5f · outbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models $” or measurement units such as “cm
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7ab0eda-b5a6-4a65-9a70-0eda9239920e · inbound
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 112900c6-95c3-433a-b8c0-8835c23ea18c · inbound
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fe9215b-8e3c-4eae-9b83-3f1e8e898893 · inbound
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.