Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2403.11703.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:13:03.169164Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T02:36:26.526436Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 7dc71625-7470-4bd5-9191-21555ca9f18b · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4160050a-63f9-4da2-a747-14b2a3401ff5 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 158
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7eb48ea3-aab4-46d7-89a3-843fd6901bcf · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a176405-3573-42a4-9762-43da5b515c7e · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bff25321-8cc8-4f57-ac0e-fab80c0b6773 · inbound
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4c9380e0-8f64-479a-9a71-12683b17e959 · inbound
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aee29cb-5595-4836-be2c-d468dbc7c14c · inbound
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ccce93-dd71-4f15-af6e-bde51b409499 · inbound
LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf3e0c0-7b68-4965-945b-5221e8622735 · inbound
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243f9db7-0ee0-452f-875f-e68730c06957 · inbound
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fd5e40-6ffa-47b7-a945-2e14b621c775 · inbound
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef32f889-edb0-484f-86f7-91a1467a4adb · inbound
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b7389b-5513-4a1d-9481-c34d3bcb9e98 · inbound
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5647cc14-a07e-43cd-b4f1-8ea8d72e7d65 · inbound
PerLA: Perceptive 3D Language Assistant LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9beaaad-9d20-4101-9fe7-7fbb08b79fcd · inbound
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a415b3-819d-41fb-977b-ce08a5b70f4b · inbound
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0472031-8831-4bed-a82c-bec1764fb727 · inbound
Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb3d671-2012-410a-884e-ad8830cd6e8c · inbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 120
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd986c29-9567-499f-ab33-9bc948658b71 · inbound
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca19e749-1942-4d8e-87fb-4d3f901b9cce · inbound
MBQ: Modality-Balanced Quantization for Large Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfcd838-d87c-49f2-bbd7-cf1f88539ebb · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97733769-5075-41b0-9ffd-ae8eb34423ac · inbound
Mitigating Hallucinations on Object Attributes using Multiview Images and Negative Instructions LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641d1f68-7aca-4ac3-b248-8dcc02586bb5 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 216
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73ca218-7d4b-46d1-83c3-fb6010d0fa22 · inbound
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b74ac38-e830-468f-a941-ae1a69344cba · inbound
Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4207f9a5-e042-4533-96ab-33b5cfcdf685 · inbound
RESAnything: Attribute Prompting for Arbitrary Referring Segmentation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7fd5c2-e049-4bdb-8a1d-5184c95a36b1 · inbound
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a948cb9a-ed4a-49d7-a9c0-9c532f81ba55 · inbound
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30a2810-ca8f-4506-841d-cab57a499d09 · inbound
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f360b7-f5ba-4a69-b6cf-8a1077010caa · inbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a51f921-683d-4c95-aaa1-d22e69acf980 · inbound
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c856420f-0375-4530-bb92-9893575ac663 · inbound
ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2f5763-d0fa-42bd-93e6-b8142f783d72 · inbound
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee24e842-e62e-47ed-ac14-e8fe6817bf62 · inbound
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1877bc03-d0cb-490a-927b-beb65ecab6fa · inbound
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653deced-ed06-4e0f-b05e-96b9ec49eb7f · inbound
Less Detail, Better Answers: Degradation-Driven Prompting for VQA LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7d1d40ee-1341-44cb-abad-283d687a1dc1 · inbound
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9de7d28e-7dd8-4052-8e29-2b5b9566c68d · inbound
Toward Native Multimodal Modeling: A Roadmap LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 08aafe95-02e6-4d43-93a5-4bcb90d34c15 · inbound
Self-Prophetic Decoding to Unlock Visual Search in LVLMs LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e7bdc8b3-0d3d-42e5-bd3d-270ddb991127 · inbound
When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d0d2ec27-054d-4647-b706-b20d7448e68f · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 228
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c955fe02-0de6-4385-83ac-b075f505bd63 · inbound
MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba25bc3d-d543-41ea-ba6b-fe2f83f6904b · inbound
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.