Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2306.17107.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:20:55.581390Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:39:56.491911Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d4bfc8e1-85d5-47e8-bb2f-fba9bc7fdeb5 · inbound
Otter: A Multi-Modal Model with In-Context Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9f0d2089-6468-4dae-8b88-6b1cde9056d4 · inbound
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f81cb006-3de7-4356-a7d2-478facbfd48b · inbound
Improved Baselines with Visual Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3d5da0f1-f352-4fc5-ad69-54e92950cc89 · inbound
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 752f4153-e37b-4e86-9fd2-0590c19295c6 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 183
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ee7bdb2-737a-4bac-bedd-c87c5c5dc7ec · inbound
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation da761b21-bc1c-4309-9a6b-0202105cea7a · inbound
Yi: Open Foundation Models by 01.AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40490174-7897-4290-ba07-fe31d5bcba88 · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7961a538-9c52-4f4e-b58f-0293a8b8c28b · inbound
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a0c9ca41-5014-4f37-9a4e-26ea794e962c · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 43929916-c215-4956-965e-51a3ddd062bb · inbound
LLaVA-OneVision: Easy Visual Task Transfer LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ae786b8b-a54e-4e8d-8a98-55d03431759f · inbound
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2db8a3-2551-487a-8100-5a7ca7283c12 · inbound
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 145
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edc6aaf-2860-41f6-8347-58a8f0931ca9 · inbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558f1a50-6dac-4dea-b57e-712b593cedeb · inbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d29eb95-245f-4fdb-a75d-d9ca6d406c35 · inbound
NVILA: Efficient Frontier Visual Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 538b0ddb-5dbf-4130-ab77-8f763dc35550 · inbound
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c205b1d0-1f6a-4798-bf3e-4ad34ccec911 · inbound
Chimera: Improving Generalist Model with Domain-Specific Experts LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a0b67e-8927-45fa-b62a-1bb7f7b13556 · inbound
FILA: Fine-Grained Vision Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889bc1e0-3991-4abf-a2ac-f4ec986508a0 · inbound
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2222b20d-9139-44d3-a3c2-0ebc94a796b4 · inbound
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ade628d-a4e4-4dfa-9dd1-ceeb3e3f2dc9 · inbound
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d98d28d9-125d-4090-b5e4-ce311eb6566a · inbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e70a1b-40d9-4db0-b8c3-58ef0c8b9815 · inbound
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e0cb043-a4ef-4d0e-8ba1-ec8070697b76 · inbound
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a793304d-b92d-4fa4-8199-35015e9dc1ae · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e374a64d-de7a-4749-b3ee-8397cfe4c77e · inbound
Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126d7e5c-0329-41f5-a982-ad7f632ef29b · inbound
Visual Large Language Models for Generalized and Specialized Applications LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 649883b4-5106-42e3-96fd-ab2226a01f29 · inbound
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861f6bd4-f1c9-45ee-a707-681fb84c0d0a · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1dfbae8-9876-4e0c-a730-6dd11911cf5b · inbound
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529a65a7-b01c-4a52-a470-69a8269c1b7d · inbound
`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f238bba-9232-4aa9-a2d4-a321631dae86 · inbound
Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 141
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83eed00-f88e-40ce-8d5a-061b7c511f41 · inbound
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation baae48cf-7a19-480a-85a7-c790be202208 · inbound
DocVXQA: Context-Aware Visual Explanations for Document Question Answering LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bc010f2-3fc0-43a9-9d3a-223b99a0245b · inbound
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f09d17-eaf0-4993-bfdf-572bab3e51df · inbound
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e354e6-b377-4bf1-90e4-2954bd1181c7 · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89728534-4ffa-413d-875c-af011ef21c29 · inbound
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a0e74d-ac7d-4915-8c45-3e16b5a9adaa · inbound
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 109
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a9aadf-9721-470a-a358-3ea30994739c · inbound
CoMemo: LVLMs Need Image Context with Image Memory LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccec615d-d13e-4de7-ba4d-53c35f5b8de5 · inbound
Synthetic Visual Genome LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24f0e5f-e49b-4afd-83d0-bc32bddfb148 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96f1011-b795-4124-9646-29a6fe338376 · inbound
Multimodal Mathematical Reasoning with Diverse Solving Perspective LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568769e4-9114-4bd8-9e2e-fe8d9b7c8914 · inbound
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b253b0-067d-4f80-b54f-90b17167e97a · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a888602-e291-42c8-9625-1461a92fcc8a · inbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8f319d-4258-4591-95f7-2899fd3f623a · inbound
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e5af482a-0791-40c7-92f8-9b122181d297 · inbound
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7999db-19ce-4621-af6c-b8935cf10a27 · inbound
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbfdfc6-adfc-4968-8047-9ba9f6bab163 · inbound
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 000e01be-5b7f-4368-870a-4cb39fd6e76e · inbound
Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d7ab97c0-7887-4c80-bb99-5e36e3e042cc · inbound
Closed-Form Spectral Regularization for Multi-Task Model Merging LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b85f4393-3752-4c7b-afd8-3bd4d3ec02eb · inbound
Vision Language Model Helps Private Information De-Identification in Vision Data LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c5bd9300-5a26-42df-9cc6-136b81cfbea2 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 228
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 422eb2e9-c7f5-49af-a622-94650ef89f12 · inbound
GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 929cea13-7215-4400-b797-b5db081d1334 · inbound
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 135
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a7a505-b829-4e90-9f55-e476e8698b32 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 282
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af204b94-4cd4-4e3f-91de-fc9cdbedc8e3 · inbound
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cab4f00-62f7-40d1-9b1f-c814d20e9488 · inbound
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0e55c6-3e45-4a69-9009-aa821fee3ac0 · inbound
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.