Pith. sign in

Paper Citation Record · LEDGER

LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2306.17107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.17107 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 62 of 62 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:20:55.581390Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:39:56.491911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d4bfc8e1-85d5-47e8-bb2f-fba9bc7fdeb5 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.915256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:c40502dab2749fa12e91cab24684a407c98bd6b58c931412cf329da2792b8065

Observation 9f0d2089-6468-4dae-8b88-6b1cde9056d4 · inbound

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets cites this paper.

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:14:10.243558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-24T08:13:18.191233Z digest=sha256:7bf4d798b87fdec69662052556925af24e508638fc09cc07bc412e5519bcb74f

Observation f81cb006-3de7-4356-a7d2-478facbfd48b · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.983508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:025f80e0bdedaa5acc4f3217ed2d56f14297a003023c755db0affd293730146d

Observation 3d5da0f1-f352-4fc5-ad69-54e92950cc89 · inbound

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models cites this paper.

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:22:04.209988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T01:22:04.035994Z digest=sha256:067e2345c7242e961ff7902a6c29573982ceffb41a5aa860d182fb9f29074feb

Observation 752f4153-e37b-4e86-9fd2-0590c19295c6 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.257255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:72d9897c1271d11be7d5d03677683c6b67d6afc620b29a220bb1f7fc913d8e0d

Observation 0ee7bdb2-737a-4bac-bedd-c87c5c5dc7ec · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.227169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:7b653f930bc1d3baf3c31c1a2cf99a6c9de3ed3c548a19787f0f253a5eb46aef

Observation da761b21-bc1c-4309-9a6b-0202105cea7a · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:47:27.946357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:7f91acd984b05937ccc966d028cf2646f79e0f599f1f94f73889aa52d803a35a

Observation 40490174-7897-4290-ba07-fe31d5bcba88 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.149695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:a51e30d85d7492a7eb298fd8272b620d6cb6a5bdf9985edbb90bd2dcd488ce9d

Observation 7961a538-9c52-4f4e-b58f-0293a8b8c28b · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.865968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:b4995f4330ac7649240103cccb00e44980692d35b071d28fa004ab7826e8858a

Observation a0c9ca41-5014-4f37-9a4e-26ea794e962c · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.151963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:c14c8821837b291b8afde956b52d2d39572b98b46e90ab5016353d1be0aa22a5

Observation 43929916-c215-4956-965e-51a3ddd062bb · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.586902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:ab1dbd7290d2aede4047175bb067ed04fb2886af68cf615ce1652fcc9c2175f7

Observation ae786b8b-a54e-4e8d-8a98-55d03431759f · inbound

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval cites this paper.

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:48:54.072395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:48:54.072395Z digest=sha256:bf5f974e28bd1c0bc5b3fe8ab6c85c81f8378d7872b07eb72db749a079f310e2

Observation 2f2db8a3-2551-487a-8100-5a7ca7283c12 · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:01.084953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:01.084953Z digest=sha256:028225f691f0260a228cfa29baffe05e3b91a2b2e74b15a1f53668a01a319598

Observation 7edc6aaf-2860-41f6-8347-58a8f0931ca9 · inbound

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness cites this paper.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.887857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.887857Z digest=sha256:b502ea87ef436759a54f201e873f31a8d519aca92fd1060ce63dbc2b20289eb0

Observation 558f1a50-6dac-4dea-b57e-712b593cedeb · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.465577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.465577Z digest=sha256:5340a73aab4b4463e59367506877cc23d398372af7ca7729e89369fe3b5af3e5

Observation 7d29eb95-245f-4fdb-a75d-d9ca6d406c35 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.046117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4833eb10738804e1cf6cb2898540431d7b29640df3a66557fed916957a0e10a2

Observation 538b0ddb-5dbf-4130-ab77-8f763dc35550 · inbound

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation cites this paper.

EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:15:45.477475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:15:45.477475Z digest=sha256:37716121873151003e13bd22640b87d8ff0e09446851d10cbda8878c1ccf0209

Observation c205b1d0-1f6a-4798-bf3e-4ad34ccec911 · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.739840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.739840Z digest=sha256:8efffceae0eb3378bda25df8b3e3584d740fdcf7f78fc1d2016994682979d762

Observation 13a0b67e-8927-45fa-b62a-1bb7f7b13556 · inbound

FILA: Fine-Grained Vision Language Models cites this paper.

FILA: Fine-Grained Vision Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:56:46.258654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:56:46.258654Z digest=sha256:df3164bc6a7382d1e4b771c91c9f5052f2bfad6a694aa0a4c53ecbe4a41c87af

Observation 889bc1e0-3991-4abf-a2ac-f4ec986508a0 · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.763872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.763872Z digest=sha256:a101c93a5457f7547491561958205088b2497b4d132eb6f3ab3260d4a7ef6b21

Observation 2222b20d-9139-44d3-a3c2-0ebc94a796b4 · inbound

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants cites this paper.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:58.137134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:58.137134Z digest=sha256:7a78fa613d99d03a47bca9e2ffc7f22b23eb7d75f70590a1fb8ea67682db97f7

Observation 0ade628d-a4e4-4dfa-9dd1-ceeb3e3f2dc9 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.090250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:27b2c605ce252e20b551ad7fe39fc0e95b3581d9e3ca2745fe7ed1da01df5c26

Observation d98d28d9-125d-4090-b5e4-ce311eb6566a · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.410853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.410853Z digest=sha256:9e625ba406d9c15f1fd4a6c38350e74d0831abdec0ba7edb03a1a5248a5f5b3a

Observation 76e70a1b-40d9-4db0-b8c3-58ef0c8b9815 · inbound

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation cites this paper.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.351843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.351843Z digest=sha256:dbd9a733ae8d0e53c4fce8f57e3c6b04a5291b5f6f3353454beff08a0935b080

Observation 1e0cb043-a4ef-4d0e-8ba1-ec8070697b76 · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.522706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.522706Z digest=sha256:6efc239d9adb6099b0ba951894b4f6d77a0360a2a91b572237ec8b9f0ea9d5bd

Observation a793304d-b92d-4fa4-8199-35015e9dc1ae · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.773900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:10740c491d0396d4b0958b25f4957d0b84fc23b1b5fd7ab3be8cca012b0f900b

Observation e374a64d-de7a-4749-b3ee-8397cfe4c77e · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.138699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.138699Z digest=sha256:298a66276a800d588dfebcaef296a2db9931c3eefb0fcb53d91864cd35c0fd5e

Observation 126d7e5c-0329-41f5-a982-ad7f632ef29b · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.281166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.281166Z digest=sha256:e4d03e680c97d2ade96e8894c1265438629582e0d6c8645b18c586a8ccecae10

Observation 649883b4-5106-42e3-96fd-ab2226a01f29 · inbound

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts cites this paper.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.653313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.653313Z digest=sha256:92320ba7894fb122a02576054126bfafd553f3cccbacb679ef0d1d62cadae946

Observation 861f6bd4-f1c9-45ee-a707-681fb84c0d0a · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.522544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.522544Z digest=sha256:0e97782c30e27e13fad0e44ffe02264cb21c7b80ac460ad03559a2a1b53d8c4d

Observation f1dfbae8-9876-4e0c-a730-6dd11911cf5b · inbound

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink cites this paper.

Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T14:29:28.328190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:29:28.328190Z digest=sha256:52bcf08bf3a25ec9e96566b867bca40cfc08fd42f9bc5491505fbc9f5d3254cb

Observation 529a65a7-b01c-4a52-a470-69a8269c1b7d · inbound

`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs cites this paper.

`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:58:57.493585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:58:57.493585Z digest=sha256:13109097da757540c0fbde74c86427d2f48b87d0dbdaad6d57be500234f9631b

Observation 9f238bba-9232-4aa9-a2d4-a321631dae86 · inbound

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective cites this paper.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:40.287500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:40.287500Z digest=sha256:8ad6db3fe2b98fef47dd6c11ed45a25bb0cd4ee30d65d18e852e39832b3f29df

Observation e83eed00-f88e-40ce-8d5a-061b7c511f41 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.814828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:4159c6769eb9b267b7555575969ffe24800d896b1fa3f77b1d3d860beea8248a

Observation baae48cf-7a19-480a-85a7-c790be202208 · inbound

DocVXQA: Context-Aware Visual Explanations for Document Question Answering cites this paper.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.581390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.581390Z digest=sha256:e5d770a95bce0b3e2d1978199483af11ac763a27d7040bb0b51bcb4380a9f719

Observation 7bc010f2-3fc0-43a9-9d3a-223b99a0245b · inbound

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? cites this paper.

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:02:42.032537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:02:42.032537Z digest=sha256:c4eabc4a31b939551d4e865e18d147ece242c48e4b1cb0136d69d51bb3869c78

Observation b4f09d17-eaf0-4993-bfdf-572bab3e51df · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:10:56.107060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:5aef8ea2e9eb4f1a286f2c60fbffc831c9a70cd8f1a16118e0cd1fbd7a6280ad

Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.258922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.258922Z digest=sha256:1e530cd5fd1d9baf92be746ec10ae27a53993064abc52175a049f085c0eafc0f

Observation b6e354e6-b377-4bf1-90e4-2954bd1181c7 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:57.989894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:57.989894Z digest=sha256:82babc627bc0aafcdf851184ba949359bfa232d60610c3d3064e949779e7efbd

Observation 89728534-4ffa-413d-875c-af011ef21c29 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.067992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.067992Z digest=sha256:7359a76bc5ee436d938b37f3ad2592ae1aff63682684f7e2fdd15249c3d4cfb3

Observation e5a0e74d-ac7d-4915-8c45-3e16b5a9adaa · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:27.922748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:27.922748Z digest=sha256:c79379396c8956930c41c5a42d7353b40a49f9762e40fc6bb67b4e3f7341199e

Observation b6a9aadf-9721-470a-a358-3ea30994739c · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.509268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.509268Z digest=sha256:8c85475c0679158cf238bc0e55db0a17ad278fa52ab84765ebac88fc96b87df0

Observation ccec615d-d13e-4de7-ba4d-53c35f5b8de5 · inbound

Synthetic Visual Genome cites this paper.

Synthetic Visual Genome LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.492581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.492581Z digest=sha256:66181f82fba92657ec18e91c6c85af7d2b3f306f52d3a735efdea31ab1d5ae7b

Observation b24f0e5f-e49b-4afd-83d0-bc32bddfb148 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:28.750521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:28.750521Z digest=sha256:48fea90e96f3f613e4464629f0c65cf3bbbd10bd6e61cf271d351640d1fbd3e5

Observation a96f1011-b795-4124-9646-29a6fe338376 · inbound

Multimodal Mathematical Reasoning with Diverse Solving Perspective cites this paper.

Multimodal Mathematical Reasoning with Diverse Solving Perspective LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:23.916976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:23.916976Z digest=sha256:9dcb7edb66fc6a1dc4c693178dc6dad2d36172592d695ecaa0dd721c5c3c356b

Observation 568769e4-9114-4bd8-9e2e-fe8d9b7c8914 · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:21.283146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:21.283146Z digest=sha256:13aa542e9bb2f4b41120532516d6a55fc3c44a999a9dd9e990986ff69aac9d5d

Observation 40b253b0-067d-4f80-b54f-90b17167e97a · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.680277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.680277Z digest=sha256:11be2e78518be112aaee481a44bf8fae8807c6488991c606e89689c10657a30c

Observation 1a888602-e291-42c8-9625-1461a92fcc8a · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.601054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.601054Z digest=sha256:3dc904077ed4c3187d2ef8cc00302ff00caf1ad45f1aa3a3c21600e9ccc00276

Observation 7b8f319d-4258-4591-95f7-2899fd3f623a · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.346263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:b21fb87541d0b2cc30204e52a59105479c4790f8f8f796586c40a4e54b321f9d

Observation e5af482a-0791-40c7-92f8-9b122181d297 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.448064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.448064Z digest=sha256:d8cba63d1b023fbb4e26f2fd56d0dfb71f7c94310ed4dc9e9a57a0e63e6effb2

Observation 9f7999db-19ce-4621-af6c-b8935cf10a27 · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:08.096024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:08.096024Z digest=sha256:f5894f2432ecc7f2463cf86be11989f9b273048ee629d65b6818b90affe9d1e8

Observation 2bbfdfc6-adfc-4968-8047-9ba9f6bab163 · inbound

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models cites this paper.

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:56:06.970215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T20:50:57.818064Z digest=sha256:f5d491bb66fe8269c243e42f0add8b52e09ce7cbc5eedea314cbee614c023ed1

Observation 000e01be-5b7f-4368-870a-4cb39fd6e76e · inbound

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models cites this paper.

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:30.856645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-07T04:07:58.031100Z digest=sha256:2a0691de79dbcda561c2b801fe2e51bd64515a6b21ece455e74e7ab891eb8c5f

Observation d7ab97c0-7887-4c80-bb99-5e36e3e042cc · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.385400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:6c889971997e33fb963a82761413387fced20647b3d892657df8f9fca2af04d6

Observation b85f4393-3752-4c7b-afd8-3bd4d3ec02eb · inbound

Vision Language Model Helps Private Information De-Identification in Vision Data cites this paper.

Vision Language Model Helps Private Information De-Identification in Vision Data LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.578905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T16:29:33.294960Z digest=sha256:f0e8e7901e73e5f0f26e5dc373cdf0ddab354eb487e6a53dc4a64e31a9cd6b7e

Observation c5bd9300-5a26-42df-9cc6-136b81cfbea2 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.219425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:2bd12e8ed3f906ccb95f16b28c409475b12a06d53c06f670540129e4c9ead457

Observation 422eb2e9-c7f5-49af-a622-94650ef89f12 · inbound

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models cites this paper.

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.493243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T01:32:40.435742Z digest=sha256:78facd1effb5fa3b7c1cd274c1a22655848b01c9b609f3c8acfe14c010295406

Observation 929cea13-7215-4400-b797-b5db081d1334 · inbound

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI cites this paper.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 135

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:2b714e46d835dc0b16f5aefaa63f1864d78afa5aa81ec26df0e593fa37f2d365

Observation a7a7a505-b829-4e90-9f55-e476e8698b32 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:11.681602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:11.681602Z digest=sha256:6c4ac8ed6f5eec37566d489ffb42f85471ee2c2cda147009ac2f36ca08703c20

Observation af204b94-4cd4-4e3f-91de-fc9cdbedc8e3 · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.244130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.244130Z digest=sha256:907d222ffc2289fa9ba2f68d2038de716b2a16ff003cd06245a89a4677c7e6e4

Observation 7cab4f00-62f7-40d1-9b1f-c814d20e9488 · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:26.039380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:26.039380Z digest=sha256:21aaf42db24784b54d03f58c39cb4a306b0c070c6adfd5610344d6914c87ac2c

Observation ab0e55c6-3e45-4a69-9009-aa821fee3ac0 · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:25.742591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:25.742591Z digest=sha256:241886bdd85d21cc7a7eba805b3de6a39b18c27b5104544d1d63a8eed89cd187