Pith. sign in

Paper Citation Record · LEDGER

LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.17107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.17107 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:13:26.039380Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:39:56.491911Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d4bfc8e1-85d5-47e8-bb2f-fba9bc7fdeb5 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.915256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:28bd195e6d27aa3b43861e8d1463c8bf8a33587ee792eebfe485eb563f3c594c

Observation 9f0d2089-6468-4dae-8b88-6b1cde9056d4 · inbound

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets cites this paper.

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:14:10.243558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T08:13:18.191233Z digest=sha256:27f398ca080bb8a9cc0c9e5cfd2d9e72484ce4505e128a83b8ebeb76c789b2eb

Observation f81cb006-3de7-4356-a7d2-478facbfd48b · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.983508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:52d2cbe6bbdd907aaa19e5b42ad1b31344d398b4088e523d82c87e952a20f22f

Observation 3d5da0f1-f352-4fc5-ad69-54e92950cc89 · inbound

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models cites this paper.

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:22:04.209988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:22:04.035994Z digest=sha256:ed5df20048714f826d39220db3f7b2487c6d7e2290c0a1dae7d9b402b8167ccc

Observation 752f4153-e37b-4e86-9fd2-0590c19295c6 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.257255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:893b7ef4ccdad82b1eb590ea20d86271e68def4aa34a1cf44fd527a750c625ca

Observation 0ee7bdb2-737a-4bac-bedd-c87c5c5dc7ec · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.227169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:4afed6d6c7a5e57a06d8d9acfb86d5062871f66ac18c1e2dd10c4c56b20eaac1

Observation da761b21-bc1c-4309-9a6b-0202105cea7a · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:47:27.946357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:579f40c573afa9608b4eb5f242ed7b7c78ea737504d554f65ab12d421a12ae94

Observation 40490174-7897-4290-ba07-fe31d5bcba88 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.149695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:a83224d4cdd55b7ad0c6928c725c22ebbc166a9e7782aa0f2456c51f8f97c7b9

Observation 7961a538-9c52-4f4e-b58f-0293a8b8c28b · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.865968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:d72793aeffdd975cd0837b2e0984a374dfca34142b7c55b1213f290f3dd41dd4

Observation a0c9ca41-5014-4f37-9a4e-26ea794e962c · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.151963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:7a393c7a51f0728d02834b5b1b0bd90e43fdbfe14ac1749b6819b3d7012e6630

Observation 43929916-c215-4956-965e-51a3ddd062bb · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.586902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:2db12a7bb88da8a04f4df6b61ed74c6957758b4d0a66ddc98bbaaf2504b9ad5e

Observation 7d29eb95-245f-4fdb-a75d-d9ca6d406c35 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.046117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:cbdcc9383b4d6f58ea44de8d0c8224c94777e05a6ed8ba89d8fd952dfaf5ab44

Observation 0ade628d-a4e4-4dfa-9dd1-ceeb3e3f2dc9 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.090250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:e9460e278c98452edbcb8ecced4a16ffc8ca7ac66bfbf48daf464cb6d93e3d20

Observation a793304d-b92d-4fa4-8199-35015e9dc1ae · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.773900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:b6e8ace97b2a7838713d29f1aa7029c82f9ef63be01757d43ca82a3d6234d042

Observation e83eed00-f88e-40ce-8d5a-061b7c511f41 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.814828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:2911f54b5e0cd7f08dfb648ef51efda6ebf4332bdbdeb5aa37f7752c503d4bbc

Observation b4f09d17-eaf0-4993-bfdf-572bab3e51df · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:10:56.107060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:0cd74cfd3b15220d07f36e16d24cd1ba9d72fa7f54720de8ae4d68b3ea011a51

Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.258922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.258922Z digest=sha256:fc27693fcffb1d80470cec685261e2c85d9132244326f526a7036808f690528c

Observation b6e354e6-b377-4bf1-90e4-2954bd1181c7 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:57.989894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:57.989894Z digest=sha256:aecf7db498d2235e4a527bd65e3fc44b3802c58fe28b0502c71746dcdb62fc49

Observation 89728534-4ffa-413d-875c-af011ef21c29 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.067992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.067992Z digest=sha256:f8fe46deccb1f792121c24642edd0e4ecac59cf74c71b71689bbe17e5edb4baf

Observation e5a0e74d-ac7d-4915-8c45-3e16b5a9adaa · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:27.922748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:27.922748Z digest=sha256:0db34efcc2d7f8b68ad29fb2d90496c6c95291838c8fbf89f3425a210882f3fe

Observation b6a9aadf-9721-470a-a358-3ea30994739c · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.509268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.509268Z digest=sha256:370a97f9f2fa7bd31a2aa8268a90bfa9bafef928b68c955554b886da933a7c43

Observation ccec615d-d13e-4de7-ba4d-53c35f5b8de5 · inbound

Synthetic Visual Genome cites this paper.

Synthetic Visual Genome LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:56.492581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:56.492581Z digest=sha256:597d8e73fbd72edced34303df4b4a260c6d2288e643e6324e3373e9edd5b1442

Observation b24f0e5f-e49b-4afd-83d0-bc32bddfb148 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:28.750521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:28.750521Z digest=sha256:5e0d343c18a48add0176f7147dd07c0e2f9e8436aaec236d69cd9646ec5ffa4b

Observation a96f1011-b795-4124-9646-29a6fe338376 · inbound

Multimodal Mathematical Reasoning with Diverse Solving Perspective cites this paper.

Multimodal Mathematical Reasoning with Diverse Solving Perspective LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:23.916976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:23.916976Z digest=sha256:e9be5268700b7334455ded2e450ae43887415394db045f1fa5621e15fc4c814c

Observation 568769e4-9114-4bd8-9e2e-fe8d9b7c8914 · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:21.283146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:21.283146Z digest=sha256:dc4cdc42e6c4145ca219c60461f8383800e433d8e7c2177c53b5027eebfd9027

Observation 40b253b0-067d-4f80-b54f-90b17167e97a · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.680277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.680277Z digest=sha256:c3f5052676b06ff8a2afa3ae683f5eb9d200cc743a8d09df9886cb5b9321ed3f

Observation 1a888602-e291-42c8-9625-1461a92fcc8a · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.601054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.601054Z digest=sha256:8e5f67f70c2cf4449f71c2e4fb6d35814f9bd488ea156f5543f17676f842c13d

Observation 7b8f319d-4258-4591-95f7-2899fd3f623a · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.346263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:fc5016ffe00c515bd10aa312ad71b92185f1fd73c6182a78c556b1b4ead02737

Observation e5af482a-0791-40c7-92f8-9b122181d297 · inbound

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos cites this paper.

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:17.448064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:17.448064Z digest=sha256:9abb1e8e14579e35bb076ebbc9a8a9f938f167172ed82ec8418db33e85816764

Observation 9f7999db-19ce-4621-af6c-b8935cf10a27 · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:08.096024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:08.096024Z digest=sha256:cf457da1782adaf736e02e9b9ea4da54aa9d53c3cf6bc64b7430e8ae98221563

Observation 2bbfdfc6-adfc-4968-8047-9ba9f6bab163 · inbound

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models cites this paper.

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:56:06.970215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T20:50:57.818064Z digest=sha256:13fc9ca24aa0d60f21e2661534e5dc51fdd3301fc2ce2b1a2f2e373aa455d92e

Observation 000e01be-5b7f-4368-870a-4cb39fd6e76e · inbound

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models cites this paper.

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:30.856645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T04:07:58.031100Z digest=sha256:6b858a8bb0288c6715dc5669b0c8da78c6a1fc6cd37f63f29849aeb3396742d5

Observation d7ab97c0-7887-4c80-bb99-5e36e3e042cc · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.385400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:04662693a0be5768f9aa405c28a0ebe9dc75de68cc18c1dfc26bafc71b7f5f8b

Observation b85f4393-3752-4c7b-afd8-3bd4d3ec02eb · inbound

Vision Language Model Helps Private Information De-Identification in Vision Data cites this paper.

Vision Language Model Helps Private Information De-Identification in Vision Data LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:31.578905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:29:33.294960Z digest=sha256:717ca77c823fd62d55f766d7e989e2b1313b5839fcf748445975079e54f0f1a6

Observation c5bd9300-5a26-42df-9cc6-136b81cfbea2 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.219425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f672a1bea445107efb87e7bfbdcc32251976a8b64035212c7237aa829ee37e99

Observation 422eb2e9-c7f5-49af-a622-94650ef89f12 · inbound

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models cites this paper.

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.493243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:32:40.435742Z digest=sha256:34707ad39b93889c431f63a0c41cff4dc2798106d1b4a48e1b3267850a30df7a

Observation 929cea13-7215-4400-b797-b5db081d1334 · inbound

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI cites this paper.

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 135

Resolution
unresolved
no resolver link, observed 2026-07-14T04:45:32.682508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:45:32.682508Z digest=sha256:4191d3e3c4771ee3d54a90158ffa496ff68d7c7112132036895269f9aa99358c

Observation a7a7a505-b829-4e90-9f55-e476e8698b32 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:11.681602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:11.681602Z digest=sha256:8ffc352d399aab5c711c85215c5a03946d97fafe79ad71793d08d8a3d0095bc8

Observation af204b94-4cd4-4e3f-91de-fc9cdbedc8e3 · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.244130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.244130Z digest=sha256:937a038c373ec1cd00c79b0a9f3be1b58a8e09841edb32dc30ed99072a8ca8e9

Observation 7cab4f00-62f7-40d1-9b1f-c814d20e9488 · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:26.039380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:26.039380Z digest=sha256:ed9b5243ca543278e223799f0c27740f04b9866489f765d9aa62f943e747fea8