Pith. sign in

Paper Citation Record · LEDGER

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

As of 15 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 2 inbound Pith citation observations for arXiv:2412.16364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16364 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:43:08.380678Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:32:48.356870Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T15:32:48.405652Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5aacc674-890a-4e79-a0c6-81092622eb59 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.015968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.015968Z digest=sha256:8e234cb01eb4589717ec8c09ceb8a245a38861fe3f26a3f07430b1364037ed91

Observation aef6de57-ada9-4950-90b9-189813c046f7 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.644327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.022473Z digest=sha256:993d483e4b23cfb8fb3d5cce544ca1f1d130c652894213b7e5b9f680f2e3dd26

Observation 7a91493e-ad36-4acd-b3f5-756b4178e5fa · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.028411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.028411Z digest=sha256:972f864990c7cb75c80840b1749d8a0a5807379a9e608382fdbb8450d16721bd

Observation dbad5b1a-de18-4cd4-95c4-376891cc6205 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.034935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.034935Z digest=sha256:9285a20e7bb701e9d1ec7c4678507455b2b97360287ce712ecd133a95d0310cf

Observation c9aced62-d86c-499a-9f61-c0bbade12ff2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.040942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.040942Z digest=sha256:6aaa1e201027ebd8f5cd5f6fe51ef0d6dc6334100fc806ec7efdd647ee6f5902

Observation 456f2e71-686d-4cbc-9525-5d55b18abacb · outbound

This paper cites Qwen Technical Report.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.047462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.047462Z digest=sha256:2b1d5e1977986cba32567b309c6a278861af0ba1adbe59c4102c1d7126c14bad

Observation b1ccb274-a20f-4c6d-b211-3fd36e23abba · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.054234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.054234Z digest=sha256:57780903536713c295f7f13523051205e98633df22d6043efd1fd10da8c7fe7b

Observation 2f1ee924-13e8-41eb-855e-e2ea59aa1595 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.059750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.059750Z digest=sha256:33da6259b31ca25beae5ece4200d3300399fc3ce00677f67f753c4a1b0fb8195

Observation 08368eb0-7df9-4183-892e-b4b7cf4c808d · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.065314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.065314Z digest=sha256:7d519a870fbb97eb1b1a1c1a90fe27440516fc7eb224d5ff1bf0ab991a78f599

Observation f5819198-0513-4c32-8114-88143f42e9d3 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.071267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.071267Z digest=sha256:dc722facd096d020f88cbe4a51beae81292b79b0f5bb3c571421745daa6f11dc

Observation 31d5aa54-06cb-499c-a7a8-d6e5c7dbd311 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.076731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.076731Z digest=sha256:89e5fecc58dc98416da8b83cc563c807e0f59df4cbdc672f8d02ed0d523bc022

Observation f68f39bf-0ae6-4b10-a708-822246bf7dcd · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.082289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.082289Z digest=sha256:0123b6166761ffdbf7f42a6f579d7f276ebd90065f6cf73386a4826ded01f654

Observation ad00a16b-2107-40b4-9d40-b3ff799b7734 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.088107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.088107Z digest=sha256:f7c779a8e85ffc56c440088d82af2b8f80952f768f086675990ba4873510a577

Observation b4953843-a2cf-4ada-84b5-63a33a282ab2 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.093445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.093445Z digest=sha256:eb0a4f4179e282241df1258f85737c284f391bb96caaf0936916188f7700875b

Observation 03b891d6-79f4-47fe-b99a-8f1b243851b2 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.099118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.099118Z digest=sha256:004c303476f9d9ce42df9233262b08bff8c59c7c0b175aade6161f784b4dd464

Observation 10e10636-9685-4993-b3c9-bb76777af97c · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.104693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.104693Z digest=sha256:f867a1f1e0472965ad47e4cadfc714652e503e45fcbc50b04924108824ac0c2a

Observation 6e901dd1-c024-4d6e-a312-6a7d435ad9e5 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.111217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.111217Z digest=sha256:a6c0fa787f1b851896a0468819be83ac0943d42632f39cb70d8f8987b3435610

Observation 1363a921-11cf-432e-9ef4-255feebfeea8 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.558346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.116972Z digest=sha256:d0a358b0987bffb07311472892fa58f918d6b77be66548a072f5f8e5c5aec078

Observation fd1c25d7-e7d1-495c-90a9-16cf49b50fc8 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.541404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.121871Z digest=sha256:8c54581d1b4c452f3d9e6b547a0822cea7088bb8cc3aa3e3e7c40de0cae6365b

Observation 747d26dd-1952-4d61-a751-8fa681cfff86 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.126593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.126593Z digest=sha256:0647e85fe601ab456b67cb4ff4bd26f378fb34380893e744369657c03c818a84

Observation 4bad5c29-2f67-490c-be1d-f7bb08f05aa9 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Building and better understanding vision-language models: insights and future directions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.131404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.131404Z digest=sha256:ae3bac46c88cfc40c1ca1f952468552aa67cd5b210c084b043aaea1c7d7715e2

Observation d522efe4-aeba-4d33-b122-5344e7deb0a1 · outbound

This paper cites What matters when building vision-language models?.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation What matters when building vision-language models?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.136042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.136042Z digest=sha256:5d4772ab95af75f6380e67890f9bee1621de48f27044afa15e50fea428d296f9

Observation 255da361-d0a9-4079-9334-10dfc9b4a91f · outbound

This paper cites Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.140756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.140756Z digest=sha256:e523820ddf410b05ee1abbfdf47d426bb69c7e5db0a7779e0ff103e6a732e72c

Observation be2a0fa4-5fcd-42d2-8c89-e5a8607e13d9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.145499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.145499Z digest=sha256:9d387dc3efab8b51141b1e705397180162db689d74437c8ce9021d4e71f173f4

Observation 9fc6c466-f7e4-4222-a81a-f9a6d9f3d38b · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.150494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.150494Z digest=sha256:e693ea3086f2c1adab729d17eec916710485f1d98ba52586f8255f6224782e53

Observation 28191912-273f-4eef-b0c1-c64b068cea95 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.155150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.155150Z digest=sha256:db00e1a8f0bc7c25a42e6cc113abc841822a29950728928db53e976c801d1701

Observation 67d16506-8c00-4edb-9dcd-180883516aa8 · outbound

This paper cites From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.159936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.159936Z digest=sha256:e6a8940815a53f746a0838ccabcb07f803c0a2fc3a6d4920e2e3cddf2db364ab

Observation 3fee9e36-0d0d-4994-9c58-de20c1335339 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.513539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.164632Z digest=sha256:9f34a959b25f85e8a58044c63dd8ebfd7b4e9d954f2ef9142cafe9af831b9c16

Observation 42e9e741-e613-40d1-912a-94582f8773d0 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.169290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.169290Z digest=sha256:dbfd2ef868100c66f65ab2881e5a0b2dbf9954ec2d4de8d3a765517c9e50d14a

Observation 8cc2fb93-52c7-4778-892d-c0421725b7e1 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.174228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.174228Z digest=sha256:fdc51c0248f17b7ffee11608cbec7aee81be58751a575e850daf0a9d36167c1e

Observation e17d3b1d-91f7-4c8b-9d5a-d2832ddd1a74 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.179262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.179262Z digest=sha256:753989210b5841fa6516673f14df495829efa4ba316107778017288393183a11

Observation 8b73e3aa-3a16-4350-bcc1-3e5a57987d8d · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.184299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.184299Z digest=sha256:1b2a6b9a2032833de85ca13b61a779747cf556a3260f1132d6853352a69aa5cc

Observation 79c41576-adb8-49a9-8249-12504a124eb8 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.189517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.189517Z digest=sha256:fddb387ffd44ae696a1eba15558fa6d6c40f4fe3b4317cdbfd2156aa6fdac0fe

Observation 823076d6-51a6-4cc2-a594-43e901f5f457 · outbound

This paper cites Visual Instruction Tuning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Visual Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.194693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.194693Z digest=sha256:882daaf08b3761b4ed016976ee4ea79d6fc1a21d67a5b298b3584a1bc6bf9c3f

Observation 9d3cded9-7343-4fba-93c2-17c02b5ab6ce · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.200561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.200561Z digest=sha256:7f13b2edfe78587fe690439c009311405f3d58ae69900f6bfb770cdf9319500e

Observation 68958a47-46d5-4079-9441-3d430f02dc99 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.206401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.206401Z digest=sha256:88363744649b035b80dbf0ab21bb9f30a784c818ac072a0e196b1c30f3874a33

Observation d7a9b546-68f6-4722-add5-1fba7b2724c1 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.212053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.212053Z digest=sha256:b96e2c11c2627b726bb519aef5d277e48b9b38e3f8c72a5aafcfd57684793469

Observation 637a3691-509c-46d7-9e79-adb2080d4e06 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.217303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.217303Z digest=sha256:c8cf674384e2bd4194ca068117ca81c9f211fec1faed6a07fd86b37e7176c907

Observation f08d9e99-b392-4428-9f5f-de5d3983499d · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.223014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.223014Z digest=sha256:b215b646cb12d2bb2f5102addbe1c3fb88978f24cb7f2075ca71d7c0e768a9e2

Observation ae2dc1fe-b717-4c00-97f1-1ca9741ab3a7 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.227921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.227921Z digest=sha256:d99bb7095aa7e47f2e44ce866c1e10ef93109e7a0a9a568588d1a975642fd3a2

Observation 5511b1bc-08cb-4d2a-a873-3ee6dc34abfe · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.233284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.233284Z digest=sha256:55a048b8e7bf6891d43b2068832593d36da177aa7987a046b7e8f0b894eef2e8

Observation 85c2ceb8-1290-4bf9-9aa0-94302cca30b6 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.238556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.238556Z digest=sha256:48599ddf23b370ae6786a8dba6a532c7e8fdd29556e829dc987e6da89f38a116

Observation 7c107ef1-2274-45ad-b018-1025e2bd7f32 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.243860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.243860Z digest=sha256:1862c1eb22c0114edaef3f4c8ae67da23d139b9881cfc8b0d58b776e729d3995

Observation f2ed88f0-b8b6-4a0c-bcda-7f20f7169886 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.248898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.248898Z digest=sha256:91562a2585bd58478699240579e25e03f3a5d3a12a54c07c621b82da280d5127

Observation de5324e3-3a32-44a0-a14c-207650ee37e9 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.254025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.254025Z digest=sha256:b778628ede816067d5dd76abfd449ef4fab53985dbbb2979d23ee5597a446e1a

Observation fa9e2839-d2f8-4f32-b1b4-61908a1ba1e7 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.259469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.259469Z digest=sha256:11a29bca2ad97d3ae599b7147016e9bd580d5552f775bcdc3f5da6aac46f43ad

Observation 83e3ba0f-5a8b-44cf-98c4-82a29bed14f3 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.264439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.264439Z digest=sha256:b8f8c8a0f772687fdcba2ab712c715c76286b7ac8a0998c0d62b48e467f94217

Observation 084ea4e5-4ca3-4848-b4e4-03a678783b6c · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.350050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.269511Z digest=sha256:f032a5e5034315d42977f215f6fd7e5bb7b5280bbf1f6f996e2c1f268829441d

Observation 6ee6884b-69f4-481a-9f67-263f29763ccb · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.274466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.274466Z digest=sha256:a1081a1568e6067ad64e82bcd5bd63eb4426998aa5a28b889abcf99723c37de2

Observation 660abd36-6da8-426d-a5df-ea81382b457c · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.279350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.279350Z digest=sha256:367c6e9c3a96e2531436e667777658e2e9be99ae733c4f34c875dec18cdbbedc

Observation db90e8f6-2ebb-4bde-8364-563eaeb3d8d3 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.284180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.284180Z digest=sha256:f1a1cf24867018cb1c0e91e1d6e52f0f9b700bfef45a64b5c9f46010c1fa292d

Observation 3504fa73-4fd3-4924-a1f6-f4c77c5adda4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.289122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.289122Z digest=sha256:7350f67cdc1da49c4a006a2f10155167a8f7a80448a51b3108ac7bd72d7fd036

Observation 37a2c492-2125-4336-b4b9-052f6a34d4bb · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.294005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.294005Z digest=sha256:0f0e4d0b5e41cb543d84057541f3e609e9e72e2d4aa32a2cf5bdf39049c4edff

Observation 819bbdeb-4eee-42cb-8852-9622a2f95d9b · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.307105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.298833Z digest=sha256:f2c5c558bc53508489b2a22dace5518488dcbe517c1201524e42a339c9561e2e

Observation 1bb51277-e1c3-4935-83b6-ec35b11cdbb0 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.304211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.304211Z digest=sha256:8338b8c40e39ab0d859fe0e82c2b921c0179e0cc1861e6fdc5bc05ad38d0fe16

Observation bd986c29-9567-499f-ab33-9bc948658b71 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.309295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.309295Z digest=sha256:df6e19bd39700ea8f49c2b999df60b92565b647966dd6e2a086f1f8f6bf693cd

Observation ef38d3f2-5405-4ae5-bee0-17e04f7452bb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.314525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.314525Z digest=sha256:27afad88e369e88aa53938134e690976c427751284bc5827a01158dfb8f4ea19

Observation bf827e25-0aed-4782-8627-8dfd9377e165 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.320223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.320223Z digest=sha256:82cae0d35c9e4bbc67737213b3eae97e6891c499fa1e33982c4a2e19dd989c56

Observation bba694e8-6c14-4073-a46e-480ac2bcdcea · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.330976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.330976Z digest=sha256:73b31119a741882cfc9d24658590cf0e12ca3edc67d7b4fec1fa3a5c3cbc7bf2

Observation 9e3dde9d-e7b0-4f54-b686-4a4033c772ec · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.336017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.336017Z digest=sha256:011db019c759a47153df55b7ee538328b6caca7786a400e9d668374381b04ecb

Observation d493ed16-3acd-4a52-b132-6fa5eed4ebed · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:43:09.288560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.341568Z digest=sha256:90ae3a348e6be6f70d3bc83ebc270f5e52fcb95eed24d132572a61bf88abfec7

Observation 039c9c1c-7118-4aa7-b692-2ad34cec58f7 · outbound

This paper cites an unresolved cited work.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.346668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.346668Z digest=sha256:71b9af5db47a27c36416db2c782bc48d1f5695ef2261fc9a27c442a816db152b

Observation 76e70a1b-40d9-4db0-b8c3-58ef0c8b9815 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.351843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.351843Z digest=sha256:9a2929fa9163a339d728aa8790db725e6d77c43d180f208de7e67cafc580b9c9

Observation 9f71f3fe-2089-4064-b1c9-612f4db72381 · outbound

This paper cites Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:43:08.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T10:43:08.357447Z digest=sha256:b30a409ca83376c78bf4952951eb3c0ce30656ef9503f97f95101055d428fdac

Observation f5749c3a-7a2a-4bff-86e9-52f2ca6bdcea · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.367718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.367718Z digest=sha256:8241ba31277f1225e77c1b7dadb4f2a823b0bdfc742d287c60e96e7245745190

Observation 746f9fb3-8af6-47cc-a585-49c169dd95e8 · outbound

This paper cites online" 'onlinestring :=.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation online" 'onlinestring :=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.374718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.374718Z digest=sha256:ba9823a9ddfd0384844770df1edb1294577ecd6628d5cb006563301f3d87d6f9

Observation 240ab25e-bb93-4ba0-9e60-afb92eeac574 · outbound

This paper cites write newline.

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation write newline

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T10:43:08.380678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:43:08.380678Z digest=sha256:eded1984128f17d5718ac370955fdd34b1ace7e736fcda6f5b67bd65c20d3024

Pith citing papers

Observation 7b2f1c54-e03b-44af-b3cf-ae100fa0aa5e · inbound

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities cites this paper.

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:32:48.411909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T15:32:48.356870Z digest=sha256:239f298682ab5822d20b78dd107ece1f385606b776200cca59754e771741975b

Observation 1a5d8708-20ca-457e-91b6-9079ea167a94 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

Reference 280

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:38.147584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:38.147584Z digest=sha256:973aeb22e8e3b631b3641db56c1f739df0f4c4d25d850776bd092795aedcecce