Pith. sign in

Paper Citation Record · LEDGER

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion

As of 14 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2505.18115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18115 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:54.579360Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46b57075-b50b-4e2b-8eda-d6c7d4a00491 · outbound

This paper cites anthropic.com/news/claude-3-family, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion anthropic.com/news/claude-3-family, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.891987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.891987Z digest=sha256:dd3c48b6a3c2e2f67f78ad35692e2d5ab4a7f5cad89e52ca0f6c46fba73967af

Observation 782c97b6-a0ee-48a1-98b5-ddaddf9a3ffe · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.984855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.984855Z digest=sha256:a7ed8b95729e531931f257a04adf684f052f47f0a1c7132197414cfb2b2c97b9

Observation 104d0814-623d-4e48-b819-5e760535fa89 · outbound

This paper cites Easyocr: Ready-to-use ocr with 80+ supported languages, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Easyocr: Ready-to-use ocr with 80+ supported languages, 2020

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.057781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.057781Z digest=sha256:d1a3daf1d82a8e8d3e9c8a4cab83e8bfea12cbdb13673856c5c97acec855be3a

Observation 796be892-b44a-4377-ac8d-ab004995c347 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flamingo: a Visual Language Model for Few-Shot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.145455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.145455Z digest=sha256:7b7c62cc6b8c1f377b699338818fb3b3b3ddfbc838d67397b3d758892b6d1d49

Observation 651eafbd-c59b-4130-9934-46bd7538a63c · outbound

This paper cites Visual instruction tuning with polite flamingo, 2023.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual instruction tuning with polite flamingo, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.204766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.204766Z digest=sha256:588264dd674039065bb8850699a2bac249f7c69f78ed7a7fda4abcff3d304a00

Observation c831eee2-e73b-417e-9027-f3269ee61fb9 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.318544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.318544Z digest=sha256:e3c5c378820d32de5ef8d6d18bacdee1d24169808693c169c86eabadd20d57c9

Observation 5cb8c7c1-401d-412e-8e3e-0e7c97190d55 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.428346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.428346Z digest=sha256:dd7395a603a4569e77e9025fc83f80ad6aa392483ee092fea701b42656f0b0ce

Observation b95a4dfe-e21a-45e8-b711-6b0a9c66063b · outbound

This paper cites A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A spatial-temporal attention- based method and a new dataset for remote sensing image change detection.Remote Sensing, 12(10), 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.525458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.525458Z digest=sha256:739d95a871966fcea6b47b0854f2612e325bfd476265d5be52611219080e97c6

Observation bd16a6c8-3983-407a-bbbc-0a74463d333d · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.643807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.643807Z digest=sha256:11060e1b940216685ddaee9b54389f1a45089a9788c62e4c1fe14f3048bda34a

Observation b005e4b9-97e1-43e2-939f-20470ea8ce43 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.747310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.747310Z digest=sha256:29769d3a57a1fad3ba7f79abfd166b9a1dd5a98d883093af12a72f12afa294e7

Observation f9834eef-a3de-46c0-862d-9ffd3fdb35b8 · outbound

This paper cites Lawrence Zit- nick.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lawrence Zit- nick

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.869048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.869048Z digest=sha256:202b2bbc2a1f446d0d276238e0fd4ba40371bbe78c589a627e648de8a122cffd

Observation 9996ac57-7c7b-4b9a-87ba-f060f5621bb1 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.991699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.991699Z digest=sha256:405a062ab30cca51767e4699f1aac5665e69b5c727c650c02cbfb97daaf66902

Observation 2099f931-f9a5-4ddb-ad83-46088d849220 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.067033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.067033Z digest=sha256:100ef020916b722c90014bdc8437aa87092a678ac28959c0f8f79f9145b8e858

Observation f14b0b32-e92c-413b-97d5-96e04a4e1db4 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InstructBLIP: Towards General-purpose Vision- Language Models with Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.157821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.157821Z digest=sha256:8fcea03a959d228a8a3e95fda8292d6e73a38eb5dfd16f6575050f11e444e087

Observation 60b60378-128d-4a51-bfab-a54d8ec06a15 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.256692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.256692Z digest=sha256:93dfd956fbcb64349154f334af3adef2f2d22d689ccc9d75f7effee113410e62

Observation 5b0978de-f67e-4d56-95cc-d7f8b6cb2cb6 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VILA$^2$: VILA Augmented VILA

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.380183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.380183Z digest=sha256:57ca416990d2522a20daf755edf0a9c00382813c17f58a3671905e423bd5d461

Observation d7ecdb0d-2492-49a6-a8cc-52313f9dbd76 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.490128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.490128Z digest=sha256:4508606b4922fccec9ac1459f95383ac9331d315454504f1dbc4e4c53bef34b8

Observation b6714ece-d05d-4107-9b94-f2872aac866a · outbound

This paper cites MultiModal-GPT: A Vision and Language Model for Dialogue with Humans.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiModal-GPT: A Vision and Language Model for Dialogue with Humans

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.592607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.592607Z digest=sha256:7b31ae95932818dcf19174cb5a609b3b8d3a27728d91c31df9c6220d2bcf9801

Observation 5f0a3843-0cf2-4425-99ad-8c4993c74386 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.728782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.728782Z digest=sha256:4ed1bfba73bf55720bb339079cfc5ed3d813adc26d3f63cc7a4959e21c3d8c91

Observation 67454761-f5da-407e-a1de-f6713a34c68e · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation, 2019.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lvis: A dataset for large vocabulary instance segmentation, 2019

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.817491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.817491Z digest=sha256:685e00ef2f5aafc398e3531e20601557085a186817c6dc789951bd49411f36b5

Observation fcce6116-1ad7-4ac3-ab54-2bfeac289962 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.902463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.902463Z digest=sha256:d0e7aba13bd9ffbc0e70fe1f9fe75bc6415a25ae65120ddc5ba4d2ebe1fceaf4

Observation 138539a2-acc1-45ab-ac6e-ffe5b680a44d · outbound

This paper cites A diagram is worth a dozen images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A diagram is worth a dozen images

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.015592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.015592Z digest=sha256:c6772837f52b62c21abfc25f7f8b2e4a81365fda3bf445eaeda4b4506617e52e

Observation 1701edd4-247a-4a7e-8207-88461434a3e7 · outbound

This paper cites A hierarchical approach for generating descriptive image paragraphs, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A hierarchical approach for generating descriptive image paragraphs, 2017

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:00.174744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.108696Z digest=sha256:2a3589226f0d379736561ca1691f5201cd7e515bfe016d6b25884cff749f7d53

Observation 52b62b25-d5d4-4d13-9baf-7587316d19f3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:00.001617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.174900Z digest=sha256:102fd9a40fa61eb42ec0a4edc9f6cfff0bbd6e39b30c7e511ad6b5435fd6c264

Observation be1b42e2-5ddb-4079-abfe-66c8fdcd20f3 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.240559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.240559Z digest=sha256:131f16a8e358e580f43dd59453b97aaa496934dabf4d146c476e37dc34ab8451

Observation 371ba48b-ccb0-4cbe-afbe-c7981a230f11 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.346808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.346808Z digest=sha256:0de1b57d12f03d301e4a01ab402c08f6a683ccb34e829859afc41340bc445c30

Observation a77ebe23-372f-4444-aed5-cd0fa947b70b · outbound

This paper cites Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.903935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.455236Z digest=sha256:27e30b25e95daa1354813761a2122b9a5997ae0d9639ca12954db34c7dc035b0

Observation a8c558c5-a72f-45e9-a9d7-1f7ebbe0173a · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.502166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.502166Z digest=sha256:0ba6279f016da54bc33c2a6f42c621cbade6b40dca598afa147886e8f8043f7f

Observation faaad85f-8620-4213-b570-1c8ea8c4ffc9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion VideoChat: Chat-Centric Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.563235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.563235Z digest=sha256:7e902ba95d6efcb10a01f843dee71ff55f849c07d1f0d6107c70ffe553399f60

Observation 14856309-26db-41ac-9566-3a4537ffc053 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.620424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.620424Z digest=sha256:8810a48c09cf394a61951d8be00a917ce4975979d8d486efb0d0736227bebc66

Observation c585b79d-269e-4c02-9308-fcf74de76a1e · outbound

This paper cites Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions, 2018

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.790604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.686490Z digest=sha256:b989c39ae308534a4e0a1ef4a2bb05a4c2920e67fa19f98046c8cf783c8f4326

Observation 94969b14-d44d-47c2-8861-5070f20e3530 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Evaluating Object Hallucination in Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.747886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.747886Z digest=sha256:f7a83f226301dcaa542e55858a6494799e93fdd718e72f8b637eb59efa2a531f

Observation 7c0382da-693a-4f3a-9776-458f24d0c6e3 · outbound

This paper cites Visual spatial reasoning, 2023.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual spatial reasoning, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.666472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.838964Z digest=sha256:8d849bd78ad7685ccf9132c1fea05427bf195192dc4cebbd850ca98d7ca5d7d1

Observation cab813d8-41d8-4017-b08c-62601713ff13 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.545183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.905071Z digest=sha256:50fb652248700f820f57433f074d96873dbc7713d5a9c0e8c3b2e38299206872

Observation 215f9bce-d532-40d2-9738-92b79cd9c2c3 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Re- moteclip: A vision language foundation model for remote sensing, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.390252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.020164Z digest=sha256:48ccf9673d0d07a79c94b8594b7a77c2e9e664c1636429cb5edbdadb02c18407

Observation 7d2fbc4c-7672-43e1-8e6d-30159bf48853 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Improved Baselines with Visual Instruction Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.139418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.139418Z digest=sha256:51cc9b18463b2a073174e7945488a2bc87167e508ad1b8a18aaedd5b1026e884

Observation e31ce97b-c664-4210-859c-2a0a52e2a0a3 · outbound

This paper cites Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual Instruction Tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.264947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.191652Z digest=sha256:fb66c6ffcd50523b04bbd8094b00bee18fcd03ffc53ff13ca28e72da406faa14

Observation 06f694ee-ae57-43f3-9b51-af7dfbf81379 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.260198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.260198Z digest=sha256:ad1009fdd3c8c25a1c583dc3c1c1ae29d17d65166e8fc9efabbcc682ab49a55b

Observation ef8d6850-20a7-4319-ab31-b52b2775110e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:50.359568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:50.359568Z digest=sha256:22770fac56c14093131be5146ef4b1af51528b6554aea51287dc010925342f39

Observation 1677e10a-078b-4e21-a9f7-49a49750125d · outbound

This paper cites Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cheap and quick: Efficient vision- language instruction tuning for large language models.Ad- vances in Neural Information Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.127343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.416318Z digest=sha256:de43620e6c166d0df0d386aaa13e735a9aec14e5f8ca7c7cadb4300c9c6bf921

Observation ca9547b3-7859-44f6-81e2-abd7d2890ecc · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:59.000551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.498036Z digest=sha256:2530d2bc198ade06184a7d318508ba48c615a1abd8a6706b8d43810c836527ec

Observation dd0db20d-f593-4607-8163-e0c6fdd3d277 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.871384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.601306Z digest=sha256:4ac5154011e4ea2ce6b269919aa7edac43127ddd98bb0cd97af5e2a30ca64b67

Observation 784ff458-d379-46e3-9ff8-67e5b2e9d52d · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.756318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.652514Z digest=sha256:5df5322b77c970dc97b8c4c0a4b1b6b62076360b48033b70c0521bd586ad0cc5

Observation d21ae786-c348-4a0a-a84f-33f36383efbd · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Docvqa: A dataset for vqa on document images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.625510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.686865Z digest=sha256:7b8aacf9abdf0250ba812166a36143ad03cb1f84daeb0de638fc5a76403ec478

Observation 06388035-1aea-42a3-be92-4b1475ccde45 · outbound

This paper cites Infographicvqa.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Infographicvqa

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.498770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.750838Z digest=sha256:13eacdbb932d3b25498063289221331dc633df8a0a31e65bc08576c4ea997466

Observation 4776cc86-5905-4a10-a79d-db2de8dde0bb · outbound

This paper cites Introducing llama 3.1: Our most capable models to date, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Introducing llama 3.1: Our most capable models to date, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.357548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.862393Z digest=sha256:0654a51cb8061f5fa0d208914f002fc4e998281efd18f324a8da4086529a15f5

Observation 66ace162-e509-4286-86b4-c8dd4db39ea8 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Ocr-vqa: Visual question answering by reading text in images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.243549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:50.960826Z digest=sha256:512ed4ed3d1766f772f18995da8b31341b85caaf3af0fd38d89cff4e15b1147a

Observation a4a480c5-8733-4000-9d67-3dfc28e2529d · outbound

This paper cites GPT-4 Technical Report.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion GPT-4 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.068769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.068769Z digest=sha256:c6c18a3b493e9bdd626ff5c3de6b7c15220bf6902f86e0420a6412a0873f1f0c

Observation 7056f086-e554-422e-a905-bbf7d132200e · outbound

This paper cites Gpt-4 technical report.ArXiv, 2303:08774,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gpt-4 technical report.ArXiv, 2303:08774,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:58.096391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.149206Z digest=sha256:8b1760411b4c2f823ed07a8749ad68863f628f52b1fd4ffe0d9363f0f67c5119

Observation 5cae0bc2-df35-4cb3-a429-ba3d5a4969ae · outbound

This paper cites Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Im2text: Describing images using 1 million captioned pho- tographs.Advances in neural information processing sys- tems, 24, 2011

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.959041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.239072Z digest=sha256:a7c68087d47dcad6b45dc4bc1ed2c063376c0d7940370aa25646a936766442ee

Observation 54bf0e9f-0d87-4f8f-b08c-d3b7863163e0 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.800017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.305339Z digest=sha256:7efb06b0a3713d3ed98469e2c9ff131544d7e95b385fc2087fdad1bfc03b2c6b

Observation cd0bd085-48de-4114-8338-4ecd557293d4 · outbound

This paper cites Connecting vision and lan- guage with localized narratives.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.426771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.426771Z digest=sha256:335029f3b8f39c17cc5f7602d68448d821958d69e71a64fed42642f7988111c3

Observation f3b8d652-7b42-461e-b27b-47389529bc39 · outbound

This paper cites Connecting vision and lan- guage with localized narratives, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Connecting vision and lan- guage with localized narratives, 2020

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.640816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.548336Z digest=sha256:c6cad9a56728cbaf8ef2a4858ffa5ef1d1b620256cdb6e250aaf01c725c6bead

Observation 8d661666-9ccd-4410-9e1f-88ff7c3cc442 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.481113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.644142Z digest=sha256:bca18921fc155cc94c089a4e99ebc8060e260fd5d32f1b0e38d1cbc29df41862

Observation 91df7c65-2a47-41ee-8a80-046a9b9afc38 · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Sam 2: Segment anything in images and videos,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:51.718046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:51.718046Z digest=sha256:4d56ecf0069dd72c05e0cb1493f3037df4d003fd37c93cefe044537b21589e58

Observation ef1219ee-359e-4e0b-b56a-f987b3b47ed1 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models, 2022.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Laion-5b: An open large-scale dataset for training next generation image-text models, 2022

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.268836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.814560Z digest=sha256:e15f1c110985ff053da58c925ec61f43c23d4a3e8860007438791380d237c9e8

Observation d9436920-017a-44a7-afda-5d1af3ec8759 · outbound

This paper cites LAION-5b: An open large-scale dataset for training next generation image-text models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LAION-5b: An open large-scale dataset for training next generation image-text models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:57.125445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.909444Z digest=sha256:823dcd0cfdab36f4ce0865ed7e77fd43aaf76362a1c1f144137c58a8b7c6f7ff

Observation 90b898b8-7c71-46e2-95cb-1566aef7fb9e · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion A-okvqa: A benchmark for visual question answering using world knowl- edge, 2022

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.954674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:51.987410Z digest=sha256:ac0f5d9fe0e0c18b8951fd042602776660933ee18f070a592e826421bd286384

Observation 3356d403-ad9b-4824-8be7-6416ca85b22c · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.061066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.061066Z digest=sha256:99fe944bd6b9b5ba55d8343452086b1ce072d328d773306a1a9329516e7bb0ec

Observation 004772e0-2b2d-4d52-b3ea-10f23bce2144 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.139069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.139069Z digest=sha256:bcb21c807e99dc7d52ff7640f4b00860d43cbde2ee55c768c2f3c7a2a0b68502

Observation d92235de-80c8-4b4f-bef1-c0820066dc7c · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension, 2020.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Textcaps: a dataset for image captioning with reading comprehension, 2020

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.806437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:52.225514Z digest=sha256:0349d59d26a49e9c2f2efb7bc1bd342897a0542a35a086b7e9b9090e35177a7d

Observation a49b2bf3-6aba-4f4c-9b1c-21bc58b08f63 · outbound

This paper cites Towards vqa models that can read.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Towards vqa models that can read

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.621102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:52.325073Z digest=sha256:5ce60625ecb54edfa4a2224ca6e5514a87f44587ec14e4b8f3b076610f882aa1

Observation fa194d73-e61a-437c-90c9-b34f3b6187a6 · outbound

This paper cites Expressing visual relationships via language,.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Expressing visual relationships via language,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.482837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:52.400976Z digest=sha256:d2754d20896de846149fe198e2a2a9d77272f33fba2f48ffcd8fc6400c2c7a6a

Observation 7869d6d7-f163-4c65-bb7d-4966d4508cb6 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages, 2021.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Vi- sualmrc: Machine reading comprehension on document im- ages, 2021

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.336625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:52.451431Z digest=sha256:1e0cc231cc67308da03c873ed44697d6ae738e18511c85c72b40197f5d51b3fe

Observation a922125a-d40a-42cd-88e0-b8e911720714 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemma 2: Improving Open Language Models at a Practical Size

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.521778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.521778Z digest=sha256:5432a9d8154a9d22f85e902480a7ec47f99618a76a30b5f2a2c6f75b50e53213

Observation 9c42d8e9-8075-4a81-9d7f-c781c5e161b8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Gemini: A Family of Highly Capable Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.582430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.582430Z digest=sha256:e8a9e6a98cfa3c44014e99c688c3b2deb4740a17676579a9111a0bfce7b1946a

Observation 38bf7cf7-0be6-46ee-b247-97608eee72ae · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.700977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.700977Z digest=sha256:2ce0de02aa0dfefa8d9c0f24578e6f3a30a8c6b268cb1eacf2f58bb0dfa81e31

Observation 380b9a90-ab99-4a3d-8c4d-52846bee1ac0 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.789583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.789583Z digest=sha256:4d6c7e1c8516625984322ecbfda6ac2f60c7590646d4da4d34dd8679918f5cac

Observation 3b1342ee-9a8c-49d0-a595-5350d486a591 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:52.880552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:52.880552Z digest=sha256:85b2cac7728d09b278ca0b94164b8b1cc443bebd636fed55fb4623f52ba4a430

Observation 198dbafe-e916-4c44-9197-d8817203973c · outbound

This paper cites Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visionllm: Large language model is also an open- ended decoder for vision-centric tasks.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.167971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:52.996429Z digest=sha256:af1cfda9828ee364ec17c534177beefcf5bfcfad550b1a831e125220f8a31170

Observation 5b49823c-2299-4c5a-b696-a4f930bba88f · outbound

This paper cites MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.072593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.072593Z digest=sha256:e99095445bbd2f1c0417d503cf989ebe6e75f646719d682aad477be06768abf5

Observation b4595a73-fa6a-4f5a-9acd-dace869b8128 · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.136267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.136267Z digest=sha256:5a5f84ac62042eee480ca4c7d109093b35a475c70f37bd1ef52b34e9115a64a1

Observation 49d01f9d-7713-4b5f-be24-b4ddca4e6e66 · outbound

This paper cites Grok-1.5 vision preview, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Grok-1.5 vision preview, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:56.031446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:53.189797Z digest=sha256:eb83f177c50d7a3b27e68ba48aa514a6dfccb0537e6ef525dedc1d0282f9619c

Observation 7b2e6457-62fd-42fb-a0e8-3d9c75c8c7c7 · outbound

This paper cites MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.282556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.282556Z digest=sha256:b5320119050f8e715c4cd4451ececaaa7d43531dea74aee74b0b1b786f30f5c6

Observation 6942175e-fc92-4847-a539-083a5e111a60 · outbound

This paper cites Depth any- thing v2, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Depth any- thing v2, 2024

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.381001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.381001Z digest=sha256:49d3a96e39f99b0d62d4f83ed05c9d33ab86a80a9a36715994d38ab41a056624

Observation a3f2b3ba-ba36-44b4-85d1-72b640d28a8d · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.457998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.457998Z digest=sha256:415c77f61f010844cd22562f0d519310f0e8bfbd5e93337ddb35ab032a7b2e5e

Observation 0e25ae2a-754a-4ccc-9765-bee5c3134768 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:53.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:53.547479Z digest=sha256:78cc1f5dbbc94d90f65d5d56cdc8256138c2409eb1e6b071d645f995d1eea90d

Observation 5b4c1e2a-b493-4099-9e37-d7226f266c96 · outbound

This paper cites Berg, and Tamara L.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Berg, and Tamara L

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.940987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:53.643876Z digest=sha256:eebac429fbe6b96421eb4f638087a79fdba3a2429ad81d7846b36a7717b459a8

Observation adab7efc-a886-42f9-bb80-f7790886c47f · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.792963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:53.717638Z digest=sha256:3b3eaeecc7200348e053cf9fe52e75e4f36dfd91366580d0138071f28caa725f

Observation ab4b3daa-e6ad-49dc-927d-d3c5aa0b623d · outbound

This paper cites an unresolved cited work.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:37:55.668988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:53.805718Z digest=sha256:496e8237035636386282f1fdd743d936fbe3a667f0efde853ff9833f8450a538

Observation b03a1f4f-3d0b-4a85-9256-5af7f73046ce · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Rsvg: Exploring data and models for visual grounding on remote sensing data

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.534422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:53.898696Z digest=sha256:d9f2e180cea54c995773bdcc746f14282638d622dcfa004c278a4722cca6fe7e

Observation a6a07919-ca3a-4726-933d-7a98911f55b3 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.393162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:53.962404Z digest=sha256:88dab6c3f2254177f6e5c8019f204d2687c2e67f724142b9640dc811dce2e3b6

Observation 15480673-59f4-46b2-b0ea-d123ad922208 · outbound

This paper cites InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.044923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.044923Z digest=sha256:7ff75289005fa253c1a8752bdcfb7a6a2921d95145f9d3ca9bb666b3e9f5ed69

Observation d2f15a56-beea-495f-9df7-5032803d62a4 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.116967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.116967Z digest=sha256:5d550d7966d7a05425f0cb55a792abb8613751c210c5884c163523dc679c5de3

Observation 1253bd24-febe-40c5-a414-cd03f10b8c9b · outbound

This paper cites Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Hierarchical and robust convolutional neural network for very high-resolution remote sensing object detection

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.246086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:54.194802Z digest=sha256:24ce6d42f6df7404d3091ba711216faea285a06f4e8f85632c55733083d4e6a5

Observation a346cf23-6cad-4a8b-867c-efcd33e76a44 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.258922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.258922Z digest=sha256:d7b44305f248e8406f03805c81179f7f879cdd0359c5ac9fccba8db4976d4c1f

Observation 05c6122e-8d8d-44a8-bd5b-e427e1374c22 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SVIT: Scaling up Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.331667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.331667Z digest=sha256:6d828b1a6927fa872db3d6facd6c713ec3ad739fff55186786d23fdfacc4810f

Observation f40fa1de-f74d-43db-8c29-353c44361fce · outbound

This paper cites ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion ChatBridge: Bridging Modalities with Large Language Model as a Language Catalyst

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.397753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.397753Z digest=sha256:3333ae3b31fbce7563206f2b97fe7415a3debf80b1f2be487ca78daf0957cd78

Observation 6619e153-8a2f-46be-a2bd-06e10ce65ca6 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.505962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.505962Z digest=sha256:dcb40517f21ca2ba7f0a630add55bccadf32c04761df141ef9eb3a05f8500618

Observation a5dbdd01-7047-4333-b9db-7191c179e2e4 · outbound

This paper cites Visual7w: Grounded question answering in images, 2016.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion Visual7w: Grounded question answering in images, 2016

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:55.062673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:54.579360Z digest=sha256:3a5043ee6ef2f4f05a354e58d422c46f0a108759c4ceaa52a433001135627812

Pith citing papers

No inbound Pith citation observations are available.