Pith. sign in

Paper Citation Record · LEDGER

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

As of 9 August 2026, this Paper Citation Record lists 100 of 141 outbound references and 0 inbound Pith citation observations for arXiv:2506.02308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02308 v3

Coverage vector

measured 100 of 141 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:13.382009Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 141 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e72a6ff-7a11-48b4-ae88-cd378c6cc080 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.527558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.527558Z digest=sha256:693c068172ab7fd8e5c8b3846e37fa8e8d7f789c0a168b10184cff389ab3a8b4

Observation 5d9e4a59-fcbb-4154-a3dc-00777f240c32 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.562670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.562670Z digest=sha256:903b159abca7452c86a46f44b5b124f929191aa3540c40e61e48d91ecf797935

Observation 9c4ecb82-6154-4acc-a4fb-6b44987a6151 · outbound

This paper cites Gated multimodal units for information fusion.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Gated multimodal units for information fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.700976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.700976Z digest=sha256:c18ba0598e64f99d56c216309129be030f14cb916f87bd6d4222e107b505cace

Observation 6ae9519e-fc8d-4d7e-8576-94f679fb99ae · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.747469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.747469Z digest=sha256:e8e6dca43d2d30ac6d10a9e0fa954b27af9758a687e0aa722d1bb999b041e16a

Observation cc05ce37-688a-45b6-a477-0b980139db4f · outbound

This paper cites TouchStone: Evaluating Vision-Language Models by Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping TouchStone: Evaluating Vision-Language Models by Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.798457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.798457Z digest=sha256:afe11cf2b5fef08a97b72ee61d7464809c9dc224a176215be8c24dc3366eae4b

Observation 5e354e26-95aa-4a78-a683-4a3af107b05d · outbound

This paper cites Multimodal machine learning: A survey and taxonomy.IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2018.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal machine learning: A survey and taxonomy.IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2018

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.843432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.843432Z digest=sha256:557b53efd0693f365b1c432c5b2d30765f79fa07ef27095b558fe0e97a83514a

Observation 6f29f9f5-bce4-4f82-b8e0-885e41700c97 · outbound

This paper cites Routledge, 2014.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Routledge, 2014

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.906604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.906604Z digest=sha256:389e0177f6d3accbb731b69535d4bb9795e1cf8c2b07790d44791fe50a6758aa

Observation e89c36fe-d388-433f-bed6-d92df087b41a · outbound

This paper cites Introducing our multimodal models, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Introducing our multimodal models, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:04.936237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:04.936237Z digest=sha256:20e778827f0a77b6206cea3d8d673d9ac271b7180ed0980bc4d539916066fb2c

Observation 6433d950-8cef-4b59-9237-4728e1b49ee0 · outbound

This paper cites Identifying beneficial task relations for multi-task learning in deep neural networks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Identifying beneficial task relations for multi-task learning in deep neural networks

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:18.420739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:32:05.006588Z digest=sha256:c333e07ca4ae80fc180c3225de4f10b843d603587e717b6ee6c0871c7bcbc21f

Observation cd280595-afce-46ff-96b7-565d47a44a47 · outbound

This paper cites VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.026455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.026455Z digest=sha256:592c3a5a5173baed9869206a442f864af700d0f403fda055bde32b83e900a515

Observation dcab9994-dcdf-4bb0-b60a-f1da440d338d · outbound

This paper cites Decimer—hand- drawn molecule images dataset.Journal of Cheminformatics, 14(1):1–4, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Decimer—hand- drawn molecule images dataset.Journal of Cheminformatics, 14(1):1–4, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.029575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.029575Z digest=sha256:a797b55d64a1054b58f97cb35e529c63f229800014ecaa6e0e343e0425f9a81b

Observation 21781edd-69dc-4733-95b0-e8b71e70acc6 · outbound

This paper cites Language models are few-shot learners.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Language models are few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.037177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.037177Z digest=sha256:0af67cdc53cde04be9a27942feb0ca0dd698ac6eaea50c82e46223da00b31546

Observation ad46f82c-7692-4788-ad1a-1e9f16dd11ad · outbound

This paper cites Multi-modal sarcasm detection in Twitter with hierarchical fusion model.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multi-modal sarcasm detection in Twitter with hierarchical fusion model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.151820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.151820Z digest=sha256:cd1d8be82cb930b4f5d5d93d39dfc9b5ad27219fbbab819b85ffc09c545fb46b

Observation d774d44a-756b-4322-ae24-040c575f5fd8 · outbound

This paper cites Multitask learning.Machine learning, 28:41–75, 1997.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multitask learning.Machine learning, 28:41–75, 1997

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.213593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.213593Z digest=sha256:45f35566aca56291cc9daa8df1abe99243a5e18da34a86daf4db9280c2da69e5

Observation 76fe4df9-9b3b-4f6e-907d-f5b74a3c32d4 · outbound

This paper cites Uniter: Universal image-text representation learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Uniter: Universal image-text representation learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.268085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.268085Z digest=sha256:724a0432e58529ca5b9a0a655dae877a7e1668feba88cc55c97dad0d1ad08de6

Observation 0f224f8d-8a5b-49b5-afee-cc1d273be619 · outbound

This paper cites Remote sensing image scene classification: Benchmark and state of the art.Proceedings of the IEEE, 105(10):1865–1883, 2017.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Remote sensing image scene classification: Benchmark and state of the art.Proceedings of the IEEE, 105(10):1865–1883, 2017

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.304510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.304510Z digest=sha256:50eb8e7abbc4a9df2b281bcd668df3040fc7a0fc639687b203c6a2a7bcf1ad03

Observation 32516e51-e7bb-4f50-b8e0-fb3821c725fa · outbound

This paper cites Scaling instruction-finetuned language models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Scaling instruction-finetuned language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.334394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.334394Z digest=sha256:75bb3dc1ea105d6493c2c521327be9dac9c33e5e99c56dcfe6a31395ee227d94

Observation a342a713-9f91-4ec9-8763-5e592fbf8f3b · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.365315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.365315Z digest=sha256:767a54be61ecb3f04654aac71db7fb7ffc28022d553a4a308bf23f39608d500a

Observation 173bcff0-9f5c-4431-bf94-66f0292c5824 · outbound

This paper cites CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.400551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.400551Z digest=sha256:65db2f4385614643738eb051fe62b17b0b0feab048a57581ebe0840ed0338891

Observation 9902121c-7620-4247-a60c-fac86ab9ee71 · outbound

This paper cites Rico: A mobile app dataset for building data-driven design applications.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Rico: A mobile app dataset for building data-driven design applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.480514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.480514Z digest=sha256:8de4f65386950a0c530b92592eb345dbdd0fc8b6d1a6780e16990b7691cb27cd

Observation 1fa4f4d7-5be0-4a03-9d1c-55e166867a6e · outbound

This paper cites Multi-task learning for contextual bandits.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multi-task learning for contextual bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.627088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.627088Z digest=sha256:2bafaace91d25c097e6d5a062159a07219319fc2d164eea3435a0bb4bad034c1

Observation e9e143eb-ebcd-49c2-99d8-933fd35ba783 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.719625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.719625Z digest=sha256:bc580334825302d5eece2edef993dc152bb7d9a51c9412d83fdd588f6fa6e70e

Observation a7a989c2-da13-49c6-a0bb-e1c9fd05f7b3 · outbound

This paper cites Regularized multi–task learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Regularized multi–task learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:05.877074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:05.877074Z digest=sha256:e9d6a359478c0804cb90e0295d2d2a0ded7da1083d6dd92247ab2bee44688e31

Observation cc85cd95-1328-4e69-a9d1-3e194e70c192 · outbound

This paper cites A survey of current datasets for vision and language research.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A survey of current datasets for vision and language research

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T11:32:16.878332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:32:05.967677Z digest=sha256:e1bb4c90cad070f875228ef5e6ab588426c8e7477848d7a570e8dbf6c4a04cf1

Observation 768b9948-d9bd-4d19-8276-08797925237a · outbound

This paper cites Efficiently identifying task groupings for multi-task learning.Advances in Neural Information Processing Systems, 34:27503–27516, 2021.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Efficiently identifying task groupings for multi-task learning.Advances in Neural Information Processing Systems, 34:27503–27516, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.122368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.122368Z digest=sha256:ce7a0dbfc7c52ade509cfcc778df7cdf879b00ef222fb7c3f5634af4edf11de8

Observation ddc20d4c-2299-45cf-ae63-000c5ded65ad · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.278770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.278770Z digest=sha256:2c999f1597ea4812eb4fd367d1ea7775f9989e0de0d73532901105adaa02311f

Observation 290ba769-b032-4440-9962-5770afc355c6 · outbound

This paper cites Enhancing knowledge transfer for task incremental learning with data-free subnetwork.Advances in Neural Information Processing Systems, 36: 68471–68484, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Enhancing knowledge transfer for task incremental learning with data-free subnetwork.Advances in Neural Information Processing Systems, 36: 68471–68484, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.390507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.390507Z digest=sha256:6975c8bfc8443cdbd0ad69b6ba20aad33fb775a99386af27427e1292f1edb457

Observation 0af10fe6-21dd-41c2-9bbd-98163be95788 · outbound

This paper cites What makes the difference? an empirical comparison of fusion strategies for multimodal language analysis.Information Fusion, 66: 184–197, 2021.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping What makes the difference? an empirical comparison of fusion strategies for multimodal language analysis.Information Fusion, 66: 184–197, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.461953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.461953Z digest=sha256:7ee2bc611b394550fa321af39043556d8f48e09db59650688908ddaff7d0434a

Observation 1fdc778f-7ce6-47a4-a7ae-e42df988440c · outbound

This paper cites Challenges in representation learning: A report on three machine learning contests.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Challenges in representation learning: A report on three machine learning contests

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.595058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.595058Z digest=sha256:1a96175198f695591e6859deb4b9fc597b7a4d029e6610b65ed2d26c74484f30

Observation d046d532-f878-482a-ba56-fd7bf699edb0 · outbound

This paper cites Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.772655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.772655Z digest=sha256:cc6add556f76ca653895c37f37c0c471555d8f258f919516ba438a31f4f0d375

Observation a3f12151-0bfb-48b8-9b09-a99a9d9051e3 · outbound

This paper cites The Llama 3 Herd of Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:06.924038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:06.924038Z digest=sha256:4a88617c69b9bd8ff3fa7db2d1e2321116e64194e25105884d2dc638a3e1b170

Observation bbf73a16-98ee-4cde-b53b-9078408175a8 · outbound

This paper cites FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping FuseMoE: Mixture-of-Experts Transformers for Fleximodal Fusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.087252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.087252Z digest=sha256:a50ebb6cbd9aeb1ebd350d713c3bdcb5956707f075e57c1951fd3e44965687d5

Observation 02835773-a720-4aa7-aa41-6c85c98c5c77 · outbound

This paper cites The movielens datasets: History and context.Acm transactions on interactive intelligent systems (tiis), 5(4):1–19, 2015.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The movielens datasets: History and context.Acm transactions on interactive intelligent systems (tiis), 5(4):1–19, 2015

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.200221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.200221Z digest=sha256:8d6748fc216df3796b7fa20ad087ff29a522443dc84fa3a361ae3d1125bec116

Observation fe890e94-9343-4d1e-a87d-64e1f16ff1e2 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.313230Z digest=sha256:a1267bef4a03bc3e5673aee64ba4ef850c3b55270111febced6f79c1998145f5

Observation 5d8ce1ed-9b3e-4a1a-b565-7b8c9039a1ce · outbound

This paper cites Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.406001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.406001Z digest=sha256:a08726100172f23c6fbb3c68bfb687489be178a1805be5ee1aa7e22962edd157

Observation f73ceaf9-c10b-4494-a4e0-5da38ec49ce0 · outbound

This paper cites Framing image description as a ranking task: Data, models and evaluation metrics.Journal of Artificial Intelligence Research, 47:853–899, 2013.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Framing image description as a ranking task: Data, models and evaluation metrics.Journal of Artificial Intelligence Research, 47:853–899, 2013

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.530750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.530750Z digest=sha256:2e0fb26af52c243b8a9de77a02865d6d59f187c17abd37a8364795f11f75fb31

Observation 9107c440-7526-485b-83cf-399bf2592576 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.603688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.603688Z digest=sha256:65aa15b9e7498d7103490ffbe7bc1bf0ab3d78b51e45775026bd6eab61f980d9

Observation 79624a44-ee99-45fe-8a2f-a6ccd07ffcd1 · outbound

This paper cites Language is not all you need: Aligning perception with language models.Advances in Neural Information Processing Systems, 36, 2024.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Language is not all you need: Aligning perception with language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.706753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.706753Z digest=sha256:7f8e08f6698e4b3525a96a013b39ee6b101582690f0f4feb5ecf05c78f702a45

Observation e07bd960-e60c-4927-b867-ed39c1f51aca · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.801704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.801704Z digest=sha256:1e2487235351d11bac00c07692058a9989ee57a67598160cd76be1c8350f0a6a

Observation 9d117f85-ae8e-4a0e-ba64-07e09f48724c · outbound

This paper cites MemeCap: A Dataset for Captioning and Interpreting Memes.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MemeCap: A Dataset for Captioning and Interpreting Memes

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.886072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.886072Z digest=sha256:47ca3c3b380f4ac8657c7a508f665556ac0c3acda657bd46cd0d8f8f1d1813c7

Observation 3e07980a-456e-42a0-9235-cebecffaea29 · outbound

This paper cites Grounding, meaning and foundation models: Adventures in multimodal machine learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Grounding, meaning and foundation models: Adventures in multimodal machine learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:07.962103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:07.962103Z digest=sha256:1008db780caee5e75657817fd0b4070771c07b4813643e1eae8fef736ba31e07

Observation 5b6b26b7-74c5-4438-b0ad-ca60c5e80416 · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.057119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.057119Z digest=sha256:92b1451e60249af172ed5a783f056fbf9ff50e36b8da5a3fd3c04deb05f5c7fc

Observation 6ba146e0-c0e7-482f-9c0d-e1241aa6cfcc · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Grounding language models to images for multimodal inputs and outputs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.134800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.134800Z digest=sha256:ba734e183f18bdbb039f47b94dc56db56194da4915227f44e590729e99d81edb

Observation ebbdc4c9-9f27-46cd-8635-f465165eca7f · outbound

This paper cites Integrating text and image: Determining multimodal document intent in instagram posts.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Integrating text and image: Determining multimodal document intent in instagram posts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.213341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.213341Z digest=sha256:dc35160246b0017a3175fd9f5602c9f864fce4576d10fec4dbe4c981f1aefd7a

Observation 81367ec7-4bf7-4d58-8c13-4d395ffa3b38 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.289284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.289284Z digest=sha256:9220b31c58c0e1da4c278648b31cbc8d2152594fd5a1bb3870a3425d8caeb907

Observation 7bf7604b-2f7c-40fc-a505-bc6a1c7de0f0 · outbound

This paper cites Visual question answering in radiology (vqa-rad), Feb 2019.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Visual question answering in radiology (vqa-rad), Feb 2019

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.393758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.393758Z digest=sha256:2264759a1dec8d34034d7d8eb3bd63154fe516292c88715e3ee7dd353e1d9957

Observation 0e1cea30-994e-4a38-829b-4aa36efec105 · outbound

This paper cites Instruction matters: A simple yet effective task selection for optimized instruction tuning of specific tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Instruction matters: A simple yet effective task selection for optimized instruction tuning of specific tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.494742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.494742Z digest=sha256:c079c867efd21dcee9bee2327d62b829e52e9e45098055c1e80ef1d2ac2329e1

Observation 12000cb8-bd3f-4ebb-bdd0-71b7a6c84ed9 · outbound

This paper cites Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.590738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.590738Z digest=sha256:2fbd9156fd02bc86f3659a57f150667c23d56925dd29632e11334a004d32802a

Observation 3339b701-7b54-4569-b8b0-1a67bf608d74 · outbound

This paper cites Holistic Evaluation of Text-To-Image Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Holistic Evaluation of Text-To-Image Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.663980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.663980Z digest=sha256:150e614cef11364b69b03abda3a553132da5dc91455b9054085d6087a235ca80

Observation c37fffa3-e70f-4ba6-97e2-2061b0cf5c92 · outbound

This paper cites Enrico: A dataset for topic modeling of mobile ui designs.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Enrico: A dataset for topic modeling of mobile ui designs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.716998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.716998Z digest=sha256:d508e9d5b0db7f567958f5c0da3023f0fd38dd510e8525cfb70905a40220ca23

Observation 90f00e84-b636-4dd1-947c-a54aa87792ed · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.822173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.822173Z digest=sha256:227fddfabbd07db08681ae24ab6154105b8495a6506b942a5db94ff1ff9b60b7

Observation 8a5b6c77-ff9a-45c1-a1cf-45f3239321ec · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:08.910355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:08.910355Z digest=sha256:9f782d7c312df44f2791f17d173901ae895d69c79278254690d06d0b0f99634b

Observation e19ecf89-d1a1-4719-b260-ce494f80e05b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.002796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.002796Z digest=sha256:23db1201c8624cc13d31e9c8a4e76780fa6cea90ea27b7f7de372969afc693c0

Observation 8675a51e-c40c-4633-b0ee-8bd0ec2ec42a · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.111934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.111934Z digest=sha256:96127e07bfe78a9162c0eeb12619750594eb2becbb20bb81bdcc128791fc15d5

Observation a2fa1d85-5890-4ecd-9287-59c6b632f8ef · outbound

This paper cites Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:18.162639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:32:09.209151Z digest=sha256:195ccb4d9b087665a205ac3eb891acb0cdf5470ff692c74184929450a6ffab18

Observation 7251f186-dc2b-4614-a44f-c7f06692135f · outbound

This paper cites ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented Benchmarks

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:17.935739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:32:09.315467Z digest=sha256:b2e2c13dbb7a0a1cde1faabe2fb9b5bb094265d86a63cc55ae0b777a4819c7a6

Observation a3406c66-9c72-4b2f-8325-7026a08654c5 · outbound

This paper cites LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:17.755225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:32:09.418694Z digest=sha256:c723c48711616f0f57bb73b8cf1486f096e512e24dfd888c1e23b4d125c2c443

Observation 72297dd7-f86d-46e7-bd84-320cc9fdadf0 · outbound

This paper cites Multibench: Multiscale benchmarks for multimodal representation learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multibench: Multiscale benchmarks for multimodal representation learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.534609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.534609Z digest=sha256:85f072a2ca5eb1a4b98ca860cf1ca1792efa90f407be297d7cfd3426636264a6

Observation 70d246e9-cf60-4749-819d-90584ff1b2f8 · outbound

This paper cites High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping High-Modality Multimodal Transformer: Quantifying Modality & Interaction Heterogeneity for High-Modality Representation Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.629828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.629828Z digest=sha256:3d2a00f6bebf1ab3bddd77f84345c44ff8c8bf67faf850031100a0f81bf8b0c0

Observation 9f41e533-ab7a-411b-ab1c-e99c48761650 · outbound

This paper cites Quantifying & modeling multimodal interac- tions: An information decomposition framework.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Quantifying & modeling multimodal interac- tions: An information decomposition framework

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.749214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.749214Z digest=sha256:74c3c19efa2748f65bbf7850c3039ae414cac849c7f61c6925f3a882c431b369

Observation 6b22747e-c6cf-4576-8539-59fe22518dc3 · outbound

This paper cites Foundations & trends in multimodal machine learning: Principles, challenges, and open questions.ACM Computing Surveys, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Foundations & trends in multimodal machine learning: Principles, challenges, and open questions.ACM Computing Surveys, 2023

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.887313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.887313Z digest=sha256:b96b334942384c49db28615be5fe10e6adbddb596e13392b71542775f80c2339

Observation 6b6ca40b-4c35-436d-9848-aaaf7681f027 · outbound

This paper cites Hemm: Holistic evaluation of multimodal foundation models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Hemm: Holistic evaluation of multimodal foundation models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:09.983756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:09.983756Z digest=sha256:1397274f1d7b2c25a347891dc6dec8e4efa28f8b06cdcb0fdd5448514570f8e9

Observation c7a293f6-c050-41ea-937e-61a38f8899be · outbound

This paper cites Microsoft coco: Common objects in context.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Microsoft coco: Common objects in context

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.038516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.038516Z digest=sha256:ecf4caa9d8c5b4b02d8cc4a82c53b8a6326ee560782c724d7efa877555b02e57

Observation c5a4df77-f579-494e-a5d5-dd4f17f3f023 · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.106641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.106641Z digest=sha256:9c8631c670c3ee0d867038d09a171b42ed7cab90881874674e66274e51285977

Observation a369bfd1-54e3-47bf-b801-6809145c68af · outbound

This paper cites Visual Instruction Tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Visual Instruction Tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.199856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.199856Z digest=sha256:6435e745b1a2caa0bf7c12063c36f057277fed926e80888ddfb683385283eeae

Observation acb3b8f2-44d4-479b-81ce-0e829631ad69 · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.270829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.270829Z digest=sha256:32ac566d8aea3a9839a5cb912b9f921ec4faf83cb784c0b7004718e906a77387

Observation 24f00f28-3cf7-493e-b113-c6dd6326232f · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MMBench: Is Your Multi-modal Model an All-around Player?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.345380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.345380Z digest=sha256:b157dc26dc262353b0c6fbbc05c3a9703b6dc5a39e377da5c03844609b494837

Observation bc85cb85-975e-4319-8afc-377612f6b484 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping NVILA: Efficient Frontier Visual Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.410916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.410916Z digest=sha256:d8d0312740f97a82a5ab98da94f8b93bdc97435c06cb2ed5bbf9a0d2bb34634e

Observation 358804ef-e94f-4b19-b305-1c4f584a089b · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The flan collection: Designing data and methods for effective instruction tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.465143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.465143Z digest=sha256:922a65d7370e0bb03835fc12e278becba7a6ef668819d0e525d214681f8aa459

Observation 6138fea2-d956-455b-88ca-fe92967b3a2d · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.596512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.596512Z digest=sha256:3f942d3fa0df0f52f0e5cc492962c24f0fb3e1cb486aba3cb43009f4f3ed142f

Observation e22de362-4a5e-46c7-9914-ff50055fd6af · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.690316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.690316Z digest=sha256:816dacd802e05a040370bac129bb0e12e67a1140abefdae30e2880557ba21ac0

Observation e82fa458-0e23-448a-9146-d01f3efe5fc3 · outbound

This paper cites Modeling task relationships in multi-task learning with multi-gate mixture-of-experts.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Modeling task relationships in multi-task learning with multi-gate mixture-of-experts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.751352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.751352Z digest=sha256:57ca08dafd128515a7e7cfd0ebf9cb87777e57ddd881d7b095bbc4fedc74f9e5

Observation e4f8709e-9072-4134-ad22-b450a213f0ee · outbound

This paper cites A taxonomy of relationships between images and text.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A taxonomy of relationships between images and text

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.833909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.833909Z digest=sha256:f35f5584e2097d35182d0ad2a3825f47f291885b753dfdb5da81becd664408cb

Observation 3a885e73-e46e-44de-937d-a8f1211fb382 · outbound

This paper cites Multi-Task Learning as a Bargaining Game.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multi-Task Learning as a Bargaining Game

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.907531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.907531Z digest=sha256:1a36ecb01a9480b0b1dddfa399465bb5e10d151a2b4ce97e7e18b85f0ba32255

Observation 38a13155-3f7d-4e6a-9d2a-f4d4255497fd · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.004205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.004205Z digest=sha256:f5e57ea5496f6a557bbe04ae7f59649f5718bd9d82c20fd4e85f2f35013ee42b

Observation 50ca30bf-cb60-4064-8692-675a772afe83 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.089611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.089611Z digest=sha256:65a3ba4ab1a273f293d60f94c7c85a4eebb28556a609fbba7c053f80a5da1e97

Observation 589ed68d-a688-407b-97c8-126f0f7d2be0 · outbound

This paper cites To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.173110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.173110Z digest=sha256:6690f86af6d32be3771481dcd1b21d94797e01dea9d60295a614ff01aa87f618

Observation b16046a3-756e-449e-9dfc-c7732505f577 · outbound

This paper cites Connecting vi- sion and language with localized narratives.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Connecting vi- sion and language with localized narratives

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.335880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.335880Z digest=sha256:8ff0347a8b26d548446ceb14379aba5aa787d1f9178523bb9986cd84d6f41bc2

Observation fbd1ea9b-431c-42ee-8189-12cbe4615956 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Learning transferable visual models from natural language supervision

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.434785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.434785Z digest=sha256:ea21f6b31ee00c69265e37c7e3c6a8f6a7846f60781561c258f7e53286ebcdfa

Observation 5d51db8a-ffcd-49a5-9f94-83ecae176c6c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.524533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.524533Z digest=sha256:824d12bce2a8d85644c40f395d1c664216ab1dc7a6c450a14c2f8e6b12c03df7

Observation 905d2572-20f2-4d9f-bc59-472f188a60e4 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.619517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.619517Z digest=sha256:b050844a5e79df4e659718506a9f0972e875b7afde0832b1498bdef4d8ba4372

Observation 7c8118fb-a7cd-4d4f-b76a-d8fedd103c6c · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.690608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.690608Z digest=sha256:8d07bd7ebc63e9a75b96f4b6962b8d6bd35e72571d3945716835d21556890090

Observation 472364db-c53c-4439-b64d-349bfecdc59d · outbound

This paper cites an unresolved cited work.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.777240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.777240Z digest=sha256:54584863b870adab53cd35706acdce0774dd976696b1951d867523b2056cb658

Observation e61748c8-ef1e-46a9-b396-87264baf6307 · outbound

This paper cites MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.897837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.897837Z digest=sha256:5431ce5abf85399ba5db5cb1183ffb53d8eb51c316b74a7ac434698cbc36e453

Observation 55dbad53-eca4-4db8-b5f5-5471953cdc26 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:11.968805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:11.968805Z digest=sha256:9d98fe60aef0954641a15c9114afb0dcce1ae5ede07f6e647d77aee9e802a58f

Observation 474173e2-3c80-4f77-a82d-8fa0fbb37145 · outbound

This paper cites A Principled Approach for Learning Task Similarity in Multitask Learning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A Principled Approach for Learning Task Similarity in Multitask Learning

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:32:17.428702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:32:12.056052Z digest=sha256:430023cf0c32fb87e2096661fc58b7fdbf10b45c5b98b1f6682ba3c18cb82df8

Observation f507c611-7c9e-4302-aa20-d2f80dbf868e · outbound

This paper cites Guibas, Jitendra Malik, and Silvio Savarese.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Guibas, Jitendra Malik, and Silvio Savarese

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.174376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.174376Z digest=sha256:9c39c3135d0ce1d1580af36c45b620c87fffd93206dac7668ad1fd141d492b88

Observation ddd338b8-b98e-4bee-9eb0-11616d61f666 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.247002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.247002Z digest=sha256:3fddb034fef18b8c38c9c30bc9db29d1886fffa4db39d0dcb9696b5d8c5adf7a

Observation 0db03e46-08c1-4b7f-af66-5fd88c72a835 · outbound

This paper cites A corpus of natural language for visual reasoning.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping A corpus of natural language for visual reasoning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.321003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.321003Z digest=sha256:c397b020e99273b54ba66b42a1cad97bb7403ec5ab154a53cca8ebfff97484c4

Observation f4b9e7d4-e50b-4cd9-bb20-3f36bc1f00dd · outbound

This paper cites Multimodal transformer for unaligned multimodal language sequences.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal transformer for unaligned multimodal language sequences

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.460206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.460206Z digest=sha256:a1ffb429ca792dc8d088827400f5d5f582d413977e814b70192c02827d8ebf54

Observation 50ce583a-c8f9-4df7-b10a-bb0b97aa119c · outbound

This paper cites Multimodal few-shot learning with frozen language models.Advances in Neural Information Processing Systems, 34: 200–212, 2021.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Multimodal few-shot learning with frozen language models.Advances in Neural Information Processing Systems, 34: 200–212, 2021

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.554019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.554019Z digest=sha256:c536cecfa9fc14e95fdbd924b5d9250d448fcc0d041de06f80739a2a0b629565

Observation c0fd2c97-e468-422d-b884-af96401a9895 · outbound

This paper cites The inaturalist species classification and detection dataset-supplementary material.Reptilia, 32(400):1–3, 2017.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping The inaturalist species classification and detection dataset-supplementary material.Reptilia, 32(400):1–3, 2017

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.652313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.652313Z digest=sha256:04e06d06cd0182f4d670b29670e76d75ffd0b149362ff1edcb47fdbf1b2f12c3

Observation cc27e696-39cc-4348-8753-55e5d6d6395a · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.736468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.736468Z digest=sha256:6ba3272d77d01d6c371fb3b5af00337cc240564ac9a32cdb7fc1ce55e2388e4e

Observation a06b68e7-3638-4111-af79-e9f040f8cfa8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.860372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.860372Z digest=sha256:b9ffbef9f7bac1539c5da9c9d5e1643ae90ce631ab393e26d2f232e555a9cbff

Observation c74420d2-fe47-45a7-8127-cbf5d76b1431 · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:12.947608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:12.947608Z digest=sha256:98d9d665e6f415279bb8c5568df6fc9495d05ec0ee2b16adb077f2b26f6de62a

Observation 5b734aa5-7731-4c40-9e97-9d368c1ba4cd · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Finetuned Language Models Are Zero-Shot Learners

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.042859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.042859Z digest=sha256:11d412097bb7f11d7d630c5b899301b2015ee2126dfd738c58d842e887f20498

Observation d905f998-4953-4734-96dc-1096ee352a4a · outbound

This paper cites On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving, 2023.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping On the road with gpt-4v(ision): Early explorations of visual-language model on autonomous driving, 2023

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.109666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.109666Z digest=sha256:fb337697e44471fd4ea7940dc297744e4b19d0ed62062df0dae213445375df2c

Observation ea17165a-e1cd-4a8a-955a-ef3e918e3154 · outbound

This paper cites Nonnegative Decomposition of Multivariate Information.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Nonnegative Decomposition of Multivariate Information

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.263231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.263231Z digest=sha256:16a141177fb3da212f8a30e097ba413de9dac3f9acc982886ca2c300e64cb133

Observation c984e5dc-6606-476f-be84-b2f351141a27 · outbound

This paper cites Vision-Language Dataset Distillation.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping Vision-Language Dataset Distillation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.340080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.340080Z digest=sha256:9c697300b732b22f748410f437a5e7051f021f5632ec42e811f81a2dbcb0fd34

Observation 3d9636fd-d20f-46ff-aef4-26bb12c40bef · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:13.382009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:13.382009Z digest=sha256:9d511195d21145819eea5d8436b844ebc9911605ec6c23b42594fe802fafe9bf

Pith citing papers

No inbound Pith citation observations are available.