Pith. sign in

Paper Citation Record · LEDGER

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning

As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2507.21924.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21924 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:41.518126Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved38
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5478395a-a3e2-4910-ab88-b83667be0c36 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.323457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.323457Z digest=sha256:f8d0945fb5308c1cba573ecf790c7236f059ad516066c32a40c1a3024c3cd1f6

Observation 0462b592-f0b1-40a1-ab79-724f553592da · outbound

This paper cites Qwen2.5-VL Technical Report.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.561550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.561550Z digest=sha256:14855f339de19fa63d571b111b26d09e36db17aaf248e0df86ad87e4082d6a6d

Observation d43985dd-d259-4ece-bddd-04e480c521b6 · outbound

This paper cites Jawahar, Ernest Valveny, and Dimos- thenis Karatzas.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Jawahar, Ernest Valveny, and Dimos- thenis Karatzas

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.688929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.688929Z digest=sha256:440784972fd8ca06e9676515293261373af12f77382cbb79c87e80833e05ef05

Observation 5e75490b-058d-4a84-81f8-52a64d95e4fb · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text en- coding.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning An augmented benchmark dataset for geometric question answering through dual parallel text en- coding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.821625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.821625Z digest=sha256:1f7dff9c748316599f1305254b3b28900449ecedc1dcff15bdd307367b7fd9d2

Observation 3baeeb8c-e4ef-4928-8cff-4e32bae54d62 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning FireAct: Toward Language Agent Fine-tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.887391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.887391Z digest=sha256:bbe8b1a10a941400984558fb59412df2aba9cba4337335a32b2923e225d33b27

Observation 463991f0-cc00-40b0-8798-d05ee1827c01 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.962339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.962339Z digest=sha256:a9c62cd994aaf4aa4cea8cc7cf27353e0de8e125d4b65adda0392bdb4008723c

Observation 6705916d-5dbc-4a8f-8611-f80eeaf2b8f1 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.056475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.056475Z digest=sha256:84d0f1baaa392dbdab617a92b6d04cf9bf0ebeb536b44f145a988b585eb1a6fc

Observation c5d95906-1dc0-47af-9f83-a6d76b8c75b1 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.202517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.202517Z digest=sha256:26e47c3a29fff510bd9ee6d86f04b9dc6a1d6635b27c2bdb93107a75e8cdaea4

Observation b2b48a63-6335-45e0-b0b6-35b2fc97d4c6 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.356535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.356535Z digest=sha256:bea013f71167c06a5e25f58b5080dd114947695f8ae66620c7c2464bdb5c5158

Observation b41b9cc9-cd38-4c58-9845-d34dd9504529 · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.447571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.447571Z digest=sha256:a0ef64a786a84b56f24b491133db114202f8b980aaeba921e091c46a4174adbd

Observation 92c9f48c-c1af-4cba-885d-a314261a25bf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.551884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.551884Z digest=sha256:50e2702de9bbb8214ac5e63dc6b651c4e0393968f7392ae40988f860ec4330f8

Observation 15d9bcc5-8cbf-46d3-96b5-c0b01a6075ea · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.676917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.676917Z digest=sha256:50e69730ba3158ec78eb53583061b934a10bf84c69eeedd88b6d9ceb02b43c0a

Observation a80648dd-7b1d-458c-ab83-446322737701 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.810532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.810532Z digest=sha256:ef4da0dbda990c8c0113f01c514ca1699d0078fee09ee6a1c53fe4614c3de9f6

Observation 971d3421-3b28-423f-9ecc-977122187998 · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Clevr-math: A dataset for compositional language, visual and mathematical reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.899714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.899714Z digest=sha256:46672d699d7a3a7c651c36c1ff81eabf6484811bf848761eb454e767a496f7ba

Observation 4f5df2a3-68f0-4fb0-9d97-992800e1cd89 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.133368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.133368Z digest=sha256:6999aff39e6f501a18a809f7975a9ad078a0a54b85f6e1de74e6e85f3d178725

Observation b5973dad-5e46-4e7a-9168-241b722a9d96 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:53.445321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:33.235402Z digest=sha256:0ea9fdff7dabaec83e0c88457ad6a920645a8c300effaa2af95af3bf0a6ed4d6

Observation b764f13b-cc2a-4907-983f-74e1c22f8fe0 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.353518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.353518Z digest=sha256:c53cb6da9393002050ee77d215eab2baaff65bc012c32bafce8daceb184e92ff

Observation 0b4a3181-5f0a-43df-9a38-1b3f554a584c · outbound

This paper cites Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:53.080042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:33.503296Z digest=sha256:31fb0b0ad412be70f42b345c372e1143c084eb816ba23f5f4d0e5885e078b999

Observation e63cb2f6-1627-44cd-8a77-77aa298bd6b9 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.871053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:33.635653Z digest=sha256:f3ec130cdac9264112dc7e1a1919be92237f2c750dfbe35e140b95788020dd03

Observation 1d71c6f4-9663-4f51-bedf-903cc3f88412 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lora: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.799912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.799912Z digest=sha256:d897b34d9119e13d55a0f08540854722574f2de31aa0f485e6248afa4cb39cf3

Observation eae346d7-dea3-4869-a0fd-8ff980354b34 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.676694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:33.926347Z digest=sha256:6506f4967b6292d166f4c6718522bfbd162a74661f668597e668b7057505e579

Observation 3877697a-21a2-4942-8ffc-e5c49755e270 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.425932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.070741Z digest=sha256:e67c46c26e2b98602c3bb8f7f41bf2ad815dcb16751d0f7cc86643a38b021403

Observation 7624ac31-6ceb-4462-b1f2-92689753961f · outbound

This paper cites GPT-4o System Card.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:34.176052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:34.176052Z digest=sha256:5042fb86f1c4b0fc85d9220ab50b70cbe5705cd4649960e6c54c0a912c21e662

Observation e6c32f4b-80b9-4a92-ba9d-501d5d813541 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lawrence Zitnick, and Ross Girshick

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.105623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.324440Z digest=sha256:a0c38735f1f7a785796c7482f725487bdae47fe8024de92b076f04d7b0155a4a

Observation 3cde0bb2-3704-4693-af26-65dc37e6cc7b · outbound

This paper cites A diagram is worth a dozen images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A diagram is worth a dozen images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:51.795759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.473275Z digest=sha256:5a5a7763190d675876abb2ce1c834b971cdb39d16b65fb0e7775d66a44bfeed9

Observation f257a9f9-e7c1-4d5c-a8b0-4517f4a3ed53 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:51.269990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.736523Z digest=sha256:7cb7361cc7d3f2e1fd98a54e151d4fa77f7076250f1d7ebe216cc11b3d131038

Observation 71b10f2c-bc23-4add-b219-882ce0430e63 · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.997805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.828859Z digest=sha256:1128edace2a9a68438d7aa85721609e57c1a15154cd32a53390057af3ec876d2

Observation 126b78f9-f3f7-4151-8c5e-c6987d0b1265 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.784892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.949585Z digest=sha256:9224d0bad0dfb01709c827c2125deb26cf18ba9f96d9aff90ee4fb7fe0925a7b

Observation cb91d659-f428-42cd-9432-dd29556578f5 · outbound

This paper cites What matters when building vision-language models?.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning What matters when building vision-language models?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.087650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.087650Z digest=sha256:34c4e3685190f12da886effdd14e5662fd2bebef767ebd41bdbdbf42ee4a04fa

Observation 32ee5be9-6523-403c-8df5-f6171b4f14aa · outbound

This paper cites Kankanhalli.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Kankanhalli

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.510363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:35.219442Z digest=sha256:0c498e167d8effd34ca92c146baea1417001f8502c482a0d7ef91ce33a0c34cb

Observation bf4c79f0-08e5-4ae4-9fb8-731fff0b5e2b · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.368071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.368071Z digest=sha256:b15d5f02531bc7817dd426ab6f4ef09a174924bce4e8300bf2bee9abce6682fe

Observation e6dbb72a-e8be-4707-98fa-1a65ca8b36d9 · outbound

This paper cites Visual spatial reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual spatial reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.268677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:35.520950Z digest=sha256:c081353a0f4390cb72860d0d673a63a39fa235277010d33615818215a33562ff

Observation 150b32c7-24de-4cd7-a868-11d75292721e · outbound

This paper cites Visual instruction tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.688521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.688521Z digest=sha256:10f11d1399fda6b9687801bb9a0f03c282a0e2acfffc6a9fdada9ba0523b0867

Observation c24c4f2e-b304-45f7-a192-054771395861 · outbound

This paper cites Improved baselines with visual instruction tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.794468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.794468Z digest=sha256:9e7aa4d08134533ddabc96c45ca16da83f7cf1d4587643182cb2f7dba3e5a03a

Observation da4acb6d-e861-4fe5-bc60-f8c68030414d · outbound

This paper cites Llava-plus: Learning to use tools for creating multi- modal agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-plus: Learning to use tools for creating multi- modal agents

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.028090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:35.884051Z digest=sha256:cbe22cf9a79c767e5514061591ed6d36c3aed7322b480d9d67cf149e7275ca24

Observation 3c4785de-f488-46e6-894f-42253a8b2587 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.759709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.050970Z digest=sha256:2a21643ee0f89c74a19a7f7195c29d17e461756107d234dfd1efeecfd65ef6b6

Observation 17eb8dc7-fd91-4c74-a714-e93a0aae3cc0 · outbound

This paper cites On the hidden mystery of ocr in large multimodal models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning On the hidden mystery of ocr in large multimodal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.507460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.188259Z digest=sha256:d6d1d568d652c1342e8292dd5162141da79b0419edc7f730ee8119bdf3f3a55d

Observation e8520f35-a2c8-4be5-be7d-da8fdc5baeea · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.280302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.349318Z digest=sha256:6007707144b76f6e09422f2af327e7a810f22fae8e6807b3da1f6b0bbc872d4e

Observation eec69d49-8b0a-48c7-834b-7b86943d7753 · outbound

This paper cites Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.930150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.452978Z digest=sha256:6c749b2574c0de250430bd51367d55c6cdd7fd88dc0bf129b53d6e5af6176211

Observation ce7ff31f-0d94-4bb9-8493-408472c1e8f2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.664559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.570393Z digest=sha256:8f68685b0734496414242825d046d8f6844f9c2338942a1b310fd17b17f2d1d5

Observation 319aec87-0af3-4f86-bfe8-be9ab2818488 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:36.719126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:36.719126Z digest=sha256:0916ee5b73ea04e12d9e6adc6b535f10e2c2980d1be1e45c0b4a9bbe2e00aa3d

Observation 2ab389bd-043c-404e-8d5f-fce2f8f5d213 · outbound

This paper cites Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.351154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.825822Z digest=sha256:eb0d1a6a6d8342aa11f8ccacb99c3e88376e476f629d502768c558618287e7a6

Observation 0780e952-e4e3-4756-b3fb-50787697b212 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.243615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:36.926908Z digest=sha256:1514f3abfc89c23f5b96f120848cdb138f90024bb35555716dbd3e7e4d7627d8

Observation 6f07a7d3-977d-4b42-a5e1-42ec6e1e802e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Docvqa: A dataset for vqa on document images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.168702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.040447Z digest=sha256:1e5f203be30b1d9319236fc396f87c4be4a8089042d6fcdb1cbe775f4cb4a92e

Observation 1c296ffe-614f-4b16-b661-1b1a212f80d4 · outbound

This paper cites Infographicvqa.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Infographicvqa

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.004564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.177685Z digest=sha256:e6a214213bdd3389f223d72192e31471ed305c5e3f1298e382464857f2209f6c

Observation 6ae9bfaa-2a9c-4804-aa2a-dc2b857795ce · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.920183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.259659Z digest=sha256:556605abfa17ef4dc90da5356ddf7a0b9dfc81e90425cab187b25e5634c2c976

Observation 71d02cb0-f2c3-4195-be9c-e0e90823e35f · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional chain-of-thought prompting for large multimodal models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.797006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.377593Z digest=sha256:f184d283b28af0b6c527d6df910d8eba0eb88ed62dd57b9f78f81ebafef506df

Observation 89d54ed3-7d75-46e4-9c19-469c7336160e · outbound

This paper cites Compositional semantic parsing on semi-structured tables.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional semantic parsing on semi-structured tables

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.676234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.497595Z digest=sha256:ec89a210aab215d8c6d6219854a8f339222b43efb3c97aaf03a910072865202f

Observation 53b939d8-6681-442e-976a-2f9777000580 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.533383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.575202Z digest=sha256:8876ac27f9e718e6d2ff80c0a73d1558d51d38606ad97f09d77a6f7b9a8d3a91

Observation 67ba3997-6b2a-4bf6-96b9-68bb29ab71f5 · outbound

This paper cites A benchmark of facial recognition pipelines and co-usability performances of mod- ules.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A benchmark of facial recognition pipelines and co-usability performances of mod- ules

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.367570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.702678Z digest=sha256:d02f5b044f1d603b505294575fcde9eb8f67b63d738d7f8580c931b5bd57f9a7

Observation 65e6ac54-3195-45f1-bb0a-adf7d0b85a17 · outbound

This paper cites Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.277223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.818888Z digest=sha256:b6b6a44bff133d5cb6dcf2b4b36f9342dfaf3e163f12df0880b2fa914234b557

Observation 1494792a-ecc3-4427-8057-b946dd8698eb · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.182352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:37.933670Z digest=sha256:2dd3a2c0f33b362f7ebf560d653f3b9e9ca149b348be009bca881970ffe2594a

Observation be5ce8d3-57c9-458f-9ab8-cb9c930acd9a · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.079081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.079081Z digest=sha256:3795924111cb024e5030197214b1b097ba87f065fcaaa195209a4a9d04e3172f

Observation d7963086-d156-4656-8cc0-45f01486735e · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Textcaps: a dataset for image caption- ing with reading comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.200474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.200474Z digest=sha256:0a02984a40fae80280a528db495cf24891ccdfb38c1699d2864e52679e9fbffa

Observation 031f965e-c81e-47ce-990d-88f3dc516a55 · outbound

This paper cites Towards vqa models that can read.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.982901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:38.315186Z digest=sha256:b4244f5537552c7b40c38feb4882085376be001e5aef425f9430bd5a959e3f64

Observation 957c6707-7fcb-4195-84a1-e72f0b84820a · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.403343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.403343Z digest=sha256:8380f967d9e34bcad10bbdd58ac7a84927e614612b9cd36a7230367706bc4f6e

Observation cc460c53-afcf-45cc-bb4e-231079fc7c88 · outbound

This paper cites Tang, Angie Boggust, and Arvind Satyanarayan.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Tang, Angie Boggust, and Arvind Satyanarayan

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.756525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:38.500912Z digest=sha256:f9c4e19d7b1e92d4924de4d74ba28609495ab221f36e1e6073171f4af15e906f

Observation e04a0778-012f-48fb-91de-e1f8e17e9509 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.622584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.622584Z digest=sha256:0c48176ec31cbd5d654fbe4e473e734c1dd937aa6df44a6f367bd85a11639d72

Observation 771ed215-2e01-4c95-beb8-e346183c90ee · outbound

This paper cites Document understanding dataset and evaluation (dude).

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Document understanding dataset and evaluation (dude)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.522644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:38.703728Z digest=sha256:56d18fcf06a56dfe2db044e60659fdd5a00e49c1cdab187c50601d7698da2956

Observation 8a635456-6fae-42b1-b5e6-1c6d2232609a · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The caltech-ucsd birds-200-2011 dataset

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.288111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:38.811721Z digest=sha256:e43fcfbf0c8542d01977311caddf8667286ea6ad9637a77ba46f74fa1aaceae8

Observation 17cea784-ef6c-4838-ba07-805935d5b36f · outbound

This paper cites Screen2words: Automatic mobile ui summarization with multimodal learning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Screen2words: Automatic mobile ui summarization with multimodal learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.064600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:38.899604Z digest=sha256:cdecf3313378e836c35397c727297bc3db07809a90a607de45012d2d2bd0381f

Observation 3d2fe1c4-0a8c-4ee2-bfd4-ebe8f53cfd67 · outbound

This paper cites LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:18:41.874604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:39.029360Z digest=sha256:904df32e0ee717e19004db73bc95723acf89c464fa82911f090d6761e8a3d078

Observation 7509c407-ae30-4486-8ab7-12cd21de8e94 · outbound

This paper cites MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.118189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.118189Z digest=sha256:da052d9d7f621442cf9eb05baf1be90f604539416768bc3b71ad55a5f18e57ce

Observation 29515936-2adf-4b96-8dcd-d245ff7c8558 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.869578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:39.207676Z digest=sha256:d8c3de5cdf2e5d3b4ebe1f6a461f371f02312bad313d0250f584bc81abd7493c

Observation 5e426e2b-2a9f-4ee5-b01b-71570168c1a5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.306583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.306583Z digest=sha256:864eac20eaf1be592ddc752a0631544cc232c42a18b9b60d30a2c2d92305083a

Observation b2630173-c5e9-4b66-b4b2-b531121a1309 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.399123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.399123Z digest=sha256:43b53cfb0f1049a2dd5d2eb3c01e05ca3d99d889723ca643168d2ea8318d3fa5

Observation 4197aa3f-1e21-4c07-b7c3-0bd0709b8881 · outbound

This paper cites Grok-1.5 vision preview.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grok-1.5 vision preview

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.603526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:39.504236Z digest=sha256:0fc0e37614cd3d7ffa954335276597050c1ef90ba0e62ebef0a1afbdf5659896

Observation 744db546-4e65-46b4-96d3-0275004cc8e8 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.605249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.605249Z digest=sha256:0fe7f7c10241c8a9be5a72cd9fae7c22697683ca12e0bd44edc9690155f88d6d

Observation 6af0f074-14a1-49a9-a077-7722867c6ca3 · outbound

This paper cites Llava-cot: Let vision language models reason step- by-step, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-cot: Let vision language models reason step- by-step, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.387347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:39.749545Z digest=sha256:1a27df6d686e49fd2fd512965897a65e9ec221ff1c94419a2f739a21416c1e6b

Observation 6f935308-2bd6-42db-a8bb-dbd13a588ae8 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gpt4tools: Teaching large language model to use tools via self-instruction

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.106833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:39.829944Z digest=sha256:19126495d8e8f420f1068f5db00eb453519a8d2ea2a981d256672b4f44b3c3f3

Observation fc6a8b88-5b9e-4155-82de-657a638a8d80 · outbound

This paper cites React: Synergizing rea- soning and acting in language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning React: Synergizing rea- soning and acting in language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.865434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:39.900854Z digest=sha256:880e4f79084fe14edb48ab609c9b836be4b70ab86e9accb343da4def4e719a34

Observation 6aed4166-fd33-4939-ae15-4e8f720b5c6a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.013192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.013192Z digest=sha256:389dda9f942cf56df1372d2287affb76489db7a3d9ddc9c77414a37bc290a760

Observation bed3ff95-f0d1-42b0-b8b5-0604add61aee · outbound

This paper cites Agent lumos: Unified and modular training for open-source language agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent lumos: Unified and modular training for open-source language agents

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.640480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:40.111923Z digest=sha256:25217a360c3ef831c0bf1668f4cf423325f0075fc2395249280351626a7c93b9

Observation 09ac949a-d7ff-4981-a42b-e5acb84a3f31 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.437025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:40.209708Z digest=sha256:5699ac843fe0c195342527f6eff3f9bde2ed55ca49681b5b966f83b6cc6c8d0e

Observation 565e3f9f-ecd4-4be0-a9e7-b16f64f828c3 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.313560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.313560Z digest=sha256:ad9ba2cbbdb8e255093eae0f1cba9095b3177fc846c7fe40a871dd034508dbd6

Observation 0967ca33-3478-437a-911a-75eb29b5b5d5 · outbound

This paper cites Raven: A dataset for relational and analogical visual reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Raven: A dataset for relational and analogical visual reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.189547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:40.418499Z digest=sha256:62569f4b13cb6b08d39116bd248d1cfe47cb2582d2607e0465e239e27eb64ac9

Observation 0cc7d3f4-c98e-406c-8321-11f839aad7c7 · outbound

This paper cites Swift:a scal- able lightweight infrastructure for fine-tuning, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Swift:a scal- able lightweight infrastructure for fine-tuning, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.965856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:40.563398Z digest=sha256:f726d4bd491b4e73ef708367fe9e7f8326cdb428c43a0b07b8627d3065728a85

Observation 4411d2fb-c374-493c-8caa-d4770086b41b · outbound

This paper cites Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.750174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:40.636458Z digest=sha256:0819df567e05d7070431f71c924db990e01c288baf94dec7dcb61fc1f0e30edd

Observation de5ef537-70bf-40f3-9766-c216ab3a5fd5 · outbound

This paper cites Visual7w: Grounded question answering in images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual7w: Grounded question answering in images

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T12:18:43.524971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:40.695193Z digest=sha256:224dc4dce235c7b869be6986dce5efce2d6c1e4a642a68b2e2975732db173ce2

Observation 5256053a-33dd-46de-8251-152ba8f8af31 · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.808887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.808887Z digest=sha256:a825d7f40a2c77d61a20a705b38c8e860cd6f70d80f67e624aca5aefbd703d2c

Observation c81563ba-84aa-460c-a47c-c3f38c58dd28 · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.922255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.922255Z digest=sha256:72df6f70891a96a6fb7292828bdc49fdc90551cd1ee75bf5edc74cba3cc244f3

Observation 8c1b6981-1d26-4138-b7c8-99894dea8acf · outbound

This paper cites objects": [.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning objects": [

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.246936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:41.042512Z digest=sha256:d07393ceb1c2b612cf3201fa6fc6d8178c26e58d587e9284e99dcc2d2209b081

Observation 05128088-b15e-47bb-ac2d-8c23e84e2e5c · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:18:42.999932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:41.141614Z digest=sha256:03fe52c36af62ce62f62d785820c9c5ba5ded3c9d90b832b3f9a683b982f30d6

Observation 3f7bae2a-c3bb-48b8-98f5-9085610ac3ff · outbound

This paper cites image_caption.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning image_caption

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.706027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:41.259999Z digest=sha256:a7229a3fdf2e6697e2cb3d6b7213bb571210e3481fb4567ed7a6869a67a3718c

Observation d5f08797-6684-43ed-88c0-c5610b21feb7 · outbound

This paper cites needed": true,.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning needed": true,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.471282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:41.390516Z digest=sha256:c1e64b8c35860e69ef6f23a855b4db3f3956dde6767099afe9a555a36645dd1e

Observation 72e7adf1-f915-45ba-b014-9a7760bf60b4 · outbound

This paper cites continue.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning continue

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.265587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:41.518126Z digest=sha256:46c673e6d3335d22aece14bcde5322712ae643e080f7893b840c4a4ae56c04eb

Observation d63ce0b8-1ee8-4fc6-a71d-aea9de96110b · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.027058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.027058Z digest=sha256:c47b2d6119e03b63b76332b52bab641a7bf8804e4d33fdf7ff95d4d47a199c0d

Observation cf2204af-5146-43f9-9318-0ccc42de42dc · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T12:18:51.528283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:18:34.612419Z digest=sha256:ea488b1ae0c47ce0f794fbcd4db96fddb02f2d7fb6d6e8bfd91dcbd717b09aa0

Observation 846a591b-e1c1-4d71-bf21-87c1e714b30f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.416961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.416961Z digest=sha256:c2c7f7b6f1f5379bca78813d389b6f8936f7fe7872c42b50cca8cf4c94b763c2

Pith citing papers

No inbound Pith citation observations are available.