Pith. sign in

Paper Citation Record · LEDGER

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning

As of 19 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2507.21924.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21924 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:18:41.518126Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved38
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5478395a-a3e2-4910-ab88-b83667be0c36 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.323457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.323457Z digest=sha256:df3825cdca1d68511edde35d58c9dd1967024e455212d5112d12c742353a0273

Observation 0462b592-f0b1-40a1-ab79-724f553592da · outbound

This paper cites Qwen2.5-VL Technical Report.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.561550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.561550Z digest=sha256:c9ec6549a08c047faf26d4285c168e01a79f64708500d2ec72e90c54b2e3dce4

Observation d43985dd-d259-4ece-bddd-04e480c521b6 · outbound

This paper cites Jawahar, Ernest Valveny, and Dimos- thenis Karatzas.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Jawahar, Ernest Valveny, and Dimos- thenis Karatzas

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.688929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.688929Z digest=sha256:a5f85a876fe8d5ab633a95ed9690b492e610ed45bd32a52da485b4d28dd6ffbe

Observation 5e75490b-058d-4a84-81f8-52a64d95e4fb · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text en- coding.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning An augmented benchmark dataset for geometric question answering through dual parallel text en- coding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.821625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.821625Z digest=sha256:98f2dbf79633ae361586d0875a9b092aa7f6316af2c6e72b1a9d99d58f9b623d

Observation 3baeeb8c-e4ef-4928-8cff-4e32bae54d62 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning FireAct: Toward Language Agent Fine-tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.887391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.887391Z digest=sha256:2405fac82f85c0eed8f22094bb03b3516d017defa36c116f3ecae9967e1530b1

Observation 463991f0-cc00-40b0-8798-d05ee1827c01 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.962339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.962339Z digest=sha256:51bed020edea3417c5d9565b8f18628aefc16a5743eaf427b1419dc7d6b5e3ed

Observation 6705916d-5dbc-4a8f-8611-f80eeaf2b8f1 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are we on the right way for evaluating large vision-language models? In The Thirty-eighth Annual Con- ference on Neural Information Processing Systems, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.056475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.056475Z digest=sha256:561a0fa85f17bf4961bbc8e80d4ff551904bb89338bd23460d8bbbd77dcfad3d

Observation c5d95906-1dc0-47af-9f83-a6d76b8c75b1 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.202517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.202517Z digest=sha256:26f744442da5cc3133d9e74d1729686e6256ca03e143c7976b3e4d6c34239756

Observation b2b48a63-6335-45e0-b0b6-35b2fc97d4c6 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.356535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.356535Z digest=sha256:af8b992a666a4086a925566a2b29c2554d945bc48275adc2d34e0eb223498d48

Observation b41b9cc9-cd38-4c58-9845-d34dd9504529 · outbound

This paper cites Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.447571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.447571Z digest=sha256:462a7f5469a92fed8ec2838454ab2aea2fdb054cb097df958fe0befb38694b4c

Observation 92c9f48c-c1af-4cba-885d-a314261a25bf · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.551884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.551884Z digest=sha256:b78db8f7ed9c5afde1ca6c5a4d7d9f823f00de6a6c5e5db608a28867a59765da

Observation 15d9bcc5-8cbf-46d3-96b5-c0b01a6075ea · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.676917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.676917Z digest=sha256:4f4c214984ef02ea02fd92a9e2301207843a395d7a368bcd8010b1112a439a2f

Observation a80648dd-7b1d-458c-ab83-446322737701 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.810532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.810532Z digest=sha256:2f984da267dfde6077c8d17e35c3839d3a51c038157a2fd1dacb3def53e7836d

Observation 971d3421-3b28-423f-9ecc-977122187998 · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Clevr-math: A dataset for compositional language, visual and mathematical reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.899714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.899714Z digest=sha256:2f143ad982380657979aa1622f9fae97d5411de8387a1cf6c0bb0c73ecf4de9c

Observation 4f5df2a3-68f0-4fb0-9d97-992800e1cd89 · outbound

This paper cites PP-OCR: A Practical Ultra Lightweight OCR System.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning PP-OCR: A Practical Ultra Lightweight OCR System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.133368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.133368Z digest=sha256:906612e35cba5aeac6217e738ef0f51fdaa75fd014ff17d827bd62ce42f9aff7

Observation b5973dad-5e46-4e7a-9168-241b722a9d96 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:53.445321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:33.235402Z digest=sha256:b80be10eb74af67edb6226282cbda00725ad07f63e24305d7be58dfcc9129c6d

Observation b764f13b-cc2a-4907-983f-74e1c22f8fe0 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.353518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.353518Z digest=sha256:4ec9dc4e6fe30df9254569225525ed832e672ea04d5f25dd472c494d44b17d0b

Observation 0b4a3181-5f0a-43df-9a38-1b3f554a584c · outbound

This paper cites Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Multi-modal agent tuning: Building a vlm-driven agent for efficient tool usage

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:53.080042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:33.503296Z digest=sha256:6c3d4eaaeab84665b2bbe71cc9299f85890c5b467b8a46746dc26bfd5c8218f5

Observation e63cb2f6-1627-44cd-8a77-77aa298bd6b9 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.871053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:33.635653Z digest=sha256:e900be6ece6c0ba7f0b5e789080bb1ca727defbaccfe18bd8ebafa5c00246b40

Observation 1d71c6f4-9663-4f51-bedf-903cc3f88412 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lora: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.799912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.799912Z digest=sha256:754b5d4e4d596a7c1ca003472e5815770c76432ce8278333ddd121c0f3f5c645

Observation eae346d7-dea3-4869-a0fd-8ff980354b34 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.676694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:33.926347Z digest=sha256:3f5ddfbacd2c30e7ae19ffce700460a3d610647e79eefb18aec27a8ac2780748

Observation 3877697a-21a2-4942-8ffc-e5c49755e270 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gqa: A new dataset for real-world visual reasoning and compositional 9 question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.425932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.070741Z digest=sha256:62cde6d7531989241a29a3c0f955c91ed9026cf0bbb2e5e2b3a32b4a0c7d77dd

Observation 7624ac31-6ceb-4462-b1f2-92689753961f · outbound

This paper cites GPT-4o System Card.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:34.176052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:34.176052Z digest=sha256:3353d71101cdd51ab1d34876f45dba5eecdf943f2f3c43003f2244b214c8a00f

Observation e6c32f4b-80b9-4a92-ba9d-501d5d813541 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Lawrence Zitnick, and Ross Girshick

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:52.105623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.324440Z digest=sha256:89718767359aa2e8e23d616ac7b967faab91c2b33a11239636e81559c32c82e4

Observation 3cde0bb2-3704-4693-af26-65dc37e6cc7b · outbound

This paper cites A diagram is worth a dozen images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A diagram is worth a dozen images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:51.795759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.473275Z digest=sha256:0076c1fd85693f51c3a14600fd980557a0f16ce9fc722c43d8bbfab189391898

Observation f257a9f9-e7c1-4d5c-a8b0-4517f4a3ed53 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:51.269990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.736523Z digest=sha256:e4a06d7d5a0eef2aa8020fb0e67a9c07ffd2e125a32a8a02ef74a7cab519044d

Observation 71b10f2c-bc23-4add-b219-882ce0430e63 · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.997805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.828859Z digest=sha256:6ab386a7b4722174b7abc29ec23a84128e75841c14a3fc043166b8c94ad3441a

Observation 126b78f9-f3f7-4151-8c5e-c6987d0b1265 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.784892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.949585Z digest=sha256:80fc2d8703a0ff3eee2b603fbc03b1c32998d2cd906549e3a7517c594ea638eb

Observation cb91d659-f428-42cd-9432-dd29556578f5 · outbound

This paper cites What matters when building vision-language models?.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning What matters when building vision-language models?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.087650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.087650Z digest=sha256:f86e5e13475f39164b7028d1bd07d5b90e3ce751e059bdf3ec15c71f6a6db515

Observation 32ee5be9-6523-403c-8df5-f6171b4f14aa · outbound

This paper cites Kankanhalli.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Kankanhalli

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.510363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:35.219442Z digest=sha256:0915722d20068201a9c67a58d37c5473b0c82a0f31e1fa6844cfa3ac56e1855c

Observation bf4c79f0-08e5-4ae4-9fb8-731fff0b5e2b · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.368071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.368071Z digest=sha256:5ffb0ac04dc4e9665c2271669139e308c6e44229a8ac37bce39aae82085a324f

Observation e6dbb72a-e8be-4707-98fa-1a65ca8b36d9 · outbound

This paper cites Visual spatial reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual spatial reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.268677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:35.520950Z digest=sha256:271de033746205a46b25c0073e0e44180bb3ed0f28e82b183603d3449e333545

Observation 150b32c7-24de-4cd7-a868-11d75292721e · outbound

This paper cites Visual instruction tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual instruction tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.688521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.688521Z digest=sha256:8138bdca221eb6dcd15ac149c22f0278af1488eb80ec58b0306cdfc77b429d10

Observation c24c4f2e-b304-45f7-a192-054771395861 · outbound

This paper cites Improved baselines with visual instruction tuning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:35.794468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:35.794468Z digest=sha256:f05574d90420e1c4cd8e7911c9b2a8e524fe0157c2189ee6ad6ef7d1baf61508

Observation da4acb6d-e861-4fe5-bc60-f8c68030414d · outbound

This paper cites Llava-plus: Learning to use tools for creating multi- modal agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-plus: Learning to use tools for creating multi- modal agents

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:50.028090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:35.884051Z digest=sha256:9ab6e981059f2b0fceb3c0b8909255aeb72d101a2a9cb9e5676b04ed319b3acb

Observation 3c4785de-f488-46e6-894f-42253a8b2587 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.759709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.050970Z digest=sha256:adb2a3315e593f995653895bd3d90e5ea57f56f74c3c4a2bdcf13719594444de

Observation 17eb8dc7-fd91-4c74-a714-e93a0aae3cc0 · outbound

This paper cites On the hidden mystery of ocr in large multimodal models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning On the hidden mystery of ocr in large multimodal models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.507460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.188259Z digest=sha256:33b48b3f54f366c6a1daa7ac312807a1f25dfaa53bdfa589ad6dcd0f482b9d2d

Observation e8520f35-a2c8-4be5-be7d-da8fdc5baeea · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Inter-gps: Interpretable geometry problem solving with formal language and sym- bolic reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:49.280302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.349318Z digest=sha256:e0d08b440a86fb013ecc640702b53aba87175729819316830d305b14a19a5d80

Observation eec69d49-8b0a-48c7-834b-7b86943d7753 · outbound

This paper cites Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.930150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.452978Z digest=sha256:6a9cca661ceec773ca7bcab0933049de0f54758d6c48d3610a82b1ee6ed92815

Observation ce7ff31f-0d94-4bb9-8493-408472c1e8f2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.664559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.570393Z digest=sha256:8431a55d46b3f699de5c8b29c3988cb292d49ca6c431b7ee2b20894f2043a2c6

Observation 319aec87-0af3-4f86-bfe8-be9ab2818488 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:36.719126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:36.719126Z digest=sha256:c549126adf6a72fbef4abfba3d41e6e6e00abdd8082e481909cc28e37576c0c8

Observation 2ab389bd-043c-404e-8d5f-fce2f8f5d213 · outbound

This paper cites Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mathvista: Evaluating mathe- matical reasoning of foundation models in visual contexts

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.351154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.825822Z digest=sha256:a6a43e171800e2aad01dd1b7aee9397dd7549d4c119682e61733d5655770f799

Observation 0780e952-e4e3-4756-b3fb-50787697b212 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.243615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:36.926908Z digest=sha256:0caf7f0efe29aece7b07006c53da9f740fa30df298237136f4272d5030c50411

Observation 6f07a7d3-977d-4b42-a5e1-42ec6e1e802e · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Docvqa: A dataset for vqa on document images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.168702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.040447Z digest=sha256:637cb02c875a64bdbfa87a6075a66cd90ffcc14fb47b3e5f2052a54708f0dd44

Observation 1c296ffe-614f-4b16-b661-1b1a212f80d4 · outbound

This paper cites Infographicvqa.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Infographicvqa

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:48.004564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.177685Z digest=sha256:4523bd215a80b087154924b351f28c59d8a91c68663006ee7fba5f031294cc43

Observation 6ae9bfaa-2a9c-4804-aa2a-dc2b857795ce · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.920183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.259659Z digest=sha256:2bc900ec90a836159cc4f98afbbf605cdecfefcf840e883d86594d22e97a0bfc

Observation 71d02cb0-f2c3-4195-be9c-e0e90823e35f · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional chain-of-thought prompting for large multimodal models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.797006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.377593Z digest=sha256:c289d89074fc4cbbddeabc6f1b0046ab29f33796f3d4b640a3d3015f807d982c

Observation 89d54ed3-7d75-46e4-9c19-469c7336160e · outbound

This paper cites Compositional semantic parsing on semi-structured tables.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Compositional semantic parsing on semi-structured tables

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.676234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.497595Z digest=sha256:03498116589ffcaa3f8366ddfe896faf4101dec9456ccde7ce33ac2d18f2ac76

Observation 53b939d8-6681-442e-976a-2f9777000580 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.533383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.575202Z digest=sha256:559ec833d0a8b758b6b5f78e1c08ec7c0716ed9a285371e210199dbe17db0c5b

Observation 67ba3997-6b2a-4bf6-96b9-68bb29ab71f5 · outbound

This paper cites A benchmark of facial recognition pipelines and co-usability performances of mod- ules.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning A benchmark of facial recognition pipelines and co-usability performances of mod- ules

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.367570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.702678Z digest=sha256:ac0650d8779cee2401b918290910df374709b857402ee3bdd7ce38029387d787

Observation 65e6ac54-3195-45f1-bb0a-adf7d0b85a17 · outbound

This paper cites Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.277223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.818888Z digest=sha256:537f0be7b1c1c330f25301d134a5cb470aefca3cd4ab7057cf68b1d1e7ae0b0e

Observation 1494792a-ecc3-4427-8057-b946dd8698eb · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:47.182352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:37.933670Z digest=sha256:5550c324f64b5cdf00a29eebd956ebbc0fc702d97940b2635b1b4149840df95c

Observation be5ce8d3-57c9-458f-9ab8-cb9c930acd9a · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.079081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.079081Z digest=sha256:87a22c17e894176858be7c4a95e3c3260897be5dd95e7b83cd6ccb3b58aaee1d

Observation d7963086-d156-4656-8cc0-45f01486735e · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Textcaps: a dataset for image caption- ing with reading comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.200474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.200474Z digest=sha256:c734f2e47ab5906bbfe135b562f0204327178389d4efbe7cd87319106fc92637

Observation 031f965e-c81e-47ce-990d-88f3dc516a55 · outbound

This paper cites Towards vqa models that can read.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Towards vqa models that can read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.982901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:38.315186Z digest=sha256:ce4c2ab750b3ff1207230b11448f5d8c242baa1f261fa8b30a556316944eb16b

Observation 957c6707-7fcb-4195-84a1-e72f0b84820a · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.403343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.403343Z digest=sha256:5ccfbd87e84682e331f9bc45abf4a085bcb46d92eb13d1e541c6d6ffb4171827

Observation cc460c53-afcf-45cc-bb4e-231079fc7c88 · outbound

This paper cites Tang, Angie Boggust, and Arvind Satyanarayan.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Tang, Angie Boggust, and Arvind Satyanarayan

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.756525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:38.500912Z digest=sha256:1e384035e7759b915fa8c548894b3bd12ce140db1bcc915ac43c72bf8a5bd829

Observation e04a0778-012f-48fb-91de-e1f8e17e9509 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:38.622584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:38.622584Z digest=sha256:11ee09b2c3e9ace3c1f138ddce8c4e5165d3a15a5b5756ab2f3c0841bab45b49

Observation 771ed215-2e01-4c95-beb8-e346183c90ee · outbound

This paper cites Document understanding dataset and evaluation (dude).

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Document understanding dataset and evaluation (dude)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.522644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:38.703728Z digest=sha256:28762d2d51773c28698347e8d476b994cdf8a9528e9059a885dfdaa98e37a657

Observation 8a635456-6fae-42b1-b5e6-1c6d2232609a · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning The caltech-ucsd birds-200-2011 dataset

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.288111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:38.811721Z digest=sha256:d41dcc10a34c2e3c8be87208bbcdc247fe5e5aed23f90d89bc00277ec6e87fe1

Observation 17cea784-ef6c-4838-ba07-805935d5b36f · outbound

This paper cites Screen2words: Automatic mobile ui summarization with multimodal learning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Screen2words: Automatic mobile ui summarization with multimodal learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:46.064600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:38.899604Z digest=sha256:73161d7da2f07f2b5228edafe97b8b6814d840afcbe81d81f831bf44ab16fdd9

Observation 3d2fe1c4-0a8c-4ee2-bfd4-ebe8f53cfd67 · outbound

This paper cites LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:18:41.874604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:39.029360Z digest=sha256:7aef587152c598ccf3c6fb29729a3c633c173671fe2ab5fbbfd3f1c8774d2425

Observation 7509c407-ae30-4486-8ab7-12cd21de8e94 · outbound

This paper cites MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.118189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.118189Z digest=sha256:bedb1c64bebc9de5404a6408217e21aeda375071343e9317c018de34783bb36c

Observation 29515936-2adf-4b96-8dcd-d245ff7c8558 · outbound

This paper cites Mea- suring multimodal mathematical reasoning with math-vision dataset.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mea- suring multimodal mathematical reasoning with math-vision dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.869578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:39.207676Z digest=sha256:ab75c876116c464f05fe2210df62a8ef82c1d0159fa795c4786e94288e489202

Observation 5e426e2b-2a9f-4ee5-b01b-71570168c1a5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.306583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.306583Z digest=sha256:e69a9b7a751e04627b61261de589665965f11ca06febfce672a95539c9ea5e3a

Observation b2630173-c5e9-4b66-b4b2-b531121a1309 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.399123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.399123Z digest=sha256:54b6067796a554cf977be26a35dafee3c2f5b1cc5127b8896a026a9809645026

Observation 4197aa3f-1e21-4c07-b7c3-0bd0709b8881 · outbound

This paper cites Grok-1.5 vision preview.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Grok-1.5 vision preview

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.603526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:39.504236Z digest=sha256:8f3236008f874860d36cb6d221a658cab2ea7b335cde61936affeff026577bf9

Observation 744db546-4e65-46b4-96d3-0275004cc8e8 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:39.605249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:39.605249Z digest=sha256:71594780f08304077239a62a97b412aa7fdee8e57268d767e57be9dd54527d83

Observation 6af0f074-14a1-49a9-a077-7722867c6ca3 · outbound

This paper cites Llava-cot: Let vision language models reason step- by-step, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Llava-cot: Let vision language models reason step- by-step, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.387347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:39.749545Z digest=sha256:e891970fe5d532e9665d7b4d4e7ada17e7e7735f5535f30db8b9ab695755339b

Observation 6f935308-2bd6-42db-a8bb-dbd13a588ae8 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Gpt4tools: Teaching large language model to use tools via self-instruction

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:45.106833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:39.829944Z digest=sha256:8137cdc96b0f94972a4e59c5e11579c6a9c1a2da36f814dbe6217e222d4b1403

Observation fc6a8b88-5b9e-4155-82de-657a638a8d80 · outbound

This paper cites React: Synergizing rea- soning and acting in language models.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning React: Synergizing rea- soning and acting in language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.865434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:39.900854Z digest=sha256:2ddc52394e47f0b9a43c63300507c254f820d21de49295fb2f811db840fc46cf

Observation 6aed4166-fd33-4939-ae15-4e8f720b5c6a · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.013192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.013192Z digest=sha256:d68242dfaf88b4af3c6225e8e1484092f6849a775010d4a7c3f012c9a3e373a5

Observation bed3ff95-f0d1-42b0-b8b5-0604add61aee · outbound

This paper cites Agent lumos: Unified and modular training for open-source language agents.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Agent lumos: Unified and modular training for open-source language agents

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.640480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:40.111923Z digest=sha256:e30d40169aa279355e7b5a68ba39892711b482c9c2a9e3d7e70748b53849967d

Observation 09ac949a-d7ff-4981-a42b-e5acb84a3f31 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.437025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:40.209708Z digest=sha256:bffac190a0bf0e8c6b409df3c6bda6233e2b374c32f997825f1d36be20c4d1de

Observation 565e3f9f-ecd4-4be0-a9e7-b16f64f828c3 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.313560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.313560Z digest=sha256:2935da5070fc09074e0093af9b3c069fb38063ff8202642bf185c5d307d6f5be

Observation 0967ca33-3478-437a-911a-75eb29b5b5d5 · outbound

This paper cites Raven: A dataset for relational and analogical visual reasoning.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Raven: A dataset for relational and analogical visual reasoning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:44.189547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:40.418499Z digest=sha256:765560dea158f33d5c6b07e4e3eb1f28436042fdeeceff2d78c2562d826ba47a

Observation 0cc7d3f4-c98e-406c-8321-11f839aad7c7 · outbound

This paper cites Swift:a scal- able lightweight infrastructure for fine-tuning, 2024.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Swift:a scal- able lightweight infrastructure for fine-tuning, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.965856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:40.563398Z digest=sha256:bfa69ad2337cfc8fb6af7e725e4008d9c80c8e06c6af7e929ef515edc50c4cb2

Observation 4411d2fb-c374-493c-8caa-d4770086b41b · outbound

This paper cites Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Seq2sql: Generating structured queries from natural language using reinforcement learning, 2017

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.750174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:40.636458Z digest=sha256:64e65d6f078d5598938e61cd50d31eddfa25cbf4915980c88de262cb26a9bcad

Observation de5ef537-70bf-40f3-9766-c216ab3a5fd5 · outbound

This paper cites Visual7w: Grounded question answering in images.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Visual7w: Grounded question answering in images

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T12:18:43.524971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:40.695193Z digest=sha256:43399c0302241516f0e56d65125a892cca6c7ad2d5df801228909033ec056842

Observation 5256053a-33dd-46de-8251-152ba8f8af31 · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.808887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.808887Z digest=sha256:f15cc99770fc5b9004e50b2517f2446a652f19996e50c00c16f9b89224b2fcb4

Observation c81563ba-84aa-460c-a47c-c3f38c58dd28 · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:40.922255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:40.922255Z digest=sha256:14c71ad7c280d5e7aaa2b20c12b22ee460142e8fc513cd145597343ad87484a0

Observation 8c1b6981-1d26-4138-b7c8-99894dea8acf · outbound

This paper cites objects": [.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning objects": [

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:43.246936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:41.042512Z digest=sha256:84e7b682b3fb7c10e30a974ad49922321be4380add535aeec418f49184b79189

Observation 05128088-b15e-47bb-ac2d-8c23e84e2e5c · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:18:42.999932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:41.141614Z digest=sha256:89c405bd1e77d2f4e2f21f1b35d08286f2dc8c4a5c35eb201f62e2faf48ddf9e

Observation 3f7bae2a-c3bb-48b8-98f5-9085610ac3ff · outbound

This paper cites image_caption.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning image_caption

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.706027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:41.259999Z digest=sha256:a7401f91cb50ace10761cb857619f2fe027afcb11d24a5f60bfddb836bfc6138

Observation d5f08797-6684-43ed-88c0-c5610b21feb7 · outbound

This paper cites needed": true,.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning needed": true,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.471282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:41.390516Z digest=sha256:0c9dd6cfeb6a2b5059f64962aac0f5d9386c5fdb687292c7e751f9b803e64b94

Observation 72e7adf1-f915-45ba-b014-9a7760bf60b4 · outbound

This paper cites continue.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning continue

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:18:42.265587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:41.518126Z digest=sha256:f0d777f09108d633fe1d142956af6bc306245087ccada0d33b469e3ca272f0ef

Observation d63ce0b8-1ee8-4fc6-a71d-aea9de96110b · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:33.027058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:33.027058Z digest=sha256:682d2159f39b4083a6ba755deea0362f06ad3d2d3a8cecf7c24ff1c87423f9d4

Observation cf2204af-5146-43f9-9318-0ccc42de42dc · outbound

This paper cites an unresolved cited work.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Unresolved cited work

Reference 251

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T12:18:51.528283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:18:34.612419Z digest=sha256:41414f4e74e02baec7b916aac2fe8d5e333871d49aef1b452e407cb3f595d59e

Observation 846a591b-e1c1-4d71-bf21-87c1e714b30f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:31.416961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:31.416961Z digest=sha256:b130d1868598490e874fe5f35b9a19f8a950ddeb52d1c52acc1edd07d4b41d5b

Pith citing papers

No inbound Pith citation observations are available.