Pith. sign in

Paper Citation Record · LEDGER

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

As of 16 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 41 inbound Pith citation observations for arXiv:2412.05237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05237 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:56:42.553926Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.013019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.154168Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7e3bc64-1496-4982-83c9-d097eb620ec9 · outbound

This paper cites Consider factors like the amount of detail, depth of content, and how well the conversation and image complement each other in conveying comprehensive information.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Consider factors like the amount of detail, depth of content, and how well the conversation and image complement each other in conveying comprehensive information

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.388029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.338938Z digest=sha256:b4a1c98a4b5c107f66f3c9020b0e7ff7a699a3adce4033b1a98d4a660553e348

Observation 054f2f9d-603c-4a8a-93be-addcd88cacf1 · outbound

This paper cites 1: Very low relevance, the conversation and image are almost unrelated.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 1: Very low relevance, the conversation and image are almost unrelated

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.374735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.343675Z digest=sha256:49adb6f419d34e499b4ce2d5085ead4b00a0c7ddeb5dde29c52a320b575aa9cb

Observation cc6065a3-20e9-46f8-8459-6ed942434c6f · outbound

This paper cites - Instructions should require the responder to infer and utilize visual information that may not be explicitly stated in the instruction.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Instructions should require the responder to infer and utilize visual information that may not be explicitly stated in the instruction

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.280677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.375026Z digest=sha256:e73d15c30dceb1bcd10718fe64b9c989871e8ff6ec0ef43c2449e2a8a3959bf2

Observation 490add22-3660-420b-837b-e23c3c1befcd · outbound

This paper cites HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.304858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.304858Z digest=sha256:2d30d76b7a82725e27a34e87c864ea04681d6f984f70f2cb40ecce34278437a2

Observation e3ca2bd3-d272-46fd-90f0-d7db42ebe564 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.201089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.401975Z digest=sha256:57e37c1398de5ce48219d517ad8b474081eb8f95c10c75ee6176fffcdf96759e

Observation b4673661-6c40-4a8d-9570-0215e2079d1e · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.187488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.406176Z digest=sha256:149e4761917e347560dc392caa6c598c81d66ae1dc3900ab55ea40b508539b17

Observation 980537b2-f3a8-42a7-b1fc-9a4aea3d719d · outbound

This paper cites The AI assistant’s tone should be neutral and professional.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale The AI assistant’s tone should be neutral and professional

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.174186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.410370Z digest=sha256:a523a36c9bc104872730b7234ec27710ab5cf4722119a6b70b98a85806cc4bed

Observation 3ef78003-fb41-4d07-9f99-2ff4665ba8e7 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.161023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.414758Z digest=sha256:1eb4c354865da9635e39457ddc7533f89d4212dc30ed14a826206575fa5364de

Observation aca66697-c216-4715-85eb-db3bbc45ccd8 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.147364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.419054Z digest=sha256:25d20135574ca95063e6bb301fd2dc3cb36df53ef7658efae4064ed7e52e167f

Observation eed179a7-2c03-447d-9158-87fb19a44461 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.361513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.348300Z digest=sha256:ac68622713dc9044021e14a3caafb2d2c553d39b72315cb7ff3494e7d86abfa9

Observation 87f8965e-ac06-4893-8272-94b9129eb785 · outbound

This paper cites ##Instruction##:.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale ##Instruction##:

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.348657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.352922Z digest=sha256:08e22427e37d374a87545db6d60fb37ae7e0703bc1c87fc78cf70ee123f0b33d

Observation b148c89e-0327-4551-9f70-ee52fd209962 · outbound

This paper cites - Ensure the response is exhaustive, covering each stage required to reach the final answer, while considering all details from the image.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Ensure the response is exhaustive, covering each stage required to reach the final answer, while considering all details from the image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.335367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.357289Z digest=sha256:23fbe7faaeaefa4d6eb29c196ec143631b6836d6f493b0ad6cff5b120a35aba0

Observation 6be6ebf7-3f10-445b-b57b-f662655085a7 · outbound

This paper cites - Do not include additional text or explanations outside of the required <response>.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Do not include additional text or explanations outside of the required <response>

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.321865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.362132Z digest=sha256:6001fec317327954f3b9689d0c952e5212f340ecb39853482b125d4fccffc0e3

Observation 64cfea06-9afa-416f-94cb-01c1057aed1f · outbound

This paper cites - Ensure each instruction is unique, complex, and related to the given caption and task type.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Ensure each instruction is unique, complex, and related to the given caption and task type

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.308063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.366447Z digest=sha256:37685075240f1fc43da3413545e974515bdfee2536ce29db92bce72d93c0cfe6

Observation f7af23c8-b25c-4102-bae2-fa3514812dc7 · outbound

This paper cites - Include in-depth explanations, multiple perspectives, or detailed steps where appropriate.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Include in-depth explanations, multiple perspectives, or detailed steps where appropriate

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.294173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.370692Z digest=sha256:5349cfb8a6eca277e874f6513ff5d17bc0ed79707e177a24e6692b8137147f4b

Observation ce8a3843-e2d8-4493-886b-a25fdc9c48fa · outbound

This paper cites - Your output should only consist of the generated <instruction, response>pairs.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Your output should only consist of the generated <instruction, response>pairs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.267094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.379111Z digest=sha256:763a5273b4493d220de8d0a4617f42476834cf837b9db04212d1d6d44f46a2e0

Observation df65a6b5-4dd1-422a-8665-cc84b0be16fb · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.253434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.383529Z digest=sha256:1ba3e356fa54e8feadd9a56d5ae7d7d0703df6c1d7eeceff339f3aadb5d5bf13

Observation 40563e0b-3ebc-49c3-b367-f67209f86610 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.240261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.388823Z digest=sha256:a5acc733dcda3e8aa3642f48446c7ab1a924ec446c88961b76702801664e8bc3

Observation fc48d593-b713-45aa-9cf9-890626096fec · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.226927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.393212Z digest=sha256:1edd9019d7bf8e6577f0ace5a3ab09f0884f07134504685c8f8b1fe01ee4433f

Observation f3d01dc4-bc1c-412b-ae5b-597e5a70fb7d · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.214043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.397418Z digest=sha256:0a7fe073faeee01f14fc6b983e1c1ff43cd0909ece643dfbb3ef76023d01da47

Observation 2e351fad-485c-4cb8-9b66-58744fa51813 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.133911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.423605Z digest=sha256:cb5741eaa7093fb2a868b44092f1ee0f91175893fe8cf3962ca1f7e1ad1f54ba

Observation 913ca96c-b55e-466b-9c47-49cc2c28c219 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.119616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.428168Z digest=sha256:a8a43fb2f35c881d62275d3b6ce6ae2c1a007562e4765fa01727e26c72f1a147

Observation 4d193954-d9f8-4d00-bb40-e2f84bb6f8cb · outbound

This paper cites Super VGA Graphics.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Super VGA Graphics

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.105600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.432381Z digest=sha256:a3e8038c36c19746a56204b151091a744a4bad26dd485db8d0075f145bef8cc0

Observation 1f77e1e9-d855-4f88-9bb8-7582f74f9eff · outbound

This paper cites Northgate Graphics Card.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Northgate Graphics Card

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.091654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.437029Z digest=sha256:f20478d08a62861ae73763f736388fdcbcf93cb57b70bce8c40b7ba65e9374d1

Observation 539d5bcb-684a-4321-8e20-43a0789c63d0 · outbound

This paper cites OS/2 READY!.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale OS/2 READY!

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.077927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.441483Z digest=sha256:e482c00ce0761c6272193915ea4a0dc6e2e74d69c6c1033654d60575f2bfbec3

Observation 6b5f87fc-730f-4293-bc35-8741f4262f05 · outbound

This paper cites 101-key Click-Tactile Enhanced Keyboard,.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 101-key Click-Tactile Enhanced Keyboard,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.064701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.445737Z digest=sha256:61b4d94f8f4379bfda5930a0df061cd1c38d4e78d7d5e6cf2c3aad0c33bf515f

Observation 05522b21-9ea1-44eb-b387-2d966eb21eb9 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:43.051481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.449956Z digest=sha256:f6f70ee8572c1ca5899fcdfb80a0c6ac60775780165e28c858ce32bae2a2b60a

Observation 904d6eff-a0b2-4b48-96fa-332ced07725d · outbound

This paper cites This mismatch leads to an incorrect average.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale This mismatch leads to an incorrect average

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.037650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.454484Z digest=sha256:38afca78dc1dbbd232c043244f43a1fed76f102f09ba892d582f5aba1ff3d5ac

Observation e900687f-ad93-4835-941a-9c0e0085164e · outbound

This paper cites average is 60% and Germany’s average is 61.75%, but the recalculated averages (60.75% for the U.S.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale average is 60% and Germany’s average is 61.75%, but the recalculated averages (60.75% for the U.S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.023454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.458802Z digest=sha256:b6a1973a0b13e7ee1c7dfc2d55df7ee7cb8963184566c85d5d6ca4579becfca4

Observation 54d70588-0b2e-4121-93f7-41eed9c3ffda · outbound

This paper cites more consistent and stable support for NATO.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale more consistent and stable support for NATO

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:43.009172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.463008Z digest=sha256:e8245acb0064bda9846fa01a2e871c4336346f0d9110c48143cb545c348d6f0e

Observation 555f2412-66db-4963-9312-45eedd7579aa · outbound

This paper cites For example, U.S.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale For example, U.S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.995561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.467078Z digest=sha256:bd1faad1b15dca78a67c0feb00de04b56e9e64b2bd85d3c7d2f015dd95222df0

Observation 3455b15b-5e85-434e-bac7-6b50a21dac3c · outbound

This paper cites and 61.75% for Germany) without clarifying that this is an approximation, which can lead to a misrepresentation of the actual differences.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale and 61.75% for Germany) without clarifying that this is an approximation, which can lead to a misrepresentation of the actual differences

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.982482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.471130Z digest=sha256:a5e68a37e714809415ed2fab4a7801394f849be0bab74dcaf199631be72c4de3

Observation 901f17e3-9726-49fd-b0e8-937bbf6e10da · outbound

This paper cites - The sum of the angles in this triangle is also 180◦.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The sum of the angles in this triangle is also 180◦

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.968056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.475269Z digest=sha256:b4961e860056ee06d7d02157a4c3820ded8f3eff6abb1b3a92995199939d429a

Observation b42b6478-f991-4be7-a496-d5e8ef540d19 · outbound

This paper cites - Therefore: ∠1 +∠2 = 180 ◦ - Substituting the value of∠1we found earlier: 70◦ +∠2 = 180 ◦ - Solving for∠2: ∠2 = 180◦ −70 ◦ = 110◦.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Therefore: ∠1 +∠2 = 180 ◦ - Substituting the value of∠1we found earlier: 70◦ +∠2 = 180 ◦ - Solving for∠2: ∠2 = 180◦ −70 ◦ = 110◦

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.954478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.479622Z digest=sha256:262fac6d0d29aacc68206a1dc0cd7654be5d27f3d5ee4e70345378495a5e2bee

Observation c2922e6a-7666-46d4-a238-e5e9f64f3d5c · outbound

This paper cites Revised Answer: The answer is D.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Revised Answer: The answer is D

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.941111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.483568Z digest=sha256:230b7742bf42a2495533ad82031b1af7c82b53c3f2386da3275240dbc7319512

Observation 81f423b2-4dc5-4184-8c7e-8229c6c45d7c · outbound

This paper cites -Runner B runs 3 times as fast as A and also starts from the origin.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale -Runner B runs 3 times as fast as A and also starts from the origin

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.927262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.488215Z digest=sha256:a0117c278a841ed7e3a7450fd18f546395b632d6c568fe8938269335ab59efa5

Observation f5c44316-db32-428b-a3cb-8efed150cfed · outbound

This paper cites - Since B runs 3 times as fast as A, if A runs a distancedin timet, B will run3din the same timet.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Since B runs 3 times as fast as A, if A runs a distancedin timet, B will run3din the same timet

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.914158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.492800Z digest=sha256:dfb44117b6808be7191ed24354d6e5066da812519bbfaafe277efcfc1024b0b8

Observation 62d4022c-fa17-4097-876f-2faf66eff398 · outbound

This paper cites - The coordinates of B are(3t,0).

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The coordinates of B are(3t,0)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.900446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.497297Z digest=sha256:f8f0b9e6fb7a3a6e71e9f62315b943331afc1cc3c90e2ce925d6696bd4ff79df

Observation d8e8f750-2c8d-4bff-98b4-480b5edc4ab6 · outbound

This paper cites - The angle formed by the line connecting the observer to A is α= tan −1 1 t.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The angle formed by the line connecting the observer to A is α= tan −1 1 t

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.886119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.501617Z digest=sha256:9ce67825072b45a68211db65f721dae0416fbb2e419daf2626e579bc1558db28

Observation b8e30063-a3cb-484a-bbd5-807401b100d7 · outbound

This paper cites - Using calculus, we can find the critical points by taking the derivative ofα−βwith respect totand setting it to zero.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Using calculus, we can find the critical points by taking the derivative ofα−βwith respect totand setting it to zero

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.872633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.505884Z digest=sha256:22786ce5abb632515dbf1e2a4ec87be2070102c146edb05b30f6c3087cf664ea

Observation b4cfc01e-b388-42e7-a6d4-037225beb783 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:42.859405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.510227Z digest=sha256:a4d43b3dac0c74f46bfc6e9ce2f4bd3342dac4a7bfb5f4535933e7722e75e655

Observation 2797f5f0-7bfd-483d-8f93-4c5eedd29495 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:42.846098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.514615Z digest=sha256:ccfd0310f8df06da6d16181d92f49495a120b4db9e8c1dbb1af2185722589689

Observation 530c6f4c-cc36-4f3c-a3a2-654484d1f954 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:42.832931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.518997Z digest=sha256:150419ad0dd674951c142edf2f08a023c0c50d1faf4d245416b8cc776859555a

Observation a8beddf4-177a-4e1d-846d-074e5c626f84 · outbound

This paper cites Thus, the maximum angle of sight between the observer’s view of A and B is30 ◦.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Thus, the maximum angle of sight between the observer’s view of A and B is30 ◦

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.819339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.523540Z digest=sha256:246369fb9f666c17af8939cc76cea9f636df91e4faa255561c911eebd642934a

Observation 8ddba92c-9056-4f49-8d3f-57b9aea6b5d4 · outbound

This paper cites The Roman line starts at a higher population and declines over time.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale The Roman line starts at a higher population and declines over time

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.805306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.528081Z digest=sha256:67fd6793a1a531e7f613ac86389a99c8173befe1b4aff6e5078dc125b076514c

Observation dc558812-2717-4c1b-9559-6d14575b7488 · outbound

This paper cites However, it does not explain why the Roman population started higher and declined more significantly than the Han population.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale However, it does not explain why the Roman population started higher and declined more significantly than the Han population

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.790869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.532415Z digest=sha256:804a56acf2d30a12d35d0761ba616fc98e2a91d97bcacbca7d18569f1d98609b

Observation 8767b64b-a356-446f-8d1e-104702836751 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:42.775144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.536634Z digest=sha256:b32f2e30966af911d1b9501dcb9dc5e990020a0a623bad6a75522387786e4c00

Observation 48a2578c-abd0-4dd5-b703-9720487a0112 · outbound

This paper cites - The sum of these two angles is: 28◦ + 82◦ = 110◦.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The sum of these two angles is: 28◦ + 82◦ = 110◦

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.761651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.541280Z digest=sha256:726460ba01098a9df4f27a55f5bac3e81ae356b04bcbe8b9b5a54a33a2c83024

Observation d7f946ae-e8bb-489d-ac2d-a034ef99c923 · outbound

This paper cites an unresolved cited work.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:56:42.746892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.545488Z digest=sha256:e3b249ef1741b80cef508f2127c4ece290d2017876859118b8a62bf6afc728a7

Observation 0ef85754-7f1e-44bd-9912-88d3a5333eab · outbound

This paper cites - The sum of these two angles is: 68◦ + 70◦ = 138◦.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The sum of these two angles is: 68◦ + 70◦ = 138◦

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.733029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.549790Z digest=sha256:767aa3f3add2bc3d64b49f83e6d617bd5c7f5462f29f7eed4aa09f7d062e38a5

Observation 55a1a368-25a4-4e9a-ab79-2de937aff977 · outbound

This paper cites Meh... Don’t worry about it. I’m a New Englander, so I’m used to it.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Meh... Don’t worry about it. I’m a New Englander, so I’m used to it

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:56:42.718088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T20:56:42.553926Z digest=sha256:955910e5b9474ad9439a036c24b971857a4792276544d061fa76ebcae3ede486

Observation ab998c48-f02e-47b6-8dd6-bd14219e26c9 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 233

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.316302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.316302Z digest=sha256:58ddc79bd2b95a98ed0586722914a52fdc140b8d9dad7ac77ed489d835d52d0e

Observation 52d23f46-1c7e-4ab5-ada8-1cdc7367bec6 · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.326943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.326943Z digest=sha256:16e48515dd3a42962430b7b2939da63b6afd6f7c1bc5105f58fc2334a509815d

Observation 51ff339f-1a4a-4af2-aad2-3e22e2e8b2d5 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Large Language Models are Zero-Shot Reasoners

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.310449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.310449Z digest=sha256:bbf9e0efbd80bb61049b13f631a1ab233f09811d9581a6747c5e454b5d004c60

Observation 4ce2f54f-6743-40d0-b54f-ed1e5658e3c8 · outbound

This paper cites Qwen Technical Report.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Qwen Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.288573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.288573Z digest=sha256:bf0891a55304c56400e78d1177bdc644aa56499aee95f786ff91c44b37ce60fb

Observation 01d11d0a-bccf-47c5-b7a6-7b1465b2282e · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Reference 2022

Resolution
malformed identifier
no resolver link, observed 2026-08-11T20:56:42.333159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.333159Z digest=sha256:d44a7e60cfa572bfc2aa7e88a7dfaa07fbe96a25df32dcf987dce0c18c14dd6a

Observation 7469164e-c094-46e2-821e-e7581b72d8a0 · outbound

This paper cites G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.299616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.299616Z digest=sha256:84b2f9a67bd610e91e79eae963f819cb5b9f2eae49b53758789a280da73a0abc

Observation dbe0194a-c4aa-44c5-aa3a-1bf5f8819249 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.294503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.294503Z digest=sha256:9abe4bf4d49317cb88f94935619d423c408740caf972678f219ae1900ade8a35

Observation 0fa03c66-b29c-45aa-8eb7-7200d57f6a63 · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 2521

Resolution
unresolved
no resolver link, observed 2026-08-11T20:56:42.321445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:56:42.321445Z digest=sha256:eef56441342c175d9de824367e0bed1a63aeb18b53068160bd65cc9feae4aa2b

Pith citing papers

Observation 37976e9a-38a0-43ba-b502-0c0f2ee8555e · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.188483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.188483Z digest=sha256:aea8a43e30598e252d2755009ce7275cea79bb567d50c1c392f31847efd00acd

Observation c46585a1-cbc9-40e3-9432-9fe4c0339963 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 294

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:10.041348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:10.041348Z digest=sha256:5b6dd77761331541dccc9db8d9f6229e32a3c1965aa02fa4228971df8e9fbd76

Observation f02c1b66-f08c-445d-a481-2881b210d20a · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.147822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.147822Z digest=sha256:bed607506e5b81fce16e3261b7f1a0e78b39e93878bb5797ea5d0b32252b71ec

Observation b93518cb-dc78-4422-81fb-50abce96ca5d · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.077389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:b72472baafaf9b87442e8a40edb69cb18c7026f6f05eb75468febc289ac81f21

Observation d37b096c-0952-4d52-aa29-37be6b437a11 · inbound

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization cites this paper.

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:19:20.527581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T00:19:20.462455Z digest=sha256:84bda6cc64832714727ccea69faf42763b1f39aa50e6cf8df653e44c53176faf

Observation 0ccbae95-297b-469a-988f-6eacb9334bb8 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.580628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:a0995f54d54073faf0c314c341f3ccfc01c9c2d87433779c8ad29a9ae65c58b1

Observation 7df89b55-a208-4c96-8197-b601b33a9e54 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.395929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:08288bc4c396ef22b54b2c5c6e7fda7d11ce8ec70b8e47dca5f03f28944efa69

Observation 8ef85ad1-efd7-4ce8-b69a-92e49a222d70 · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.709217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:939d30b17353b545e71f468950ad3ebb40cde9bf8432d51f86c800d678f7753e

Observation 25744b38-c729-4498-af55-83cf671defc5 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.013019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.013019Z digest=sha256:19fec1eff3a1fadd407ceb207181ef290bbaaa3fba2dd75ea9a47831ef5a9ccd

Observation f4cdf063-283d-4c84-8323-2a416e55d54e · inbound

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs cites this paper.

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:28.033287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:32:28.033287Z digest=sha256:cd88a1ca9813d579724d4b360009b1a27245c5e6e024d14863613e808e0a0104

Observation e98f92c9-47e8-4a08-a328-0c1cdccc6de9 · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.806428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.806428Z digest=sha256:0869b19dddf97bebf762acb47d50854e0e4f481b06fea4ae96a84c73890562d4

Observation 954eea10-c87b-43cf-9ec2-6b1c8ad09be5 · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.186601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:3f62f03862f627bc0124bd2237539ea8a9ea062f4d75a3a9be11b2299e00816b

Observation 703e8ac5-584c-48aa-bd45-2f71d209e61b · inbound

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models cites this paper.

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:21:39.771573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T14:19:34.622854Z digest=sha256:26657d7fc20acc9a39bfca7f2b3c981d947b12b4ee0d6419d402d5d49667b2cb

Observation 74420e64-e9bf-490e-9bad-8518905b5848 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.232126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:fab68895a529b03c1230ff909417d54982cb8e8be1670542cdb5ae106ecc50ea

Observation 535af07f-bec3-4081-8631-57ba21dd4d66 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.775772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.775772Z digest=sha256:184c8afa826ed3f779a9dcf1b976be0dd551e994cf91e4d1c43fa65332d7b323

Observation f099e45b-584b-4429-90f8-7bfef4ee4b09 · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.708467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.708467Z digest=sha256:ed58a36bbfaaa75969988c294f8082582467f66de324ad01cb71913cb5dd72ba

Observation 4cc616f9-4037-4362-bdce-c9897d81045c · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.530844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.530844Z digest=sha256:0096d5aa90c94053f1f149f0d5434a8f20c96cd540c1fae1dc6910381191b60a

Observation 2e5b393c-95b4-44a1-8b11-cd619c3751bb · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:59.200727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:59.200727Z digest=sha256:f9b6b55c94ad7926bb692a8b7eae3fed811eff17860a755133304f40a9c89490

Observation e5b51c26-c286-46df-9016-544557ae8b24 · inbound

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification cites this paper.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.600959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.600959Z digest=sha256:a374c0ad647920fc1ad1a3f3824f8f0066a52f7f03dfa271e75ac4312a26e00a

Observation 545bdbd5-21ae-42ef-983b-af63ca588c00 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:18.908261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:18.908261Z digest=sha256:358bbd8e602351c350d77d5b0df370399d4484027efdd041580a08b3b45702cd

Observation 456f258d-cdda-45e3-a431-c54fb8de5b8a · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:03.216795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:34:03.216795Z digest=sha256:4a0e589b05c482c33e038e942f7f0691e92e2fddbcce527f5d35007f8a8e2812

Observation de573936-c557-4175-a5a9-a506d29443bb · inbound

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings cites this paper.

MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:22.687226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:22.687226Z digest=sha256:54f355f98bb05636bf557590e1f14468cb4e01a600ed7779c70fb1ae5971b988

Observation 4a6c5b47-e72a-42bd-911d-b726c8c2e9cd · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:39.353274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:39.353274Z digest=sha256:759ab072b86fb792966fe060c501d5f2c340c332fffc1153c31ebb6a15b84aac

Observation 4d4fd369-42ff-48dd-aa16-8a3b2a588772 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.742706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.742706Z digest=sha256:1669508a6f9be6045863ff1270f6575108528da34e843166247db27aa7d2a750

Observation 48f42d94-ef92-4dc4-9138-292463bcf64b · inbound

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs cites this paper.

The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:12.605646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:43:12.605646Z digest=sha256:99fd65314073766fd0fc50d53afdfe59457a1d1c0129d3d5a8261b6139fa3ae4

Observation 6fb502f8-5f62-4b9c-91fe-4b1e13d4cbbe · inbound

A Survey on Diffusion Language Models cites this paper.

A Survey on Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:26.877320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:26.877320Z digest=sha256:7c950c8df703f784d75a4e9c9cf91f8bb64ff3ecbf104ee546c0abfeeceb0fe1

Observation b685855f-c7de-430b-a29e-3eb89d776736 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.399950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.399950Z digest=sha256:0d7a7b9128e41ebfe3f3ea046927c6b253ebab94c891225cc782218d0f7f2b9b

Observation 262f9d64-69a0-4f8f-bde8-be33d8710e31 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:35.481243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:35.481243Z digest=sha256:aaf4ace508a0f33fff4706f155859ea8bf98af9e7fd5fdc2747af6c6d86d08b3

Observation 0cdb59a1-27ed-4549-b01c-685666eba19b · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:19:00.503308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:1819df19e8d49d333eb4739c5d884bbae7e5cc790040a9a92be7f898ff162937

Observation 422dbeb9-5e7f-4fe7-90ca-88f2e6ae6291 · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.491974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.491974Z digest=sha256:ce3cade4442aec82500936dcf7922075d01f720891e9fb13143ea138a046f1c1

Observation ddbf47d8-68f0-498c-9092-dccac2e046ba · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.394620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:4a5b0a06ab282076d10bb259b26476541d51afe616d3a8d80661c0173e9d0e62

Observation 0f5eeb8b-5a5a-402f-8a76-d3fcc9128ac7 · inbound

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model cites this paper.

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:57.788020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:29:12.064221Z digest=sha256:7863d7a1fd01804a7def1efff2749397531e0cfc379d6dd1f7dc599b7fce06a3

Observation 7cd4543e-76bc-4aef-82f1-3d065e399684 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.313335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:fd0d6d98d709807e5ca12b8e2cd63d6c17e8953e4ad44db119c059af3040b2dc

Observation 76be5cd6-6450-49f2-94c9-332626a6307e · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.557358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:6fec4bed12a05216c8e6f836300be6881be02cc043ec4f03d76d1040932dac20

Observation 6e757b67-7e18-4a2b-886d-e421665c2846 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:57:09.462443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:6d4a9e16f626bf3244618154e2d1f1c5bd13850a5f49148028557bb7d3e3dfbf

Observation 17807cb5-4354-418f-bf11-a74de5b5aafc · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.594707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:c442088700b26b53f25fad66066b515e28963d1f00dcacc7ec895b2838f80c20

Observation d4359dae-d5c6-4551-9dd7-9321babad716 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:48:23.449898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:9de7e64f8bb9d32885d47ad3e2a4867e136eea109b8a8de7d0a15a1480857d18

Observation 6fce37bc-1f7d-4864-8b6a-863019af4870 · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.687479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:ac69a4e1a3ace39fa264ea1a167316b18b7c774c3f0bb3afdbdd2f20bd40967f

Observation f408383a-4775-42a0-928c-b751d250b4af · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.896690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:4432a594dffbb7c978d5a8f7f09a6bd2e94b9d3fca02964e431c29e53959556f

Observation 53f7f1b4-06df-4382-85c3-14791267f844 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.156125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:1e94d4ff72004bddb736b9996d55c0885a92ca609b568f85064427836e5f6541

Observation a48d31d2-7767-4a25-86a5-01e2722079ac · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.583796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-01T06:30:27.178950Z digest=sha256:c0e9ada559142f65c1576f1d48116520708d48f5b26dca8c5ee06a61491601b0