Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:56:42.553926Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 41 inbound Pith citation observations for arXiv:2412.05237.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:56:42.553926Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.013019Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:40:07.154168Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a7e3bc64-1496-4982-83c9-d097eb620ec9 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Consider factors like the amount of detail, depth of content, and how well the conversation and image complement each other in conveying comprehensive information
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 054f2f9d-603c-4a8a-93be-addcd88cacf1 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 1: Very low relevance, the conversation and image are almost unrelated
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cc6065a3-20e9-46f8-8459-6ed942434c6f · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Instructions should require the responder to infer and utilize visual information that may not be explicitly stated in the instruction
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 490add22-3660-420b-837b-e23c3c1befcd · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3ca2bd3-d272-46fd-90f0-d7db42ebe564 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b4673661-6c40-4a8d-9570-0215e2079d1e · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 980537b2-f3a8-42a7-b1fc-9a4aea3d719d · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale The AI assistant’s tone should be neutral and professional
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ef78003-fb41-4d07-9f99-2ff4665ba8e7 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aca66697-c216-4715-85eb-db3bbc45ccd8 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation eed179a7-2c03-447d-9158-87fb19a44461 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 87f8965e-ac06-4893-8272-94b9129eb785 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale ##Instruction##:
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b148c89e-0327-4551-9f70-ee52fd209962 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Ensure the response is exhaustive, covering each stage required to reach the final answer, while considering all details from the image
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6be6ebf7-3f10-445b-b57b-f662655085a7 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Do not include additional text or explanations outside of the required <response>
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64cfea06-9afa-416f-94cb-01c1057aed1f · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Ensure each instruction is unique, complex, and related to the given caption and task type
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f7af23c8-b25c-4102-bae2-fa3514812dc7 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Include in-depth explanations, multiple perspectives, or detailed steps where appropriate
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ce8a3843-e2d8-4493-886b-a25fdc9c48fa · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Your output should only consist of the generated <instruction, response>pairs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df65a6b5-4dd1-422a-8665-cc84b0be16fb · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40563e0b-3ebc-49c3-b367-f67209f86610 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fc48d593-b713-45aa-9cf9-890626096fec · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f3d01dc4-bc1c-412b-ae5b-597e5a70fb7d · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2e351fad-485c-4cb8-9b66-58744fa51813 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 913ca96c-b55e-466b-9c47-49cc2c28c219 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4d193954-d9f8-4d00-bb40-e2f84bb6f8cb · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Super VGA Graphics
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1f77e1e9-d855-4f88-9bb8-7582f74f9eff · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Northgate Graphics Card
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 539d5bcb-684a-4321-8e20-43a0789c63d0 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale OS/2 READY!
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6b5f87fc-730f-4293-bc35-8741f4262f05 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 101-key Click-Tactile Enhanced Keyboard,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 05522b21-9ea1-44eb-b387-2d966eb21eb9 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 904d6eff-a0b2-4b48-96fa-332ced07725d · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale This mismatch leads to an incorrect average
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e900687f-ad93-4835-941a-9c0e0085164e · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale average is 60% and Germany’s average is 61.75%, but the recalculated averages (60.75% for the U.S
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 54d70588-0b2e-4121-93f7-41eed9c3ffda · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale more consistent and stable support for NATO
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 555f2412-66db-4963-9312-45eedd7579aa · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale For example, U.S
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3455b15b-5e85-434e-bac7-6b50a21dac3c · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale and 61.75% for Germany) without clarifying that this is an approximation, which can lead to a misrepresentation of the actual differences
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 901f17e3-9726-49fd-b0e8-937bbf6e10da · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The sum of the angles in this triangle is also 180◦
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b42b6478-f991-4be7-a496-d5e8ef540d19 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Therefore: ∠1 +∠2 = 180 ◦ - Substituting the value of∠1we found earlier: 70◦ +∠2 = 180 ◦ - Solving for∠2: ∠2 = 180◦ −70 ◦ = 110◦
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c2922e6a-7666-46d4-a238-e5e9f64f3d5c · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Revised Answer: The answer is D
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 81f423b2-4dc5-4184-8c7e-8229c6c45d7c · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale -Runner B runs 3 times as fast as A and also starts from the origin
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f5c44316-db32-428b-a3cb-8efed150cfed · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Since B runs 3 times as fast as A, if A runs a distancedin timet, B will run3din the same timet
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 62d4022c-fa17-4097-876f-2faf66eff398 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The coordinates of B are(3t,0)
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d8e8f750-2c8d-4bff-98b4-480b5edc4ab6 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The angle formed by the line connecting the observer to A is α= tan −1 1 t
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b8e30063-a3cb-484a-bbd5-807401b100d7 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - Using calculus, we can find the critical points by taking the derivative ofα−βwith respect totand setting it to zero
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b4cfc01e-b388-42e7-a6d4-037225beb783 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2797f5f0-7bfd-483d-8f93-4c5eedd29495 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 530c6f4c-cc36-4f3c-a3a2-654484d1f954 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a8beddf4-177a-4e1d-846d-074e5c626f84 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Thus, the maximum angle of sight between the observer’s view of A and B is30 ◦
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8ddba92c-9056-4f49-8d3f-57b9aea6b5d4 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale The Roman line starts at a higher population and declines over time
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dc558812-2717-4c1b-9559-6d14575b7488 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale However, it does not explain why the Roman population started higher and declined more significantly than the Han population
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8767b64b-a356-446f-8d1e-104702836751 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 48a2578c-abd0-4dd5-b703-9720487a0112 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The sum of these two angles is: 28◦ + 82◦ = 110◦
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d7f946ae-e8bb-489d-ac2d-a034ef99c923 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ef85754-7f1e-44bd-9912-88d3a5333eab · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale - The sum of these two angles is: 68◦ + 70◦ = 138◦
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 55a1a368-25a4-4e9a-ab79-2de937aff977 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Meh... Don’t worry about it. I’m a New Englander, so I’m used to it
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ab998c48-f02e-47b6-8dd6-bd14219e26c9 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 233
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d23f46-1c7e-4ab5-ada8-1cdc7367bec6 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ff339f-1a4a-4af2-aad2-3e22e2e8b2d5 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Large Language Models are Zero-Shot Reasoners
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce2f54f-6743-40d0-b54f-ed1e5658e3c8 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Qwen Technical Report
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d11d0a-bccf-47c5-b7a6-7b1465b2282e · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7469164e-c094-46e2-821e-e7581b72d8a0 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe0194a-c4aa-44c5-aa3a-1bf5f8819249 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa03c66-b29c-45aa-8eb7-7200d57f6a63 · outbound
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 2521
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37976e9a-38a0-43ba-b502-0c0f2ee8555e · inbound
FastVLM: Efficient Vision Encoding for Vision Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46585a1-cbc9-40e3-9432-9fe4c0339963 · inbound
Visual Large Language Models for Generalized and Specialized Applications MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 294
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02c1b66-f08c-445d-a481-2881b210d20a · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b93518cb-dc78-4422-81fb-50abce96ca5d · inbound
Qwen2.5-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d37b096c-0952-4d52-aa29-37be6b437a11 · inbound
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ccbae95-297b-469a-988f-6eacb9334bb8 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7df89b55-a208-4c96-8197-b601b33a9e54 · inbound
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8ef85ad1-efd7-4ce8-b69a-92e49a222d70 · inbound
SmolVLM: Redefining small and efficient multimodal models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 25744b38-c729-4498-af55-83cf671defc5 · inbound
Generative AI Act II: Test Time Scaling Drives Cognition Engineering MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4cdf063-283d-4c84-8323-2a416e55d54e · inbound
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98f92c9-47e8-4a08-a328-0c1cdccc6de9 · inbound
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954eea10-c87b-43cf-9ec2-6b1c8ad09be5 · inbound
Emerging Properties in Unified Multimodal Pretraining MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 703e8ac5-584c-48aa-bd45-2f71d209e61b · inbound
Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 74420e64-e9bf-490e-9bad-8518905b5848 · inbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 535af07f-bec3-4081-8631-57ba21dd4d66 · inbound
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f099e45b-584b-4429-90f8-7bfef4ee4b09 · inbound
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc616f9-4037-4362-bdce-c9897d81045c · inbound
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e5b393c-95b4-44a1-8b11-cd619c3751bb · inbound
Multimodal Tabular Reasoning with Privileged Structured Information MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b51c26-c286-46df-9016-544557ae8b24 · inbound
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545bdbd5-21ae-42ef-983b-af63ca588c00 · inbound
VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 456f258d-cdda-45e3-a431-c54fb8de5b8a · inbound
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de573936-c557-4175-a5a9-a506d29443bb · inbound
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6c5b47-e72a-42bd-911d-b726c8c2e9cd · inbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4fd369-42ff-48dd-aa16-8a3b2a588772 · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f42d94-ef92-4dc4-9138-292463bcf64b · inbound
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb502f8-5f62-4b9c-91fe-4b1e13d4cbbe · inbound
A Survey on Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b685855f-c7de-430b-a29e-3eb89d776736 · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262f9d64-69a0-4f8f-bde8-be33d8710e31 · inbound
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cdb59a1-27ed-4549-b01c-685666eba19b · inbound
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 422dbeb9-5e7f-4fe7-90ca-88f2e6ae6291 · inbound
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddbf47d8-68f0-498c-9092-dccac2e046ba · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0f5eeb8b-5a5a-402f-8a76-d3fcc9128ac7 · inbound
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7cd4543e-76bc-4aef-82f1-3d065e399684 · inbound
Meta-CoT: Enhancing Granularity and Generalization in Image Editing MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 76be5cd6-6450-49f2-94c9-332626a6307e · inbound
ZAYA1-VL-8B Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6e757b67-7e18-4a2b-886d-e421665c2846 · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 17807cb5-4354-418f-bf11-a74de5b5aafc · inbound
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d4359dae-d5c6-4551-9dd7-9321babad716 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6fce37bc-1f7d-4864-8b6a-863019af4870 · inbound
Zamba2-VL Technical Report MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f408383a-4775-42a0-928c-b751d250b4af · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53f7f1b4-06df-4382-85c3-14791267f844 · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a48d31d2-7767-4a25-86a5-01e2722079ac · inbound
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.