Pith. sign in

Paper Citation Record · LEDGER

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

As of 16 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2506.23563.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23563 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:41.103071Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:45.693853Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:04:22.765546Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be9ea60a-6bb1-447f-9f50-bc46dd5b167e · outbound

This paper cites Vqa: Visual question answering.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vqa: Visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:36.964984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:36.964984Z digest=sha256:54d7ae526f95cba86b174bd911149b7ff22fb6b5a90fde5845f82c24432cb99b

Observation de0012f4-cfd3-446e-bb25-da7f0447338f · outbound

This paper cites Qwen2.5-VL Technical Report.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.018974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.018974Z digest=sha256:403bbaef1bb77e5d80d9a7080db5249bab952dfc407dcf5d03fdea444d238a1f

Observation f03e35c7-15c6-4412-a186-6ba364de3669 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.060542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.060542Z digest=sha256:4dcfd810f33170308854449ace431db497dbec24147d618c3f1768a3111457ff

Observation 36ffbdb7-f622-4f64-8ec5-c54f06370650 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision- language models with less than $3.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-v: Reinforcing super generalization ability in vision- language models with less than $3

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.107652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.107652Z digest=sha256:bf5a37f8c37f9097a86e758e7de5521f9230ce12812395e5978fac22da1b6a1b

Observation 7730835c-196e-46e4-8e49-79513dffde51 · outbound

This paper cites M 3cot: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI M 3cot: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:44.435523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:37.181858Z digest=sha256:6742a9275ef4244c0812586d83350eb23e8dfde7425c1e8815e9e07c372be57e

Observation 2c7e9642-a714-4ae7-96d4-65bf223d4959 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.266368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.266368Z digest=sha256:d6fd2c6e3be89a091ca802222579818215184d0221a7431fa96283cc1bacdc16

Observation a59ed741-d1c7-4846-9d49-55986827a34e · outbound

This paper cites The Llama 3 Herd of Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.363908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.363908Z digest=sha256:85725d2aca6cf57ad3b0ef574fc60fe1fdd7a81fc442ac5e67c01ae4861152a7

Observation 9ebdc35f-8e27-473e-997c-4698bd2eadc0 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:44.182233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:37.416896Z digest=sha256:e8cb63b3568b6c8314adf752f526c7cfbba978c08d3feff0cd11a47ef7137d4f

Observation 93ab7eea-85ef-4eb5-9497-ead8a6c49b62 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.956530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:37.474185Z digest=sha256:5629f5301209f36605d3d92f34ed9b8206e6ff72b8b70c97572572651f2fb00a

Observation a9715380-7d87-4d83-986f-97ae974160d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.543852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.543852Z digest=sha256:e3cc3a1bcb083f0b41d52d56ecb0a7a6a1528a79af9c797d24864fe8476e9714

Observation 77e758fd-1c0b-4b84-ad84-7a408c596d97 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.636240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.636240Z digest=sha256:b2a4af5bb868f910658935e26825a8e141efd2601182eeeca780fa8215d40297

Observation fb28a5cd-421c-445f-9007-ff80bac185c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.734988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.734988Z digest=sha256:acbce4a1a7d85230548e478cfba11744b7f967a26a25fd3f982967d780530b7e

Observation 5b0541e0-2129-48eb-9b5d-c0aaf865050a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.820286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.820286Z digest=sha256:f9dbb7d80f9a5223a7d72b187c60e66f58134d11f672ccc985c0641a2a74a6c3

Observation 0a995200-d266-4af7-b4d1-4c749bb2f0a5 · outbound

This paper cites GPT-4o System Card.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.882424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.882424Z digest=sha256:850f76df4859a5abc083a6cb8a5423f0a302da44bb41d61a5aff528a59634e20

Observation a1e9aacb-9249-43fa-a880-9ded340d290c · outbound

This paper cites OpenAI o1 System Card.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.947664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.947664Z digest=sha256:138bf9838efe95b563ac3fa9781f10d41cdcb4cae2566c285aa7f12fb522c4cf

Observation 6d2fc3ba-87a4-408e-b6d2-28629dc2b5da · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.034436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.034436Z digest=sha256:b3db80f04ee1d9b656fa8a917e49e53ece28d25c6f5e76f32279530e76bb11af

Observation ce92bb88-2148-40fe-9df0-df9730d36ce9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.107966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.107966Z digest=sha256:ed2a0e10e3323bdf4db223bf4439ce596918a8b0a197a3098ccc7c988c32c176

Observation e47b29ba-02f9-44ef-b388-744b0b1acbe1 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.218393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.218393Z digest=sha256:b8e04024868d1d900c902bd5fd9e311f7bac59567b009cbf66259717e54390b6

Observation 29fc1abd-5a76-480e-a5ca-aa060e71d691 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.323087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.323087Z digest=sha256:78eadf0938600a32879bddd0d0ac9017330612b669967b1fbe0bcfba97f4ee81

Observation fc37b8db-903b-4034-bbb2-e83c62082ef2 · outbound

This paper cites DeepSeek-V3 Technical Report.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.393410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.393410Z digest=sha256:fecbd4a62324891e536ce2abff86b2653a30d074076942e50c77c0aec3af0a17

Observation c142dbff-ded5-415c-9b20-b7d4a1992ca8 · outbound

This paper cites X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:41.520575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:38.463963Z digest=sha256:06fc32ea51944929d457cdf8dc2057bd487779fe9d6ca67b87fc5810e644b200

Observation 0000f601-0756-4c2f-8a0f-7a2471a0e18f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vi- sion, pages 216–233.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mmbench: Is your multi-modal model an all-around player? In European conference on computer vi- sion, pages 216–233

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.690894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:38.513018Z digest=sha256:5bab31a245a132c6b21087fc9837a1bce7a214c3b3502ff61220d06557246295

Observation 0e6e344c-68b5-455a-9592-b135010773f9 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.579018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.579018Z digest=sha256:f13dfd8e3aa855a6158c20aa151290f98a9c7da53e7b6b383e423ca7cbbc1654

Observation cf81a703-ce5f-46bb-910c-e78427fd777b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.631494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.631494Z digest=sha256:3fd4bf27a86a5032cf7cafbeb93c9147692f41d0922b52539bb6037c79379783

Observation 5144df1c-c94b-4336-8a75-5c547156bb85 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.452752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:38.777914Z digest=sha256:77ec48adbf596f08687669a571287017ea1a3ca019e6dfbba66281a8d8817b46

Observation 0f2f5a40-523a-4187-a9dd-eff88897d922 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.843269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.843269Z digest=sha256:a8ba2aaa8ea07a3dc8511178dfeb73a7d70a03ceba568bda27dcd291c45b601e

Observation 8f337445-6d0b-4852-9546-0e8596ac8d34 · outbound

This paper cites Claude 3.7 sonnet, 2025.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Claude 3.7 sonnet, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.186623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:38.930271Z digest=sha256:2efe95384b2ac13be8ae6d5939a8e81cb2c96e9cda3c85c97e98d0de7e68a370

Observation 38f60056-f59a-4e94-941d-da25dc558f65 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.038546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.038546Z digest=sha256:635c8dba494ed1506ec4b735001a5d9ed27a974a6473c1be24c17bc05e73cd1f

Observation 864c0936-66c9-4841-a69e-43aa9db98b84 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.192570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.192570Z digest=sha256:f94a228ad7b14cb761bbf15d6408783a32eb3ec3673dfddc18789a5ad4e4ae3c

Observation 33880566-25cf-42b6-8ffc-4764fa08ec7b · outbound

This paper cites Qvq: To see the world with wisdom, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qvq: To see the world with wisdom, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.896929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:39.353962Z digest=sha256:95ce2bcd19f296ef791aff71ee7b68d7b4f71c6c25293249f0a7e64622215cea

Observation b36a3b7a-3683-4148-9a7a-dd9fd47ab33c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qwq: Reflect deeply on the boundaries of the unknown, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.657113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:39.503756Z digest=sha256:7598e532b7007f4fd303f0216ef2d0ec165df6a8aa372b92327f133645c975bb

Observation 944a66d3-3182-42f5-bbbd-ff2df9985f9f · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.658744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.658744Z digest=sha256:6defced07ecc076ad27175c8d453ba53982f2a4b2ce6ebe9915babb15e5415fb

Observation dc1e6874-59e1-4d60-ac44-28e7c5356027 · outbound

This paper cites Open-r1-video.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Open-r1-video

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.440908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:39.758478Z digest=sha256:de672c39d12649a81569f97e79fdc40d38b5a36e2190134c2464c9c08c27ba82

Observation d176d38e-f478-4a60-859b-b22c4b999294 · outbound

This paper cites VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.855190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.855190Z digest=sha256:42d3b88cab38be912e9d7bd0157edcaaebf8c338a8c43849d8753e322356c60e

Observation d763cf7c-4589-4b5c-be57-5cf424486580 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.215107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:39.912917Z digest=sha256:18c84af417a2a81cf0fb86595b3e77b303acc46da89a873c4e1e407aa3c719bc

Observation 1e22d933-8446-44f4-b425-e1ea0b1ed865 · outbound

This paper cites Boosting mul- timodal reasoning with mcts-automated structured thinking.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Boosting mul- timodal reasoning with mcts-automated structured thinking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.982779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.982779Z digest=sha256:3ac1071e330b95c3b6f565da972dcd1027965b77910ee13c5308324f7e330261

Observation 47746207-6c7c-4ee6-8b2c-07e227d36c45 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.067346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.067346Z digest=sha256:2125e449924983f73e81f05c5503574c073f382aa5f48b43caffb75552e5cfde

Observation 58b2ef09-753d-40cc-9c41-ae15c01fabd7 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.126753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.126753Z digest=sha256:60d6767be908611482302c329bb07b116108f290167b7b86d22749e13d8d3bfc

Observation ea48561d-de45-475c-a596-68d921cd0686 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.221780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.221780Z digest=sha256:ad25f4109b094965cb06c1016715df9ff9e3c7371e74612137a15878f9305374

Observation fec71ebb-790b-405a-aa09-33befdc3f34a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.304579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.304579Z digest=sha256:a10a5c69e5f6b1258cd7f6c5bf6344c724d033a84dc769aca2e86e9b28ebb49b

Observation 2aacdaa0-ba2e-4fb9-901f-e0efd17ce28d · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.361835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.361835Z digest=sha256:25d3175c0007f2b0c70a387829f7082b88889fec67cf9b60d7781d1741b6f975

Observation 3559689b-0da9-42f6-94d4-496b6fe9deb6 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.429654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.429654Z digest=sha256:1b590403e3a528d9fd81d6400ea3fa3907038b63de5eff7c350fd02e5ff899f7

Observation 9ea9b439-fd49-40a5-9db9-5e1f2a8f0b91 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.493129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.493129Z digest=sha256:6384c4cc60afab88c143e93501e202168303b50dea5c59244bda29b95cd0eec2

Observation cf7dae34-fce7-45c6-bffe-2d2f5dfb778b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.565270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.565270Z digest=sha256:6fd688c791c34d7cea15133e5f65e739dce6fb20bbb425bb6214860d87b4a65a

Observation 19eed217-3bfb-42ce-86d9-feddd144e924 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.066729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:40.624530Z digest=sha256:10a5673ef9f6186275d14273691770fe8534073dd159bb0254d9246f80c618ea

Observation 4d8d3328-49e0-4ce5-b281-1c03ffe00383 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.702164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.702164Z digest=sha256:1819dcc0c2fec18ce24d91b18bbbca0a947aa84425dd0ccd5d0247ea9ab133e5

Observation fe7fe619-ec37-48b2-b76d-529d226dc326 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.775733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.775733Z digest=sha256:b25a9adfe33a5240d7ee403dc034c07802eb9bce1c3f3fdf9fe3b337d84b5df6

Observation ee42f373-7c73-4b4d-b7db-39c4af95d126 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.883389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.883389Z digest=sha256:bdf019e9fad86f6d4b8d1c65003296121e84c661df988ec25803395ed7484a84

Observation 1fcd04ec-7e36-4357-9dbc-10dbab8f295e · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:41.899167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:41:40.941360Z digest=sha256:daaa08b00a65308ece48cb1466e327d6216d32df1e0a29c3e60afe38fbabdbac

Observation 9ca8fd4c-e733-4ed2-ab5d-dc3e9595b0a2 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Improve Vision Language Model Chain-of-thought Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:41.040789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:41.040789Z digest=sha256:8c7d008550c27dd7fa95e77023b59bcf73fb056cb21a3cd69aa2351474f4a645

Observation 68b418f4-f233-4f31-b187-9a27584609bf · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:41.103071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:41.103071Z digest=sha256:c79d70f3e87b50c723280ffd413a8e03028464c2aceab32df00f0d49bab5a6ce

Pith citing papers

Observation 41cb0031-bb58-4119-abee-969cf25cdbb7 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.768219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:f6e54534df2878499d33948a9e0f50f3fba5c83ce237baa50d0f8745c794a2c5

Observation 54e58cba-4bf8-433d-a37c-d4188b30aaea · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:45.693853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:45.693853Z digest=sha256:5a4ce4d35fe26662e0e03d2835a48664538364830a62501244dd214d63ce94dc

Observation 8c7ca8d1-37a3-4a88-9c36-c51ee1825ca8 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.540546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:67ac47c9424bada1413bb5a9d3a07a02753471b38fadb749f8630c580b0ce8dc