Pith. sign in

Paper Citation Record · LEDGER

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2506.23563.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23563 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:41.103071Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:45.693853Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:04:22.765546Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be9ea60a-6bb1-447f-9f50-bc46dd5b167e · outbound

This paper cites Vqa: Visual question answering.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vqa: Visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:36.964984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:36.964984Z digest=sha256:2f1063444abda17402bae9d89c50b16f453b1f76abe34bd6a02cc1a526edf588

Observation de0012f4-cfd3-446e-bb25-da7f0447338f · outbound

This paper cites Qwen2.5-VL Technical Report.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.018974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.018974Z digest=sha256:a3dda1f871dc7b2c97c96ab326b2f2f59a8087b79044930b5cfc30b75b5c0a8a

Observation f03e35c7-15c6-4412-a186-6ba364de3669 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.060542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.060542Z digest=sha256:50de0188f39b25a789a8af4613afe85b3ea97b30725ff24fad5b11b471ef29ff

Observation 36ffbdb7-f622-4f64-8ec5-c54f06370650 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision- language models with less than $3.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-v: Reinforcing super generalization ability in vision- language models with less than $3

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.107652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.107652Z digest=sha256:01f92ec5443392e213939bb718473b35e24eb52455fd37a96a3e1fb4337431dd

Observation 7730835c-196e-46e4-8e49-79513dffde51 · outbound

This paper cites M 3cot: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI M 3cot: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:44.435523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:37.181858Z digest=sha256:9b2b059e99b71810dac1357f0344daad83ff095a8ed862ddff15ef28056dccda

Observation 2c7e9642-a714-4ae7-96d4-65bf223d4959 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.266368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.266368Z digest=sha256:d545bc51def2d0f2f346d903653d65b08ecefdc1888870a71e79b60af3023cda

Observation a59ed741-d1c7-4846-9d49-55986827a34e · outbound

This paper cites The Llama 3 Herd of Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.363908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.363908Z digest=sha256:a6f35352c256a2175f5cc234412084a07140c4e71a7f7c0cc50a114aed0297fd

Observation 9ebdc35f-8e27-473e-997c-4698bd2eadc0 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:44.182233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:37.416896Z digest=sha256:e07f212154bc9b5b61e721b714e46cabe5bea681bad0480ffb82c4292d2f5502

Observation 93ab7eea-85ef-4eb5-9497-ead8a6c49b62 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.956530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:37.474185Z digest=sha256:2422cf1a3c599910cf88fd1d42080d917a3f44b663c8da7521fc9bab0b252bda

Observation a9715380-7d87-4d83-986f-97ae974160d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.543852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.543852Z digest=sha256:f12c7738efed260a967e2c6a13ef244553ee79784be838ba862cc17fc8d35c74

Observation 77e758fd-1c0b-4b84-ad84-7a408c596d97 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.636240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.636240Z digest=sha256:bf957206677f5583fd4db7a4642d8fc7af04c85ebe21ce1e4877d87e529c6a79

Observation fb28a5cd-421c-445f-9007-ff80bac185c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.734988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.734988Z digest=sha256:cc4938f9580ac20cd8e87efd4785edf39ad951939ad10c6e1cde99d84d6ad029

Observation 5b0541e0-2129-48eb-9b5d-c0aaf865050a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.820286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.820286Z digest=sha256:50149a49efb852be3d27369b1f326bb1fe43ca3d015256935e8e91c4b52b9b17

Observation 0a995200-d266-4af7-b4d1-4c749bb2f0a5 · outbound

This paper cites GPT-4o System Card.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.882424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.882424Z digest=sha256:1aa556b3e09c79e659e6b85303cee04858d4526774540e689ecd4e9c3277a816

Observation a1e9aacb-9249-43fa-a880-9ded340d290c · outbound

This paper cites OpenAI o1 System Card.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.947664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.947664Z digest=sha256:c2948c7492677ed793e26f5f2a723dc9e98223a618619ffdee95f1f83156ac10

Observation 6d2fc3ba-87a4-408e-b6d2-28629dc2b5da · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.034436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.034436Z digest=sha256:223c316c1362f801bb93c95c55d37524c41df4e7f2084507f35bdce73121cf5c

Observation ce92bb88-2148-40fe-9df0-df9730d36ce9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.107966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.107966Z digest=sha256:d5df927fcf6eb3376ec9e19bb5c17e5680bf3106f944022bff01b125ddf64ecf

Observation e47b29ba-02f9-44ef-b388-744b0b1acbe1 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.218393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.218393Z digest=sha256:5bc2c11904402b6d87aaf1bb4bc16c509e5c91d5dbe979b9520661ca95b4ee44

Observation 29fc1abd-5a76-480e-a5ca-aa060e71d691 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.323087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.323087Z digest=sha256:1e7b61f075e6cfdbe7da1bc31dd603add6b07e76d06d55c2ad25e77c50ee6015

Observation fc37b8db-903b-4034-bbb2-e83c62082ef2 · outbound

This paper cites DeepSeek-V3 Technical Report.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.393410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.393410Z digest=sha256:2e55d2fbb33aab9fd3fe2e230c0f9cb8cbd6ac124ac05728f528350d60c6e35d

Observation c142dbff-ded5-415c-9b20-b7d4a1992ca8 · outbound

This paper cites X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:41.520575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:38.463963Z digest=sha256:1754cf0e0a1339d5cb1fbece66f1ab344e31c2d81d7dc3ca29e33c97cd9c4383

Observation 0000f601-0756-4c2f-8a0f-7a2471a0e18f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vi- sion, pages 216–233.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mmbench: Is your multi-modal model an all-around player? In European conference on computer vi- sion, pages 216–233

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.690894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:38.513018Z digest=sha256:3dfae1061506b345a08a8f68f797c91845a7607859c7f620d14483acf504e5e1

Observation 0e6e344c-68b5-455a-9592-b135010773f9 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.579018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.579018Z digest=sha256:eb128ddb9a57970b9bd58c0aa842ad9978e3e13b2242724e574153003f58e5b5

Observation cf81a703-ce5f-46bb-910c-e78427fd777b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.631494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.631494Z digest=sha256:6e64ef442c033870cefca97e388163592b3e0b553923dbe93edac556a53b3099

Observation 5144df1c-c94b-4336-8a75-5c547156bb85 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.452752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:38.777914Z digest=sha256:a42b7a67bd0f5a7570a474c8c4f703ec6cf16dde49fd479a916d4f8485e16a73

Observation 0f2f5a40-523a-4187-a9dd-eff88897d922 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.843269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.843269Z digest=sha256:19ecceae76d3ed51e623a2c53ce61789c36313d167a56e4f38df17aeaf3c7b04

Observation 8f337445-6d0b-4852-9546-0e8596ac8d34 · outbound

This paper cites Claude 3.7 sonnet, 2025.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Claude 3.7 sonnet, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.186623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:38.930271Z digest=sha256:d21fb953a309867632a3ea1cdfb75e0e99dc966cc2dbea0b881bdae8d135e02b

Observation 38f60056-f59a-4e94-941d-da25dc558f65 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.038546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.038546Z digest=sha256:99ca5a9120dd00d5ea43398c5b8af61104c24c40078201ed4435e87c9562079f

Observation 864c0936-66c9-4841-a69e-43aa9db98b84 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.192570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.192570Z digest=sha256:66e4fcab31ae932a2e6157186187cb7150c00f3ef99f837ada1609afa5ed79c7

Observation 33880566-25cf-42b6-8ffc-4764fa08ec7b · outbound

This paper cites Qvq: To see the world with wisdom, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qvq: To see the world with wisdom, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.896929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:39.353962Z digest=sha256:bbdb1d3b76bea029615f0791eebf3349c8c88df2231b18670451eb1361bd52d5

Observation b36a3b7a-3683-4148-9a7a-dd9fd47ab33c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qwq: Reflect deeply on the boundaries of the unknown, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.657113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:39.503756Z digest=sha256:dc1cf91991e197f39bd6d90aef4a9384de16edc86fc2743822dd86d500281986

Observation 944a66d3-3182-42f5-bbbd-ff2df9985f9f · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.658744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.658744Z digest=sha256:624996032f9f4215f5ce666fa03bfbc947b5d743891e71c8c8ef2e2d89c6e0f7

Observation dc1e6874-59e1-4d60-ac44-28e7c5356027 · outbound

This paper cites Open-r1-video.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Open-r1-video

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.440908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:39.758478Z digest=sha256:ad50f01396c7f95f15ee8c3043e6c724da4fc6b9e69563f3aeacd097fddc4c40

Observation d176d38e-f478-4a60-859b-b22c4b999294 · outbound

This paper cites VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.855190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.855190Z digest=sha256:42e30ae9b3127f221b94945e19b5bc6a5baf7557f0af746e9b1cce155190f500

Observation d763cf7c-4589-4b5c-be57-5cf424486580 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.215107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:39.912917Z digest=sha256:80931fdcfec2af0b563a4d73ec8b3c7b0793ca8f21957845e327ab77a46285e6

Observation 1e22d933-8446-44f4-b425-e1ea0b1ed865 · outbound

This paper cites Boosting mul- timodal reasoning with mcts-automated structured thinking.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Boosting mul- timodal reasoning with mcts-automated structured thinking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.982779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.982779Z digest=sha256:70c23b70e1d43f64da442060e72fb2b724557d1e1ef54be6d17c0dfcb24a3697

Observation 47746207-6c7c-4ee6-8b2c-07e227d36c45 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.067346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.067346Z digest=sha256:bd335fc0059b96e89ddfd768d3b793c93a0f140f13f3767a900903de1ae38050

Observation 58b2ef09-753d-40cc-9c41-ae15c01fabd7 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.126753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.126753Z digest=sha256:9e5a90a5d787ab4e74153e8538d4686fe81cb1ddbb22fd68fa8a4dbc53b9bcbd

Observation ea48561d-de45-475c-a596-68d921cd0686 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.221780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.221780Z digest=sha256:f6d0dd265d6f77f4f66a3f7f5cb875b4a1231b77c88fa9685a7650655b8a2424

Observation fec71ebb-790b-405a-aa09-33befdc3f34a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.304579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.304579Z digest=sha256:4c709e8da7ecc4e11f9ad2e4cff05b692b0127dc771fa6dfa499024dba58c489

Observation 2aacdaa0-ba2e-4fb9-901f-e0efd17ce28d · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.361835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.361835Z digest=sha256:8d872f3582d334ae2058f02f431869aeecc433c6afbe7a213ceeaf2bf92e5375

Observation 3559689b-0da9-42f6-94d4-496b6fe9deb6 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.429654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.429654Z digest=sha256:473b3b9967f8186568fd57cccc560afa15ef45960529da4e36eaf6f13f6fc86e

Observation 9ea9b439-fd49-40a5-9db9-5e1f2a8f0b91 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.493129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.493129Z digest=sha256:94d437385508f8d95223c884e6b8d44ee83397ba3e37bf5ba9d2c0c31d3b6445

Observation cf7dae34-fce7-45c6-bffe-2d2f5dfb778b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.565270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.565270Z digest=sha256:55bd25ac6be93074cc65665b48950e9da3be42d40af194e475f6b7700fc2c280

Observation 19eed217-3bfb-42ce-86d9-feddd144e924 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.066729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:40.624530Z digest=sha256:cb27da14885ebce559c8f6a30246095413f0d33e7693afe10040722b6946f635

Observation 4d8d3328-49e0-4ce5-b281-1c03ffe00383 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.702164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.702164Z digest=sha256:9f12c31bee1f29e168282fadb96d01da05d02c79a7d7a8f2a0db90f20c1e1451

Observation fe7fe619-ec37-48b2-b76d-529d226dc326 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.775733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.775733Z digest=sha256:b5433682ed0915c46bc17503c7b86afec0b1644e52db5643ff4c2da58b483015

Observation ee42f373-7c73-4b4d-b7db-39c4af95d126 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.883389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.883389Z digest=sha256:139389d1746f4cbb8edafc1ce40be66964c018d3bec071f4dde71eea894e2155

Observation 1fcd04ec-7e36-4357-9dbc-10dbab8f295e · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:41.899167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:41:40.941360Z digest=sha256:129ae8e576a0505aebaf217ae6f8f370c5e6574d2425033759f2684eefe97983

Observation 9ca8fd4c-e733-4ed2-ab5d-dc3e9595b0a2 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Improve Vision Language Model Chain-of-thought Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:41.040789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:41.040789Z digest=sha256:51cc9869caffaf21949d1c455e2d42863d5a1a08dd04750a2081a8d51a7a1d73

Observation 68b418f4-f233-4f31-b187-9a27584609bf · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:41.103071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:41.103071Z digest=sha256:e3639ea3e45ddf2cf9bd462917fea010c15e4f3ec9f93060aff371172d635f26

Pith citing papers

Observation 41cb0031-bb58-4119-abee-969cf25cdbb7 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.768219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:d19baeac651e9d9be8fc3035200277be3de96882a798ee3c35cddd18011a075f

Observation 54e58cba-4bf8-433d-a37c-d4188b30aaea · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:45.693853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:45.693853Z digest=sha256:e577639ce030b532df5e240e8828950ea6d18e7ba7f19b5d971413ba8d6af612

Observation 8c7ca8d1-37a3-4a88-9c36-c51ee1825ca8 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.540546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:7ea9ded0c4cf12c3ce71b3567f058e9fc2264e25ef83df1d251aa7fdf5f49376