{"id":"3d5101f8-12b9-4b7c-8dda-a685307e7e67","arxiv_id":"2505.12766","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reasoning-OCR is a new bilingual benchmark of 150 image-based logical reasoning questions; the best tested model, GPT-4o, scores 68.1% and open-source models stay below 63%.","lead":"This paper introduces Reasoning-OCR, a 150-question benchmark that tests whether large multimodal models can perform multi-step logical reasoning from text visible in images such as charts, labels, and documents. It evaluates 10 models and finds they often fail, especially on decision-style questions, suggesting a gap between reading text and reasoning about it.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that decision reasoning is the hardest category rests on an unvalidated 11-question subset; one ambiguous item can flip the ranking.","rationale":"The reader's CONDITIONAL verdict already captures this concern, and my stress-test does not move it. The benchmark is a genuine contribution: the overall result that strong open-source LMMs drop from roughly 93% on DocVQA to 47.8% on Reasoning-OCR is credible evidence that complex OCR-cue reasoning is not saturated. What is not established is the more specific quantitative ranking across reasoning types, especially the claim that decision reasoning is the weakest category. That claim is computed from about 11 items with no demonstrated ground-truth uniqueness and no human ceiling. The proposed annotation and confidence-interval check would settle whether the per-type findings are artifacts. This is an evidential gap, not an internal inconsistency, so a conditional accept rather than a reject is appropriate.","tokens_in":21139,"tokens_out":11039,"duration_ms":107634,"concrete_test":"Recruit at least two independent annotators who are not authors to answer all 150 questions and to flag any question with more than one defensible answer, unclear wording, or an incorrect ground truth; compute per-question agreement and human accuracy. Then recompute Table 2 (a) excluding flagged questions and (b) with 95% exact binomial confidence intervals per reasoning type. If the decision-reasoning gap between GPT-4o and the 9.1%-scoring models remains significant after filtering, and if human accuracy is high (e.g., at least 90%), the headline concern is resolved; otherwise the per-type comparative claims should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 3.1 asserts that all questions are 'objective-type' with answers 'concise and unambiguous,' but the paper reports no independent annotation study, no inter-annotator agreement, and no human accuracy baseline; the datasheet (D.3, A4) states the data were 'collected and verified by the authors.' The headline quantitative claims therefore depend on an unmeasured assumption that the intended answer is the only defensible one. The 'decision reasoning is hardest' finding (Finding ❸, Table 2) is especially fragile: Fig. 3(a) assigns decision reasoning 0.073 of 150 questions, i.e. about 11 items. GPT-4o's best 36.4% is 4/11, and a 95% Wilson interval for 4/11 is roughly 15%-65%, overlapping intervals for models scoring 9.1% (1/11) and 18.2% (2/11). One ambiguous or mislabeled decision question changes the ranking by about 9 percentage points. Without a human ceiling, low model accuracy could also mean the questions are hard for humans rather than that LMMs are deficient.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Reasoning-OCR, a benchmark of 150 bilingual (English/Chinese) questions over 140 text-rich images drawn from ChartQA, DT-VQA, DocVQA, and the web, covering six visual scenarios and six reasoning types (data comparison, data statistics, mathematical, conditional, temporal, and decision reasoning). The authors evaluate ten large multimodal models, including GPT-4o and nine open-source models, under five answering settings (CoT, no-CoT, cross-linguistic, with hint, and task-specific instruction). They report that all models are far from ceiling, that text-centric LMMs lag behind generic LMMs, that decision reasoning is the hardest category, and that CoT and provided hints generally improve accuracy. The benchmark and code are promised to be publicly released.","tokens_in":21352,"tokens_out":5533,"duration_ms":61042,"significance":"If the benchmark is validated, it fills a real gap: existing OCR and visual-text benchmarks mostly test text extraction and simple VQA rather than multi-hop logical reasoning from OCR cues. The design choices are sensible, including diverse scenarios, six explicit reasoning categories, bilingual questions, and multiple inference settings. The qualitative conclusion that current LMMs are far from ceiling on such questions is plausible and would be useful to the community. The paper also provides a detailed error analysis and makes the benchmark, instructions, and code available, which are concrete strengths. However, the quantitative and comparative claims currently rest on 150 unvalidated questions with no inter-annotator agreement, no human accuracy baseline, and no confidence intervals; several internal inconsistencies further undermine the reliability of the reported numbers. The central qualitative finding is likely robust, but the finer claims about relative difficulty across reasoning types and about text-centric versus generic models need additional support.","major_comments":[{"comment":"The benchmark validation is not reported. Section 3.1 asserts that all questions are objective-type with answers that are 'concise and unambiguous,' but the datasheet (D.3, A4) states only that the data were 'collected and verified by the authors.' There is no inter-annotator agreement, no independent annotation study, no pilot validation, and no human accuracy baseline. This matters because every accuracy number in Tables 2 and 3 depends on the assumption that the intended answer is the only defensible one. Without a human ceiling, low model accuracy could partly mean the questions are hard for humans rather than that LMMs are deficient. I request a human evaluation on the full set (or a justified sample), a report of ambiguous or multi-answer questions, and a discussion of how synonymity and format variations were handled.","section":"§3.1 and D.3"},{"comment":"The claim that decision reasoning is the hardest category rests on about 11 questions (the 0.073 proportion in Fig. 3(a) corresponds to ~11 of 150). In Table 2, GPT-4o's best accuracy of 36.4% is 4/11, and the 95% Wilson interval for 4/11 spans roughly 15%-65%, overlapping the intervals for models scoring 9.1% (1/11) and 18.2% (2/11). One ambiguous or mislabeled decision question changes the category accuracy by about 9 percentage points, which can flip the ranking of models and even the conclusion that decision reasoning is the most difficult. The paper should report confidence intervals, increase the number of decision-reasoning questions, or substantially soften Finding ❸.","section":"§4.3, Finding ❸, Fig. 3(a), Table 2"},{"comment":"There are internal annotation inconsistencies that call the data quality into question. In Fig. 1, the English question says 'early March 2023' while the Chinese version says '2003年3月初' (early March 2003); since the evaluation includes cross-linguistic reasoning, a date discrepancy can change the correct answer. Additionally, Fig. 9 is captioned 'An example for mathematical reasoning' but its question is a temporal-reasoning train-ticket problem, and Fig. 10 is captioned 'An example for temporal reasoning' but contains a mathematical expense-sum question. These issues suggest that the bilingual questions and the type labels were not carefully audited, which is load-bearing for a benchmark whose purpose is precise evaluation. The authors should correct these errors and describe a systematic consistency check for all 150 items.","section":"Fig. 1, Figs. 9–10"},{"comment":"GPT-4o is used as the answer extractor for all models, including GPT-4o itself. The paper follows prior work in doing this, but no evidence is provided that the extraction is unbiased or accurate. If GPT-4o is more lenient toward answers that match its own output format or if it silently normalizes incorrect answers, then the reported accuracies of open-source models could be affected, and GPT-4o's comparative advantage could be inflated. The authors should either measure extractor agreement against human judgments, use rule-based/string matching combined with a fixed normalization step, or report extraction-error statistics.","section":"§4.1"},{"comment":"The comparison between text-centric and generic LMMs is confounded. TextMonkey and mPLUG-DocOwl2 differ from Qwen2-VL-7B and InternVL2.5-8B not only in training specialization but also in base architecture, parameter count, and training data scale. The claim that text-centric training itself limits reasoning ability is not established by these comparisons. I suggest either adding matched-scale text-centric and generic models with comparable base backbones or recasting Finding ❷ as an observation about the specific evaluated models rather than about text-centric training as a general principle.","section":"§4.3, Finding ❷, Table 2"}],"minor_comments":[{"comment":"The abstract contains 'underscoring the urgent to improve the reasoning performance,' which appears to be missing a noun ('urgent need'), and §3.2 says 'Expect the questions in English' where 'Except' is intended. Please proofread these passages.","section":"Abstract and §3.2"},{"comment":"The column headers 'Datac', 'Datas', 'Reasoning m', 'Reasoning c', 'Reasoning t', and 'Reasoning d' are hard to parse. Please define the abbreviations in the table caption or use the full names of the six reasoning categories.","section":"Table 2"},{"comment":"The symbols ACC, ACCn, ACCl, ACCh, and ACCt are clear from the caption, but the notation should be introduced in text right before the table for readability.","section":"Table 3"},{"comment":"The Limitations section discusses scaling and scenario breadth but does not mention the absence of validation, inter-annotator agreement, or human baselines. Adding a sentence acknowledging these limitations would help readers calibrate the claims.","section":"Limitations"},{"comment":"The GPT-4o response in the 'Question Misunderstanding' example does not follow the required concise <a>...</a> format. Since the paper uses GPT-4o as the answer extractor, this example makes it particularly important to report how such verbose responses were scored.","section":"Fig. 5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear Colleague,\n\nThe short version: this paper gives the field a genuinely new evaluation instrument—complex multi-hop reasoning questions that must be answered from text embedded in images, spanning six visual scenarios and six reasoning types, with bilingual (English/Chinese) versions and hints. That combination doesn't exist in ChartQA, DocVQA, OCRBench, or MathVista. The dataset is public, the authors include a datasheet, and the evaluation covers ten LMMs across five inference settings. The qualitative finding—current LMMs, including GPT-4o at 68.1%, are far from ceiling on this task—is credible and likely holds. The appendix's attempt to get GPT-4o to generate similar questions and its failure is an honest and useful data point.\n\nWhere I part company with the paper's confidence is in the per-category quantitative claims. The dataset has 150 questions in total; 'decision reasoning' has about 11. The claim that decision reasoning is the hardest category (best 36.4%, i.e., 4/11 for GPT-4o) could shift with one or two ambiguous items. The paper reports no inter-annotator agreement, no human baseline, and no confidence intervals; the datasheet says the data were 'collected and verified by the authors.' That's not a fatal flaw, but it means the type-level rankings should be presented more tentatively. The overall accuracy numbers are more robust because the gaps are large.\n\nThere are also some sloppy details that should be caught in review: the Chinese question in Fig. 1 says 2003 while the English says 2023; the appendix captions for Figs. 9 and 10 appear swapped. These are minor, but they undercut the 'meticulously designed' claim. Using GPT-4o as the answer extractor for all models is a real concern, though standard practice; a sensitivity check with a different extractor or exact-match scoring would strengthen the results.\n\nBottom line: this paper is for people building or evaluating LMMs on text-rich images, and it deserves a serious referee rather than a desk reject. The central measurement instrument is useful and publicly available. I'd ask the authors to add a human accuracy baseline, report confidence intervals or otherwise temper the type-level comparisons, and clean up the inconsistencies. Even then, the 'hardest category' finding may not survive, but the benchmark itself would.","headline":"A genuinely new evaluation probe for complex OCR-cued reasoning, but the per-category claims rest on a small, unvalidated sample and should be treated as tentative.","tokens_in":21845,"tokens_out":3474,"would_cite":true,"duration_ms":32307,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large multimodal models are far from solving complex logical reasoning problems that require reading and interpreting text inside images, with the best model scoring 68.1% on a new 150-question benchmark.","keywords":["large multimodal models","optical character recognition","logical reasoning","benchmark","chain-of-thought","decision reasoning","visual question answering","OCR cues"],"falsifier":"Re-annotate all 150 questions with at least two independent annotators and measure agreement; then add, say, 50 new decision-reasoning questions and rerun the same models. If agreement is low, or if the best model exceeds 50% on the larger decision-reasoning set, the paper's main conclusions about ambiguity and the decision-reasoning bottleneck would need revision.","tokens_in":20966,"feed_emoji":"🧠","tokens_out":7609,"duration_ms":68153,"temperature":0.7,"pith_summary":"The paper introduces Reasoning-OCR, a benchmark of 150 hand-designed logical reasoning questions built from 140 text-rich images across six visual scenarios. The central claim is that current large multimodal models (LMMs) are far from solving these problems: the best proprietary model, GPT-4o, reaches 68.1 percent accuracy, and the best open-source model reaches 63.0 percent. The benchmark is designed to isolate reasoning from OCR cues rather than specialized knowledge, so the low scores point to a genuine deficit in multi-step logical inference over text found in images. This matters because reading text in images is a core skill for document understanding and embodied agents, and existing OCR benchmarks have largely saturated.","feed_headline":"Best multimodal model scores 68% on OCR logic benchmark","feed_subtitle":"GPT-4o still fails a third of 150 hand-crafted questions that need reading plus multi-hop reasoning.","key_machinery":"Reasoning-OCR itself is the central object: a benchmark of 150 expert-written questions over 140 images drawn from ChartQA, DocVQA, DT-VQA, and the web. Each question requires extracting textual cues from the image and combining them through multiple inference steps; the six labeled reasoning types are data comparison analysis, data statistical analysis, mathematical reasoning, conditional reasoning, temporal reasoning, and decision reasoning. The benchmark also includes Chinese question versions, a per-question hint, and a task-specific instruction template, which are used to probe whether performance changes when the model is given linguistic scaffolding. The design forces a model to do more than read: it must decide which textual cues are relevant, relate them to the question's constraints, and perform a multi-hop logical chain.","core_discovery":"The paper's central discovery is that reasoning over OCR cues is a largely unsolved capability for LMMs. On the 150-question Reasoning-OCR benchmark, GPT-4o achieves 68.1 percent accuracy with chain-of-thought, InternVL2.5-38B achieves 63.0 percent, and most open-source models stay below 50 percent. The benchmark categorizes questions into six reasoning types; the hardest is decision reasoning, where the best models reach only 36.4 percent. The paper also finds that text-centric models specifically trained for OCR lag far behind generic LMMs, that chain-of-thought consistently helps, and that giving a hint improves accuracy. These results are presented as evidence that current OCR benchmarks are saturated and that complex logical reasoning from textual cues remains a bottleneck.","pith_inferences":["Because only about 11 of the 150 questions are decision-reasoning items, the 'decision reasoning is hardest' ranking carries wide error bars; a larger decision-reasoning subset could change the ordering.","A natural testable extension is to compare LMMs against a pipeline that first extracts text with an OCR model and then performs reasoning in a pure-language model, which would separate perception errors from reasoning errors.","The benchmark's cross-linguistic condition hints at a broader question of whether LMMs reason differently when the same logic is posed in a different language, though the current sample size is too small to resolve it.","Releasing per-category confidence intervals and inter-annotator agreement would make the benchmark more useful for tracking progress over time."],"forward_implications":["Existing OCR benchmarks that report near-saturated scores may overstate progress, since the same models score below 50 percent on Reasoning-OCR.","LMMs are not yet reliable for decision-making tasks that depend on reading text, such as planning and embodied-agent instructions.","Chain-of-thought prompting, answer hints, and task-specific instructions each raise accuracy, so test-time scaffolding is a practical lever for improving OCR-based reasoning.","Text-centric OCR models should be trained with reasoning-heavy samples, not just recognition and simple visual question answering data.","The six-category breakdown offers a diagnostic: statistical and comparison reasoning are relatively stronger, while decision reasoning is weakest."],"supporting_citations":[{"why":"Supplies 60 of the benchmark's chart images and serves as the chart-reasoning reference point.","marker":"[Masry et al., 2022]"},{"why":"Supplies 50 dense-text images and is a prior OCR/VQA benchmark the paper builds on.","marker":"[Zhang et al., 2024a]"},{"why":"Supplies 20 document images and is the classic document-VQA benchmark the paper argues saturates.","marker":"[Mathew et al., 2021]"},{"why":"Provides the evaluation protocol using GPT-4o as answer extractor and a comparison for mathematical reasoning in visual contexts.","marker":"[Lu et al., 2024a]"},{"why":"The proprietary baseline and answer extractor; achieves the best overall accuracy in the study.","marker":"[OpenAI, 2024]"},{"why":"Open-source baseline family and the tokenizer used for question-length statistics.","marker":"[Chen et al., 2024b]"},{"why":"Representative OCR benchmark contrasted with Reasoning-OCR to argue that simple OCR tasks are already saturated.","marker":"[Liu et al., 2024e]"},{"why":"Text-centric model baseline whose low scores demonstrate the gap between OCR-oriented and generic LMMs.","marker":"[Hu et al., 2024]"},{"why":"Open-source generic LMM baseline evaluated across reasoning types and settings.","marker":"[Wang et al., 2024b]"}],"fun_headline_variants":["GPT-4o tops 68% on new OCR logic benchmark","OCR logic: best AI still fails 32% of questions","Reasoning-OCR: LMMs struggle with read-and-think","New OCR reasoning benchmark humbles open-source LMMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every question in the benchmark is unambiguous and correctly labeled, and that the per-category sample sizes, especially the roughly eleven decision-reasoning questions, are large enough to support claims about which reasoning types are hardest.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o tops 68% on new OCR logic benchmark","OCR logic: best AI still fails 32% of questions","Reasoning-OCR: LMMs struggle with read-and-think","New OCR reasoning benchmark humbles open-source LMMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1191,"prompt_tokens":897,"completion_tokens":294,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":221}},"tokens_in":513,"tokens_out":294,"duration_ms":3673,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:26:38.619261+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate all 150 questions with at least two independent annotators and measure agreement; then add, say, 50 new decision-reasoning questions and rerun the same models. If agreement is low, or if the best model exceeds 50% on the larger decision-reasoning set, the paper's main conclusions about ambiguity and the decision-reasoning bottleneck would need revision.","supporting_citations":[{"cited_title":"Chartqa: A benchmark for question answering about charts with visual and logical reasoning","cited_arxiv_id":null,"evidence_quote":"Supplies 60 of the benchmark's chart images and serves as the chart-reasoning reference point."},{"cited_title":"Docvqa: A dataset for vqa on document images","cited_arxiv_id":null,"evidence_quote":"Supplies 20 document images and is the classic document-VQA benchmark the paper argues saturates."}],"review_version":1}