{"id":"e080c359-a77f-4107-b5c1-053a864ee14e","arxiv_id":"2505.10604","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new multimodal benchmark shows that vision-language models are much worse at combining object counting with spatial-relation reasoning than at either task alone.","lead":"MIRAGE is a new benchmark of 1,710 manually annotated image questions that test counting, spatial relations, and their combination in vision-language models. State-of-the-art models lose about 20 points when counting is combined with spatial constraints, making the benchmark a useful diagnostic for where models fail.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.1's claimed ~20-point Relation-to-Combination drop is contradicted by Table 1, which shows ~48-point drops; the ~20-point drop is actually Counting-to-Combination.","rationale":"The reader's conditional verdict is reasonable, but the reader's weakest_assumption identified label correctness, difficulty-tier circularity, and tiny-subset validity rather than the more direct problem I found: the headline claim in Section 4.1 contradicts the paper's own Table 1. This is a load-bearing issue because the central claim, as quoted by the reader, is numerically false. The benchmark may still be valuable, and the qualitative observation that Combination is harder than Counting is supported by the data, but the paper's stated justification for 'fundamental limitations' rests on an incorrect comparison. Since this is a fixable wording/analysis error, I recommend retaining the conditional verdict with an explicit requirement to correct the claim and propagate the correction through the abstract and introduction. I did not find evidence of dishonesty; the inconsistency appears to be a genuine error in reporting.","tokens_in":15728,"tokens_out":8299,"duration_ms":76406,"concrete_test":"Recompute the row-wise task-type differences in Table 1 for Qwen2.5VL-72B and InternVL3-78B: subtract Combination from Relation and from Counting. Compare these differences to the wording in Section 4.1, the abstract, and Section 1. If Relation−Combination is approximately 48 points and Counting−Combination is approximately 20 points, then the paper's '~20-point drop ... from Relation to Combination' is factually incorrect and must be revised, either to the correct comparison or to the correct magnitude.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim in Section 4.1 states that Qwen2.5VL-72B and InternVL3-78B show a '~20-point drop in accuracy when moving from Relation to Combination.' However, the reported data in Table 1 do not support this. For Qwen2.5VL-72B, Relation accuracy is 85.31 and Combination is 36.94, a drop of 48.37 points. For InternVL3-78B, Relation is 82.60 and Combination is 36.10, a drop of 46.50 points. The only ~20-point drops in Table 1 occur between Counting and Combination: 56.62→36.94 (19.68) for Qwen2.5VL-72B and 55.15→36.10 (19.05) for InternVL3-78B. Thus the sentence as written misreports the paper's own experimental results by more than a factor of two. The abstract and introduction repeat this 'Relation to Combination' framing, and the reader's strongest_claim is quoted directly from Section 4.1. If the intended comparison was Counting→Combination, the text must be corrected consistently; if Relation→Combination was intended, the claimed magnitude is false. Either way, the conclusion that this drop 'reveals fundamental limitations in compositional spatial reasoning' is not supported by the data as reported.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Liu et al. introduce MIRAGE, a 1,710-question visual question-answering benchmark with three task types: Counting, Relation, and Counting with Relation (Combination). The dataset draws on egocentric and web-sourced images with manual annotation, and each item is assigned a difficulty tier (Easy/Medium/Hard) using pass@64 scores from two small VLMs. The authors evaluate a set of open and proprietary VLMs, reporting that models perform best on Relation, worse on Counting, and worst on Combination, and interpret the large drop as evidence of limited compositional spatial reasoning. Additional experiments study prompt modifications, image augmentation robustness, and qualitative failure modes. The central claim in Section 4.1 is that even Qwen2.5VL-72B and InternVL3-78B 'show a ~20-point drop' moving from Relation to Combination, revealing fundamental limitations in compositional spatial reasoning.","tokens_in":15952,"tokens_out":7242,"duration_ms":60884,"significance":"The proposed MIRAGE benchmark addresses a real and under-served aspect of VLM evaluation: compositional combination of counting and spatial relations. The task decomposition is conceptually clean, and the paper states that code and data are released, which would make it a reusable diagnostic instrument. If the numerical reporting and statistical support are corrected, the qualitative finding that combination tasks are markedly harder than either component alone across a range of open models is plausible and consistent with the pattern in Table 1. The diagnostic studies (prompting, augmentation, failure modes) are also potentially useful. However, the present version contains direct contradictions between prose and tables that affect the headline claim, and the validity evidence for the tiny subset and for label quality is not yet sufficient.","major_comments":[{"comment":"Section 4.1 states that Qwen2.5VL-72B and InternVL3-78B show a '~20-point drop in accuracy when moving from Relation to Combination.' Table 1 reports for Qwen2.5VL-72B Relation=85.31 and Combination=36.94, a 48.37-point drop, and for InternVL3-78B Relation=82.60 and Combination=36.10, a 46.50-point drop. The only ~20-point drops in the table are from Counting to Combination: 56.62 to 36.94 (19.68) and 55.15 to 36.10 (19.05). This misreport of the paper's own central result by more than a factor of two must be corrected; if the intended comparison is Counting to Combination, the prose and any related framing must be adjusted consistently.","section":"Section 4.1 / Table 1"},{"comment":"The caption of Table 1 and the text in Section 4.1 claim that 'performance trends on the tiny subset are consistent with those on the full benchmark, making it a reliable proxy.' The reported numbers contradict this. For Qwen2.5VL-3B, full-set accuracy is Counting 38.33 vs Combination 23.83, while tiny accuracy is Counting 30.00 vs Combination 40.00, reversing the ordering. For InternVL-3-8B, the full set gives Counting 44.24 > Combination 29.24, but the tiny set gives 38.00 = 38.00. Since all proprietary models are evaluated only on the tiny subset, the conclusions drawn from those models rest on an unvalidated proxy; a quantitative validation (e.g., rank correlation, per-task confidence intervals) or a reassignment of the tiny subset is required.","section":"Section 4.1 / Table 1 caption"},{"comment":"Section 4.2.1 says 'both prompt modifications lead to consistent gains over the zero-shot baseline.' Table 2 shows the opposite for the Combination task: baseline 29.24, one-shot 28.52, prompt engineering 30.69. The following sentence adds that prompt rewriting is 'particularly on the more challenging combination task,' but its gain over baseline is only 1.45 points while one-shot loses 0.72 points. The text should either be revised to reflect the task-specific direction of the effects or the analysis should be restricted to tasks where the gains are consistent.","section":"Section 4.2.1 / Table 2"},{"comment":"The NeurIPS checklist states 'We report the confidence interval in Section 4,' but Section 4 and Table 1 contain no confidence intervals, error bars, or significance tests. The main comparisons, including the claimed ~20/48-point drops and the tiny-vs-full consistency, require some measure of sampling variability, especially given the small combination subset (276 questions) and the 50-item tiny subset. Add bootstrap confidence intervals or an equivalent, and either remove or correct the checklist statement.","section":"Section 4 / NeurIPS Checklist Item 7"},{"comment":"Section 3.2 says all 1,710 questions were 'manually annotated and underwent a rigorous review process,' but no annotation protocol, number of annotators, or inter-annotator agreement is reported. Since the benchmark's value depends on ground-truth labels being unambiguous and correct, and since systematic label noise would directly affect the reported model accuracies, the paper should report at least a sample-based inter-annotator agreement (e.g., Cohen's kappa or percentage agreement) and describe how disagreements were resolved.","section":"Section 3.2"}],"minor_comments":[{"comment":"Difficulty tiers are defined by pass@64 of InternVL-2.5-4B and Qwen2.5VL-3B; this makes statements such as 'state-of-the-art models fail on Hard questions' partly definitional for those model families. Consider framing difficulty labels as 'model-relative' rather than absolute.","section":"Section 3.2.2"},{"comment":"The introductory sentence refers to 'the same counting tasks,' but Table 3 reports 'Counting with Relation accuracy' (30.94, 24.82, 28.42). The text and table should use the same task name.","section":"Section 4.2.2 / Table 3"},{"comment":"Several proprietary entries are hard to parse (e.g., '-48.00- 58.00 - 50.00' for Claude-3.5-sonnet); use separate columns or explicit null markers.","section":"Table 1"},{"comment":"The checklist gives conflicting answers about asset release: item 4/5 state code and data are released, while item 13 says 'This paper does not release new assets.' Resolve this inconsistency and, if assets are released, include a license.","section":"NeurIPS Checklist"},{"comment":"Appendix B.1 lists a 'Two-Stage Prompting' strategy but Table 2 and the text refer to 'One-shot' and 'Prompt Engineering'; clarify how these map to the two-stage and direct-sampling conditions, and provide the exact one-shot prompt used.","section":"Appendix B.1"},{"comment":"Appendix A lists images from Weibo, Taobao, Baidu, Xiaohongshu, 500px, and Google Images without license or terms-of-service information, despite the checklist's claim that licenses are mentioned. Please add this information or state explicitly that it is not available.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the NeurIPS checklist contains several claims that do not appear in the manuscript (confidence intervals, released assets, single-blind submission). This inconsistency should be resolved in revision. The paper also appears to be a NeurIPS submission while the arXiv version lists author names; the single-blind statement is unusual and should be checked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is worth a serious look, but the authors have misreported their own headline result by more than a factor of two. The stress-test note is correct. Section 4.1 says Qwen2.5VL-72B and InternVL3-78B show a ~20-point drop going from Relation to Combination. The table shows a ~48-point drop for both. The ~20-point drop is actually Counting to Combination. If the intended claim is about compositional tasks, the direction of the comparison matters: Relation is the best-performing task, Counting is mid, Combination is worst. Saying Relation→Combination overstates the compositional difficulty compared to Counting→Combination, but the underlying qualitative finding—Combination is hardest—stands up across all models in Table 1.\n\nWhat is actually new: a hand-curated, multi-source benchmark with 1,710 questions across Counting, Relation, and Counting-with-Relation, plus a 50-item tiny subset meant for fast iteration. Combining the two skill types in one question is a modest but real extension over prior spatial benchmarks, and the task design is reasonable. The paper also includes prompt and augmentation ablations that add some diagnostic value.\n\nThe soft spots are the usual benchmark-paper issues, and they are fixable. No error bars or confidence intervals appear anywhere despite the NeurIPS checklist saying they do. Difficulty tiers are defined by pass@64 on two Qwen/InternVL models, so claims that 'models fail on hard questions' are partly circular for those same families. The tiny subset is called a reliable proxy but there is no statistical validation; some numbers look noisy (Qwen2.5-3B goes from 23.83% on full Combination to 40% on tiny). No inter-annotator agreement is reported for the manual labels, and there is no training-data leakage check for the web-sourced images. None of these are fatal; they are standard missing evidence for this kind of contribution.\n\nWho this is for: anyone evaluating VLMs for spatial reasoning or building diagnostic benchmarks. The paper deserves peer review—the benchmark is useful, the main claim is likely true, and the flaws are correctable. But the numerical inconsistency in Section 4.1 needs to be fixed before publication, along with a real statistical treatment of the results.","headline":"A useful but under-polished benchmark paper whose headline number is misreported by more than a factor of two; the main qualitative finding survives, but the paper needs factual corrections and statistical rigor.","tokens_in":16517,"tokens_out":2401,"would_cite":false,"duration_ms":22355,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MIRAGE shows that state-of-the-art vision-language models lose about 20 accuracy points when they must count objects under spatial constraints.","keywords":["MIRAGE","vision-language models","spatial reasoning","object counting","compositional reasoning","benchmark","visual grounding","relation reasoning"],"falsifier":"Construct a held-out set of 500 freshly photographed scenes using the same counting and relation templates, have multiple independent annotators label them, and test the same models; if human disagreement is high, or if the Relation-to-Combination gap drops to a few points on fresh images, the reported limitation would be partly an artifact of label noise or data contamination.","tokens_in":15492,"feed_emoji":"🧩","tokens_out":6434,"duration_ms":59234,"temperature":0.7,"pith_summary":"MIRAGE is a benchmark built to test whether vision-language models can count objects, understand spatial relations such as 'left of' and 'above', and combine the two in one question. The paper's central claim is that state-of-the-art models—including 72B and 78B parameter systems—handle isolated counting or relation tasks reasonably well but lose roughly 20 accuracy points when a question requires both, as in 'How many objects are to the left of the kettle and above the red container?' If that claim holds, current models lack compositional spatial reasoning rather than simply lacking perception or language skill. The benchmark also exposes occlusion, dense scenes, and ambiguous spatial referents as systematic causes of failure, and it offers a 50-question tiny subset designed to reproduce the full ranking cheaply.","feed_headline":"Big vision models lose 20 points when counting meets spatial relations","feed_subtitle":"A new 1,710-question benchmark shows top vision-language models fail when counting must respect spatial rules.","key_machinery":"The load-bearing object is MIRAGE itself: 1,710 questions with three task types—Counting, Relation, and Counting with Relation—each paired with a JSON label and assigned to a difficulty tier. Difficulty is assigned by rule-based consensus using pass@64 on two small models, InternVL-2.5-4B and Qwen2.5VL-3B: a sample is Hard when both models succeed fewer than 2 times, Medium when they succeed 2 to 16 times, and Easy when either succeeds more than 16 times. That stratification is what lets the benchmark separate task complexity from model scale, and it turns the Relation-to-Combination gap into a single comparable number across models. The 50-question Tiny subset is designed as a fast proxy that preserves the performance ordering of the full benchmark.","core_discovery":"The paper introduces MIRAGE, a manually annotated set of 1,710 image-question pairs split into 680 Counting questions, 754 Relation questions, and 276 Counting-with-Relation questions, with images gathered from egocentric video, web search, stock photography, and original photos. On the full benchmark, Qwen2.5VL-72B scores 56.62% on Counting, 85.31% on Relation, and 36.94% on Combination, while InternVL3-78B scores 55.15%, 82.60%, and 36.10%, a roughly 20-point drop that the paper calls evidence of fundamental limitations in compositional spatial reasoning. Diagnostic experiments show that adding one exemplar or rewritten instructions helps only modestly, that simple horizontal or vertical flips lower Combination accuracy by about 6 points, and that reasoning-style prompts can introduce fluent but visually unsupported hallucinations. The authors conclude that the bottleneck is visual grounding and spatial invariance, not instruction comprehension alone.","pith_inferences":["An implication the authors leave implicit: the benchmark could be converted into a training set, and the decisive test of compositionality would be whether fine-tuning on Combination questions closes the gap on held-out scenes.","Because the image pool includes web and social-media sources, some images may already appear in pretraining corpora; the reported gap might shrink on a strictly out-of-distribution set of freshly captured photos.","Difficulty tiers are calibrated to two small models, so as models improve the tiers will need re-normalization; otherwise 'Hard' will silently become 'Medium'.","The Relation-to-Combination drop is a cheap, model-agnostic diagnostic that could be tracked during architecture development as a proxy for compositional grounding."],"forward_implications":["Applications that assume spatial competence, such as robot manipulation, navigation, or augmented reality, should not rely on current vision-language models for count-within-spatial-constraint queries.","Prompt engineering yields small gains, but flips and noise still break performance, so fixes are more likely to come from spatial representations than from instruction tuning.","Reasoning-style prompting is a double-edged tool: it improves grounding in easy cases and increases hallucination risk in ambiguous ones.","The 50-question Tiny subset is asserted to track full-benchmark ordering, making it usable for fast iteration during model development.","The 13-point gap between the 3B and 72B Qwen models on Counting shows scale helps, but not enough to close the compositional gap."],"supporting_citations":[{"why":"Supplies egocentric kitchen-scene images that form one of the benchmark's data sources.","marker":"[6]"},{"why":"Underlies the Qwen2.5VL-family results, including the 72B model whose Combination score anchors the main claim.","marker":"[20]"},{"why":"Underlies the InternVL3-family results, including the 78B model used for the headline Relation-to-Combination drop.","marker":"[26]"},{"why":"Prior grounded spatial-reasoning benchmark that MIRAGE extends by combining counting with relations.","marker":"[18]"},{"why":"Prior spatial benchmark showing models struggle with 3D spatial understanding, motivating MIRAGE's 2D compositional focus.","marker":"[8]"},{"why":"Prior evaluation benchmark documenting counting and attribute-understanding failures that MIRAGE targets.","marker":"[22]"},{"why":"Prior evidence that models are poor at grounded counting, a baseline result MIRAGE reproduces and refines.","marker":"[25]"},{"why":"Spatial-temporal benchmark cited as the downstream goal that static compositional reasoning is meant to support.","marker":"[14]"}],"fun_headline_variants":["MIRAGE exposes 20-point AI drop when counting must obey space","Top vision models lose 20 points on spatial counting combo","Spatial reasoning gap: AI drops 20 points on combined tasks","Counting plus space stumps best AI models by 20 points","New MIRAGE benchmark reveals AI's spatial counting weakness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's conclusions rest on the assumption that the 1,710 manual labels are correct and unambiguous, that the difficulty tiers set by two small models are fair, and that images scraped from the web are not already memorized by the tested models during pretraining.","fun_headline_variants_meta":{"raw":{"variants":["MIRAGE exposes 20-point AI drop when counting must obey space","Top vision models lose 20 points on spatial counting combo","Spatial reasoning gap: AI drops 20 points on combined tasks","Counting plus space stumps best AI models by 20 points","New MIRAGE benchmark reveals AI's spatial counting weakness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000738,"raw_usage":{"total_tokens":3259,"prompt_tokens":873,"completion_tokens":2386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":2298}},"tokens_in":489,"tokens_out":2386,"duration_ms":17047,"temperature":1.0,"reasoning_tokens":2298,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:08:49.130129+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a held-out set of 500 freshly photographed scenes using the same counting and relation templates, have multiple independent annotators label them, and test the same models; if human disagreement is high, or if the Relation-to-Combination gap drops to a few points on fresh images, the reported limitation would be partly an artifact of label noise or data contamination.","supporting_citations":[{"cited_title":"image_caption","cited_arxiv_id":null,"evidence_quote":"Underlies the InternVL3-family results, including the 78B model used for the headline Relation-to-Combination drop."},{"cited_title":"Gsr-bench: A benchmark for grounded spatial reasoning evaluation via multimodal llms, 2024","cited_arxiv_id":null,"evidence_quote":"Prior grounded spatial-reasoning benchmark that MIRAGE extends by combining counting with relations."},{"cited_title":"Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models, 2023","cited_arxiv_id":null,"evidence_quote":"Prior evaluation benchmark documenting counting and attribute-understanding failures that MIRAGE targets."},{"cited_title":"Good at captioning, bad at counting: Benchmarking gpt-4v on earth observation data, 2024","cited_arxiv_id":null,"evidence_quote":"Prior evidence that models are poor at grounded counting, a baseline result MIRAGE reproduces and refines."},{"cited_title":"Sti- bench: Are mllms ready for precise spatial-temporal world understanding?, 2025","cited_arxiv_id":null,"evidence_quote":"Spatial-temporal benchmark cited as the downstream goal that static compositional reasoning is meant to support."}],"review_version":1}