{"id":"f8ff400f-75e8-4b12-ad1d-ebed0266841c","arxiv_id":"2505.12000","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"IQBench evaluates seven vision-language models on 500 visual IQ questions and finds they struggle most with 3D spatial and anagram reasoning.","lead":"IQBench is a new 500-question visual IQ benchmark for vision-language models, reporting that the best tested model, o4-mini, answers 61.5 percent correctly and all models fail badly on 3D spatial and anagram puzzles. It matters because it tries to score not just right answers but the reasoning behind them, a step toward measuring fluid intelligence in AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The validation of the gpt-4o-mini judge does not support the reasoning-divergence claim: Table 4 compares full-set judge scores to 100-sample human scores, and the judge over-scores 3D SPRT by 0.59.","rationale":"The reader's weakest_assumption is correct and I find no reason to move away from a conditional verdict. The accuracy half of the paper (Table 2) is independent of the judge and supports the claim that models are weak on 3D SPRT and anagram accuracy. However, the benchmark's distinctive contribution is the reasoning score and the claim that reasoning quality diverges from answer correctness; that claim rests entirely on gpt-4o-mini. The paper's own Section 3.3 validation is flawed in a concrete way: the LLM row in Table 4 is identical to the full-dataset o4-mini scores from Table 3, but the human experts scored only 100 sampled predictions. Scores over different item sets cannot establish agreement. In addition, the judge and humans differ most on 3D SPRT (0.82 vs 0.23), which is one of the two tasks in the headline, and human experts disagree sharply on anagrams (0.90 vs 0.00/0.10), so a stable gold standard for reasoning is not established. Appendix B also leaves unclear how reasoning text was elicited from o4-mini and gpt-o3, since the prompt asks only for <answer>. These issues do not invalidate the benchmark or the accuracy findings, but they do mean the central reasoning-evaluation claim is not yet supported. A per-item judge-human agreement check on the same 100 predictions, especially for 3D SPRT, is the decisive next step.","tokens_in":12157,"tokens_out":5706,"duration_ms":50610,"concrete_test":"Recompute the reasoning score for o4-mini on the exact 100 predictions used for human evaluation, using gpt-4o-mini with the same prompt, and compute per-item Cohen's kappa between the judge and the majority/adjudicated human label, separately for 3D SPRT and Anagram items. If per-item agreement is below 0.4 on 3D SPRT while human-human agreement on the same items is above 0.6, the judge is not a reliable proxy on the task that carries the headline. As a complementary check, score all 3D SPRT items with a second judge (e.g., claude-3.7-sonnet or gpt-4o); if its reasoning score falls from about 0.82 toward the about 0.23 human mean, the divergence claim collapses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim about inconsistent reasoning is carried by the gpt-4o-mini reasoning scores, but the validation in Section 3.3 is not actually a comparison. Table 4's LLM-as-judge row duplicates the full-dataset o4-mini scores from Table 3 (e.g., 0.92 implies 46/50; 0.82 implies 41/50; 0.14 implies 7/50), while the human experts evaluated only 100 randomly sampled predictions. Those two score sets are computed over different item sets, so the reported 'close alignment' (0.696 vs 0.68) does not validate the judge. Even taking Table 4 at face value, the judge is 0.59 points above the human mean on 3D SPRT (0.82 vs 0.23; humans range 0.20-0.30), precisely the task named in the paper's headline failure claim, and where the largest accuracy-reasoning divergence occurs (gpt-4o: 0.20 accuracy vs 0.56 reasoning). A judge that accepts spatial explanations human experts reject would manufacture the divergence. Human scores are also unstable: on Ana5, Expert 1 gives 0.90 while Experts 2 and 3 give 0.00/0.10. The reasoning-score rankings and divergence conclusion are therefore unsubstantiated without per-item judge-human agreement on the same predictions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"IQBench introduces a manually curated benchmark of 500 vision-centric IQ-test questions across 10 reasoning categories, with 50 items per category. Seven vision-language models are evaluated zero-shot under two metrics: exact-match answer accuracy and a reasoning score produced by an LLM-as-judge pipeline using gpt-4o-mini. The accuracy results place o4-mini, gemini-2.5-flash, and claude-3.7-sonnet at the top, and uniformly show low accuracy on 3D spatial and anagram tasks. The reasoning scores are used to argue that explanation quality often diverges from final-answer correctness, with the highest divergence on 3D SPRT and anagram tasks. A human evaluation of 100 o4-mini predictions is reported as validating the LLM judge.","tokens_in":12447,"tokens_out":5320,"duration_ms":54882,"significance":"If the accuracy tables are taken at face value, IQBench provides a useful vision-centric stress test with substantial manual curation, and the observation that all tested VLMs fail at 3D spatial rotation and anagram tasks is a falsifiable, reproducible result. The attempt to measure reasoning quality separately from answer correctness is also valuable and would be a real contribution if the measurement were valid. However, the paper's headline claim about reasoning divergence rests on a judge-validation procedure that the manuscript's own data contradict. With only 50 items per task and no uncertainty reporting, several of the ranking claims are also stronger than the evidence supports. The public release of code and data is a concrete strength, as is the use of exact-match accuracy for the accuracy scores.","major_comments":[{"comment":"The human evaluation cannot validate the LLM-as-judge because the two score sets are computed over different item sets. The LLM-as-judge row reproduces the full-dataset reasoning scores from Table 3 (e.g., 0.92 on MDRT corresponds to 46/50), while the three human experts scored only 100 randomly sampled predictions, approximately 10 per task. Comparing the resulting averages (0.696 vs. 0.68) is not a per-item agreement test and does not establish that gpt-4o-mini is a reliable proxy for human judgment.","section":"Section 3.3, Table 4"},{"comment":"Even accepting Table 4 at face value, the judge is 0.59 points above the human mean on 3D SPRT (0.82 vs. 0.23; human range 0.20-0.30). This is exactly the task named in the paper's headline failure claim and the location of the largest accuracy-reasoning divergence in Table 3 (gpt-4o: 0.20 accuracy vs. 0.56 reasoning). A judge that accepts spatial explanations that human experts reject would manufacture the reported divergence, so the abstract's and Section 3.2's claim that reasoning quality is inconsistent with final-answer correctness is not supported.","section":"Table 4, 3D SPRT row"},{"comment":"The human reasoning scores are unstable on the same task: on Ana5, Expert 1 scores 0.90 while Experts 2 and 3 score 0.00 and 0.10. With roughly 10 predictions per task in the human sample, task-level human scores carry enormous binomial error. The manuscript needs item-level agreement statistics (e.g., per-item judge-human agreement, Cohen's kappa, or confusion matrices) and confidence intervals before it can claim that the judge and humans are closely aligned.","section":"Table 4, Ana5 row"},{"comment":"No uncertainty quantification is provided for any accuracy or reasoning score. With 50 items per task, the standard error of a 0.50 score is roughly 0.07, and even the overall average difference between o4-mini (0.615) and gemini-2.5-flash (0.578) over 500 items is only about 1.7 standard errors. The paper should report binomial confidence intervals or bootstrap intervals and avoid claiming fine-grained rankings that are within sampling noise.","section":"Table 2 and Table 3"},{"comment":"A showcased sample contains an internal contradiction: the written pattern says \"Therefore option (D) is correct option,\" but the displayed answer line says \"Answer: A.\" Since IQBench's accuracy scores depend entirely on ground-truth labels, this inconsistency should be resolved and the quality-control process for answer annotations should be described.","section":"Figure 1, first example"}],"minor_comments":[{"comment":"The conclusion names \"Gemini 1.5 Flash\" as a tested model, but the experiments use gemini-2.5-flash and gemini-2.0-flash; the conclusion also lists \"Claude 3.5 Sonnet\" while the tables use claude-3.5-sonnet. Please correct the model names and align them with the experimental section.","section":"Section 4, Conclusion"},{"comment":"The gpt-o3 row is only populated for MDRT, DRTF, and Ana5, with no average score; the paper should either report the missing entries or clearly state that gpt-o3 was evaluated on a subset and exclude it from cross-model rankings.","section":"Tables 2 and 3, gpt-o3 row"},{"comment":"The judge prompt asks whether the VLM reasoning \"leads to the correct final answer,\" but Figure 4a presents a case where the reasoning is considered correct although the final answer is wrong. The scoring criterion for this situation should be stated explicitly, otherwise different judges may apply different thresholds.","section":"Appendix C, judge prompt"},{"comment":"The figure caption for Figure 4a says \"Correct reasoning with incorrect prediction,\" while the text says the VLM \"misinterprets the visual options.\" Clarify whether the reasoning score is meant to judge the logical chain only or also the mapping from the chain to the selected option, since this affects the interpretation of every reasoning score in Table 3.","section":"Section 2.2 and Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The core issue is that the reasoning-score framework is the paper's main novelty but is not validated by the reported human evaluation. If the authors can recompute or recalibrate the reasoning scores using per-item judge-human agreement on the same predictions, or substantially constrain the claims to the accuracy results, the paper could become publishable. Otherwise, the headline divergence claim should not appear."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my take on IQBench. The dataset itself is the real contribution: 500 vision-centric IQ questions with hand-written reasoning patterns is a reasonable resource, and the accuracy results across seven current VLMs are plausible as measured. If you only need a compact visual IQ battery, this is worth a look.\n\nThe soft spot is exactly where the stress-test points: the reasoning score. Section 3.3 claims the gpt-4o-mini judge is validated by human experts, but Table 4 compares the judge's full-set scores (which duplicate Table 3's o4-mini row) against human scores on a 100-sample subset. Those are different items, so the 0.696 vs 0.68 'close alignment' isn't a validation. Worse, on 3D SPRT the judge gives o4-mini 0.82 while human experts average 0.23. That's the task the paper highlights as the biggest failure and the biggest accuracy-reasoning divergence (gpt-4o: 0.20 accuracy, 0.56 reasoning). If the judge accepts spatial explanations that human experts reject, the divergence claim is manufactured. The human experts also disagree wildly on Ana5 (0.90 vs 0.00/0.10), so the judge can't be calibrated against that unstable ground truth either.\n\nThere are also smaller but real issues: the paper's own examples contain annotation errors. In Figure 1, the quadrant problem says 'option (D) is correct' but lists 'Answer: A'; the folded-paper problem says 'Option A is correct' but 'Answer: B'. That undermines confidence in the gold labels. The conclusion names Gemini 1.5 Flash, which was not evaluated (the table has Gemini 2.0/2.5 Flash). gpt-o3 is reported only on a few tasks with dashes elsewhere, yet the abstract counts it among 'leading VLMs' tested. No confidence intervals are given, which matters with 50 items per task.\n\nSo my position: the accuracy half of the paper is fine and the dataset is a legitimate contribution. But the reasoning half, which is the advertised novelty, is not supported. The fix is not trivial: the authors need to run per-item judge-human agreement on the same predictions, correct the annotations, and likely re-do the reasoning scores with a better-validated judge. That is refereeable work, not a desk reject. I'd accept it for review with the expectation of major revision. If the annotations and judge validation are cleaned up, this could be a useful benchmark paper.","headline":"A useful 500-question visual IQ dataset, but the reasoning-score claim is undermined by a flawed judge validation and internal annotation errors.","tokens_in":13001,"tokens_out":2829,"would_cite":true,"duration_ms":25577,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new 500-question visual IQ benchmark finds top vision-language models fail 3D spatial and anagram reasoning.","keywords":["IQBench","vision-language models","visual reasoning benchmark","fluid intelligence","LLM-as-a-judge","reasoning evaluation","3D spatial reasoning","anagram reasoning"],"falsifier":"Have independent human experts score reasoning on all 500 IQBench questions and compare with gpt-4o-mini task by task; if agreement on 3D spatial and anagram tasks is near chance, as the 100-sample comparison hints (experts 0.23 vs judge 0.82 on 3D SPRT), then the reasoning rankings and the claimed reasoning-versus-accuracy disconnect would not hold.","tokens_in":11979,"feed_emoji":"🧩","tokens_out":8552,"duration_ms":80099,"temperature":0.7,"pith_summary":"IQBench is a 500-question, visually centered IQ benchmark built to test whether vision-language models (VLMs) can reason from images rather than fall back on learned text. The paper reports that the best model, o4-mini, reaches 0.615 accuracy and a 0.696 reasoning score, while gemini-2.5-flash and claude-3.7-sonnet follow at 0.578 and 0.548 accuracy. Across all seven models, 3D spatial reasoning and anagram tasks remain near the bottom, and reasoning scores often exceed accuracy, meaning a model can justify a wrong answer coherently. The authors conclude that answer-only evaluation overstates VLM intelligence and that reasoning quality must be scored alongside correctness.","feed_headline":"Top vision-language models still fail 3D and anagram IQ tests","feed_subtitle":"A 500-question visual IQ benchmark shows accuracy and reasoning quality diverge in today's vision-language models.","key_machinery":"The load-bearing mechanism is the dual-metric evaluation framework. Accuracy is exact match between the model's final answer and the ground truth; reasoning is scored by gpt-4o-mini, an LLM judge that receives the question text, the ground-truth reasoning pattern, the ground-truth answer, the VLM's explanation, and the VLM's final answer, and returns 1 or 0 depending on whether the explanation is logically sound and consistent with the pattern. The other central object is the dataset itself: 500 manually collected and annotated, vision-centric IQ questions, 50 per topic and 390 open-ended, designed to minimize textual priors and prevent data leakage. The judge converts explanations into a scalable numerical score, and that score is what carries all the paper's claims about reasoning quality and its divergence from accuracy.","core_discovery":"The central claim is that IQBench exposes three things: leading VLMs are strong on number-series and deductive-figure tasks but weak on 3D spatial and anagram reasoning; the reasoning score, produced by an LLM-as-a-judge (gpt-4o-mini) comparing each model's explanation with the annotated reasoning pattern, frequently disagrees with answer accuracy, so models can reach correct answers through flawed reasoning or wrong answers through coherent reasoning; and human experts, while close to the judge on average (0.68 vs 0.696), diverge sharply on specific tasks such as 3D spatial reasoning (0.23 vs 0.82). The paper presents this divergence not as a failure of the benchmark but as the reason reasoning must be measured separately from accuracy.","pith_inferences":["An implication the paper leaves implicit is that the anagram failures may come from the model's tokenization rather than from a lack of reasoning, since scrambled-letter rearrangement is exactly where subword tokenizers behave irregularly.","The paper's own human-evaluation table warns that the reasoning-score rankings for 3D spatial tasks are the least trustworthy: the judge gives o4-mini 0.82 while the three experts average 0.23, so any claim about spatial reasoning quality should be read with that caveat.","A direct testable extension is to re-score all 500 questions with human raters or task-specific rubrics; if per-task agreement with gpt-4o-mini stays high, the reasoning-score story strengthens, and if not, reasoning scores should be reported task by task instead of as one average."],"forward_implications":["Answer-accuracy benchmarks will keep overstating VLM intelligence, because on IQBench coherent reasoning and correct answers do not move together.","Evaluation of future models should report a reasoning score alongside accuracy, crediting models that justify answers correctly and withholding credit from lucky guesses.","3D spatial understanding and anagram manipulation are the two task families that separate current models, making them useful focused probes for VLM architecture improvement.","Because the benchmark is vision-centric and manually curated, a high IQBench score would indicate image-based reasoning rather than memorized textual knowledge.","The reported human-judge agreement supports scaling reasoning evaluation with an LLM judge, provided per-task agreement is checked rather than only the overall average."],"supporting_citations":[{"why":"MMMU is the large multimodal benchmark the paper positions IQBench against; its reported saturation motivates the need for a fluid-intelligence benchmark.","marker":"[17]"},{"why":"MathVista is the visual mathematical-reasoning benchmark that, with MMMU, the paper says does not target fluid intelligence or reasoning interpretability.","marker":"[18]"},{"why":"The abstraction-and-reasoning corpus supplies the definition of fluid intelligence and abstraction that IQBench is built to measure.","marker":"[10]"},{"why":"MM-IQ is a prior multimodal IQ benchmark that focuses on answer accuracy, the limitation IQBench's reasoning score targets.","marker":"[34]"},{"why":"Verify is the benchmark of visual explanation and reasoning fidelity that motivates IQBench's dual-metric design.","marker":"[33]"},{"why":"Reports a multimodal model scoring on MMMU without images, supporting the paper's claim that text-heavy benchmarks overestimate visual ability.","marker":"[31]"},{"why":"Documents shortcut learning in neural networks, supporting the premise that high accuracy can arise without genuine reasoning.","marker":"[23]"}],"fun_headline_variants":["IQBench: VLMs falter on 3D spatial and anagram IQ tests","Vision-language models: smart on numbers, weak on 3D and anagrams","Accuracy and reasoning diverge in vision-language IQ benchmark","Human experts and AI judge clash on VLM reasoning scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reasoning scores all rest on one judge, gpt-4o-mini, agreeing with human experts about what counts as correct reasoning; if that judge is biased on a task, the headline comparison between reasoning quality and answer correctness weakens.","fun_headline_variants_meta":{"raw":{"variants":["IQBench: VLMs falter on 3D spatial and anagram IQ tests","Vision-language models: smart on numbers, weak on 3D and anagrams","Accuracy and reasoning diverge in vision-language IQ benchmark","Human experts and AI judge clash on VLM reasoning scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1803,"prompt_tokens":1064,"completion_tokens":739,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":662}},"tokens_in":680,"tokens_out":739,"duration_ms":6961,"temperature":1.0,"reasoning_tokens":662,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:43:05.685442+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent human experts score reasoning on all 500 IQBench questions and compare with gpt-4o-mini task by task; if agreement on 3D spatial and anagram tasks is near chance, as the 100-sample comparison hints (experts 0.23 vs judge 0.82 on 3D SPRT), then the reasoning rankings and the claimed reasoning-versus-accuracy disconnect would not hold.","supporting_citations":[{"cited_title":"Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts","cited_arxiv_id":null,"evidence_quote":"MathVista is the visual mathematical-reasoning benchmark that, with MMMU, the paper says does not target fluid intelligence or reasoning interpretability."},{"cited_title":"Shortcut learning in deep neural networks","cited_arxiv_id":null,"evidence_quote":"Documents shortcut learning in neural networks, supporting the premise that high accuracy can arise without genuine reasoning."}],"review_version":1}