{"id":"f786fb0d-7395-451d-866f-e19d4c2b3d26","arxiv_id":"2607.08317","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 235-item multimodal stress-test shows frontier closed models outpace open-weight peers by ~10% and leaves shared failures on counting, spatial, and character-level tasks.","lead":"The authors release Blind-Spots-Bench, 235 human-easy tasks that still trip modern language, vision-language, and image models. Closed frontier systems beat open-weight models by about 10% even when both look similar on standard leaderboards, and no model wins every subtask.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The closed–open gap and “persistent blind spots” claim rest on a student-adversarial sample that may mainly rediscover quirks of the October-2025 chatbots students probed, not a general diagnostic.","rationale":"The reader correctly isolates the load-bearing assumption: student-proposed failures against then-available frontier chatbots, after cleaning, are treated as a general blind-spot diagnostic. That is the condition under which the ~10% closed–open gap (Fig. 1, Table 1) and the “hard for all” subtask story (Tables 8–9; perceptual counting / attribute recognition) support the abstract’s claim rather than a construction artifact. Other risks (AI grader, small/imbalanced n, no human baseline, tool-use mixed effects) are real but secondary: grader agreement is high and not pro-Google in a way that would invent the closed–open gap; cost and scaling analyses are secondary. No formal proof is claimed; code and data are public, which strengthens the empirical report but does not fix selection. I therefore leave the verdict CONDITIONAL and agree with the reader’s weakest_assumption. The concrete test is a targeted re-sampling / re-probe experiment that would falsify or support transfer of the gap and taxonomy ranking without requiring a full new benchmark.","tokens_in":25686,"tokens_out":700,"duration_ms":7307,"concrete_test":"Hold out a stratified 40–50 item subset and re-collect an independent parallel set of equal size using the same “easy for humans, hard for AI” brief but with a fixed, disclosed probe panel that includes both the original Oct-2025 chatbots and current open-weight leaders (e.g., GLM-5.2, Qwen3.5-397B, DeepSeek-V4), then re-run the full leaderboard and taxonomy means. If the closed–open text-only gap shrinks below ~5 points or the relative hardness of perceptual counting / attribute recognition reorders, the central diagnostic claim does not transfer.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s strongest claim is that closed-source frontier models beat open-weight models by ~10% on Blind-Spots-Bench even at comparable AAII, and that taxonomy subtasks (e.g., perceptual counting ≤~60%) expose shared, persistent blind spots. That claim requires the 235 items—collected as “five questions frontier chatbots failed” from students in Oct 2025, then cleaned and difficulty-thresholded (§3.1; Limitations)—to be a representative stress set rather than a set biased toward the specific failure modes of the chatbots students actually used. If the pool is dominated by those models’ quirks (tokenization length tricks, particular counting/spatial prompts, image-gen attribute binding), then: (i) the closed–open gap can be an artifact of later closed models having been patched on similar public failure modes while open models were not, and (ii) the ranking of “hard for all” subtasks may not generalize beyond this construction process. The paper itself flags this sampling bias and the missing human baseline, but the headline comparative results still treat the cleaned set as diagnostic of modern models generally. Grader validation (96.6%/90.9%) and AAII correlation support measurement quality, not sampling representativeness.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces Blind-Spots-Bench, a 235-item multimodal benchmark of open-ended tasks that are intended to be easy for humans but hard for current AI systems. Items were collected from graduate students asked (around October 2025) to propose failures of frontier chatbots, then cleaned, annotated with structured reference solutions, question formats, and a three-category / 12-subtask taxonomy (object-centric, abstract reasoning, language-and-knowledge). The authors build an Inspect-AI grading pipeline (Gemini-3-flash grader with code execution), validate grader–human agreement (96.6% text, 90.9% image) and same-provider bias, and evaluate 32 LLMs/VLMs plus 6 image-generation models with mean@k / pass@k, cost, and token reporting. Main empirical claims are: (i) closed-source frontier models outperform open-weight models by roughly 10% on text-only items even at comparable Artificial Analysis Intelligence Index scores; (ii) open models can be more cost-effective; (iii) tool use is not uniformly helpful; (iv) no single model dominates all subtasks, and fine-grained visual perception (e.g., perceptual counting, attribute/pattern recognition) remains hard for all systems.","tokens_in":26066,"tokens_out":1066,"duration_ms":16947,"significance":"If the sampling and evaluation hold up, the work is a useful diagnostic complement to saturated aggregate benchmarks: it ships a public dataset with structured solutions, a reproducible grading harness, multi-format coverage (text-only, multi-to-text, image-gen), cost–accuracy trade-offs, tool-use ablations, and taxonomy-level breakdowns that show complementary model strengths rather than a single ranking. The grader validation and same-provider bias check are stronger than typical LLM-as-judge practice. The closed–open gap at matched AAII and the shared weakness on counting/attribute tasks are concrete, actionable findings for robustness research. Significance is tempered by modest size, subtask imbalance, and student-adversarial construction, so the primary value is as a stress test and analysis framework rather than a definitive measure of general capability.","major_comments":[{"comment":"§3.1 and Limitations: the central diagnostic claim—that the ~10% closed–open gap (Abstract; Fig. 1; Table 1) and the ranking of “hard for all” subtasks reflect persistent blind spots in modern models—depends on treating student-proposed failures of October-2025 frontier chatbots, after cleaning and difficulty thresholding, as a representative stress set. That process can over-weight quirks of the specific systems students probed (e.g., character-length tricks, particular counting/spatial prompts). The paper flags this but still presents headline comparative results as general. Please either (a) report which models students primarily failed against and analyze item difficulty stratified by that provenance, or (b) reframe claims as results on this adversarial construction and add a small held-out or independently authored item set to test whether the closed–open gap and subtask hardness or","section":null},{"comment":"§4.3, Table 8 / Table 4: fine-grained conclusions (e.g., “even the strongest models obtain only 41.67% and 57.14%” on attribute/pattern recognition and perceptual counting; “no single model remains top-1 across all tasks”) rest on very small and uneven subtask counts (attribute recognition n=6; constraint reasoning n=9; several image-gen abstract cells n=1–2). With mean@4 and stderr, these percentages are unstable and can flip rankings. Either pool rare subtasks into coarser categories for primary claims, report bootstrap CIs / exact counts in the main text, or clearly mark which subtask comparisons are exploratory only.","section":null},{"comment":"§3.1 Review and Quality Control and Limitations: the premise that tasks are “almost trivial” / “easy for humans” is not quantified. Difficulty thresholding removed items “easily solved by models or overly difficult for humans,” but no human accuracy, time, or agreement study is reported. Without a human baseline (even on a stratified subset), the human–model gap that motivates the benchmark remains asserted. A modest human study on a representative sample would substantially strengthen the central framing.","section":null},{"comment":"§4.1 / Table 1 vs Fig. 1a: the claim that closed models outperform open models “even when they attain comparable performance on existing benchmarks” is important and only partially supported. Fig. 1a shows a positive AAII correlation with a visual open/closed separation, but there is no matched-pair or regression analysis controlling for AAII (or cost). Please quantify the residual closed–open gap at fixed AAII (e.g., regression with family fixed effects or nearest-neighbor matching) so the “even at comparable AAII” claim is statistical rather than visual.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a competent evaluation paper, not a theory paper. What you get is a public 235-item set with structured reference solutions, a dataset-specific taxonomy (object-centric / abstract / language-knowledge, 12 subtasks, multi-label failure notes), an Inspect-based grading harness with human agreement (96.6% text, 90.9% image) and a same-provider bias check, plus a broad closed/open multimodal leaderboard with cost, tokens, tool-use ablations, and subtask breakdowns.\n\nWhat is actually new is the package: student-mined “easy for humans, hard for frontier chatbots” items from Oct 2025, cleaned and annotated for automatic verification, then run across LLMs, VLMs, and image-gen models. The strongest empirical claim—that closed frontier models sit ~10% above open-weight models on this set even when AAII is comparable, and that no model owns every subtask while counting/attribute recognition stay hard—is supported by mean@k/pass@k tables with stderr, not by hand-waving. Public data and code matter here; that is real credit.\n\nThe soft spot is sampling, and the paper already names it. Items came from students asked for five failures of then-available chatbots, then difficulty-thresholded. That can rediscover quirks of those systems (length tricks, particular counting/spatial prompts, attribute binding) rather than a general diagnostic of “persistent blind spots.” If later closed models were patched on similar public failure modes, the closed–open gap and the “hard for all” ranking may not travel. Subtask n is also small and uneven (e.g., attribute recognition n=6), and there is no human baseline. Grader quality and AAII correlation address measurement, not representativeness. Those are real limits; they do not erase the measurements on this set.\n\nWho it is for: people who care about evaluation practice, model selection under cost, and multimodal robustness gaps that aggregate indices miss. Citation pattern looks normal for the area; methods are transparent enough for a referee to argue with.\n\nI would send it to peer review. Engage if you work on eval or multimodal failure modes; treat the headline gap as a finding on this construction process, not as settled general truth.","headline":"Useful public multimodal stress test with real leaderboard work; the ~10% closed–open gap is measured carefully but sits on a student-adversarial sample that may overfit 2025 chatbot quirks.","tokens_in":26687,"tokens_out":568,"would_cite":true,"duration_ms":6387,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Closed-source frontier models beat open-weight peers by about 10% on human-easy tasks that still stump modern AI, even when scores match on standard benchmarks.","keywords":["multimodal benchmarks","blind spots","vision-language models","image generation","model evaluation","task taxonomy","open-weight vs closed models","perceptual counting"],"falsifier":"If an open-weight model with Artificial Analysis Intelligence Index scores comparable to a top closed model matches or exceeds that closed model’s mean accuracy on Blind-Spots-Bench text-only and multi-to-text splits—especially perceptual counting and character-level manipulation—the claimed closed–open robustness gap would not hold.","tokens_in":26583,"feed_emoji":"👁️","tokens_out":973,"duration_ms":21196,"temperature":0.7,"pith_summary":"Modern AI systems look strong on established benchmarks yet still fail at problems humans find almost trivial—exact string lengths, counting objects in a photo, or drawing a dog with five legs. This paper builds Blind-Spots-Bench: 235 such open-ended questions collected from students, cleaned, given structured reference solutions, and labeled with a three-category taxonomy covering object-centric skills, abstract reasoning, and language-and-knowledge. An automated grading pipeline, checked against humans, scores dozens of language, vision-language, and image-generation models. The main result is that closed-source frontier systems substantially outperform open-weight models—roughly a 10% accuracy gap on text-only items—even among models that look comparable on broad intelligence indices. No single model leads every subcategory, and some skills, especially fine-grained visual counting and pattern recognition, stay hard for all systems. The benchmark is meant as a diagnostic stress test that surfaces concrete weaknesses aggregate leaderboards under-measure.","feed_headline":"Closed models beat open ones by ~10% on trivial AI fails","feed_subtitle":"A 235-task stress test finds gaps standard benchmarks miss, and no model wins every skill.","key_machinery":"Blind-Spots-Bench: a 235-sample multimodal set of human-easy, model-hard open-ended tasks, equipped with structured reference solutions, a taxonomy of three high-level categories and twelve subcategories, and an AI grading pipeline with human-validated agreement. That package turns student-elicited failures into comparable scores that separate models which look equal on aggregate public benchmarks.","core_discovery":"On Blind-Spots-Bench, closed-source frontier models substantially outperform open-weight models—about a 10% accuracy gap on text-only problems—even when those models attain comparable scores on established benchmarks such as the Artificial Analysis Intelligence Index. Fine-grained taxonomy analysis shows no single model dominates all task types, and some subtasks, notably perceptual counting and attribute/pattern recognition, remain difficult for every evaluated system.","pith_inferences":["Training and evaluation optimized for widely used public suites may systematically under-weight character-level control, counting, and spatial binding.","The closed–open gap on this style of constraint-heavy prompt is a natural target for open post-training experiments that would test whether the gap is architectural or data-driven.","Exact-count and inverted-spatial failures in image generation may share roots with VLM counting errors, favoring joint multimodal diagnostics over separate image and language suites.","A measured human baseline on the same 235 items would turn the “easy for humans” claim into a quantified model–human gap useful for product and safety risk assessment."],"forward_implications":["Aggregate public benchmarks can overstate robustness on underrepresented skills that humans find trivial.","Open-weight models can deliver better accuracy per unit inference cost on these tasks even when absolute accuracy lags.","Tool use such as code execution is not uniformly helpful and can lower accuracy when models mishandle tool inputs.","Scaling size within a model family does not consistently improve every subtask; larger variants sometimes regress on specific categories.","Taxonomy-structured stress tests can reveal complementary model strengths that a single overall score hides."],"fun_headline_variants":["Closed models lead open ones by ~10% on AI blind spots","Blind-Spots-Bench: closed models beat open by 10% on trivial fails","Frontier closed models top open weights ~10% on simple tasks","New bench shows 10% closed-open gap standard tests miss","No model wins all: closed lead ~10% on AI's easy blind spots"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The load-bearing premise is that student-invented questions meant to stump late-2025 frontier chatbots, after cleaning and difficulty filtering, form a fair map of persistent blind spots rather than a catalog of those particular models’ quirks.","fun_headline_variants_meta":{"raw":{"variants":["Closed models lead open ones by ~10% on AI blind spots","Blind-Spots-Bench: closed models beat open by 10% on trivial fails","Frontier closed models top open weights ~10% on simple tasks","New bench shows 10% closed-open gap standard tests miss","No model wins all: closed lead ~10% on AI's easy blind spots"]},"model":"grok-4.5","effort":"low","cost_usd":0.0047,"raw_usage":{"total_tokens":1312,"prompt_tokens":793,"num_sources_used":0,"completion_tokens":100,"cost_in_usd_ticks":47000000,"prompt_tokens_details":{"text_tokens":793,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":419,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":793,"tokens_out":100,"duration_ms":4609,"temperature":1.0,"reasoning_tokens":419,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T09:29:37.143201+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If an open-weight model with Artificial Analysis Intelligence Index scores comparable to a top closed model matches or exceeds that closed model’s mean accuracy on Blind-Spots-Bench text-only and multi-to-text splits—especially perceptual counting and character-level manipulation—the claimed closed–open robustness gap would not hold.","supporting_citations":[],"review_version":1}