{"id":"7a8d1528-44d1-4545-9920-4007ea884296","arxiv_id":"2509.25339","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Frontier VLMs perform poorly on VisualOverload, a new dense-scene VQA benchmark, with the best model scoring 69.5% overall and only 19.6% on questions that models collectively found hardest.","lead":"This paper introduces VisualOverload, a VQA benchmark built from 150 dense public-domain paintings with 2,720 human-written questions across six vision tasks. It reports that even the strongest tested model, OpenAI o3, reaches only 19.6% accuracy on the paper's hardest split and 69.5% overall, suggesting large models still struggle with fine-grained perception in crowded scenes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human baseline and an accuracy-defined hard split leave the 19.6% 'critical gap' headline unsupported as evidence of a vision-specific model limitation.","rationale":"Good-faith reading: the paper supplies a useful new resource, a broad 37-model sweep, and an informative error analysis. The broad finding that VLMs underperform in dense scenes is plausible and partly supported by task-level results (best counting 41.7%, best OCR 62.7%), which do not depend on the hard-split definition. However, the 19.6% number is the abstract's central quantitative claim, and it is not independently anchored: the split is calibrated on the same models whose failure it is used to demonstrate, and no human accuracy is provided. The reader flagged the split tautology; I agree that is a weakness, but the more load-bearing issue is that the entire benchmark's answerability is only checked by models. A human baseline would directly settle whether low accuracy indicates a VLM perception gap or a benchmark artifact. If the human baseline is high, the paper's conclusion survives and the CONDITIONAL verdict should stand; if it is low, the headline should be revised and the resource recharacterized. Hence no change to the reader's verdict.","tokens_in":20889,"tokens_out":3975,"duration_ms":35259,"concrete_test":"Run a human baseline on a stratified random sample of at least 100 hard-split items (roughly 50 freeform counting/OCR and 50 multiple-choice including reasoning), with 3+ independent annotators viewing the released 4K images and the same prompts used for models. Compute majority-vote accuracy and per-item inter-annotator agreement. If human majority accuracy is at least 90% with high agreement, the 19.6% model result is a genuine gap; if humans also score near model levels or disagree on many items, the benchmark's 'critical gap' claim fails, and the hard split should be redefined by an intrinsic criterion or dropped from the headline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract and Sec. 3.2) is that the best tested model reaches only 19.6% on the hardest split and 69.5% overall, exposing a 'critical gap'. For this to be evidence about VLMs, two conditions must hold: (i) the hard questions are actually answerable from the image by competent humans, and (ii) the difficulty split reflects properties of the questions, not just of the particular 37-model set. Neither is established. Sec. 2.1 defines hard as [0,20]% model accuracy, so 19.6% is by construction near the ceiling of that bucket; this part of the headline is tautological. More importantly, no human baseline is reported anywhere. The only answerability checks are model-based: quality control runs 37 VLMs, ablates images for three strong models, and uses Gemini 2.5 Pro to detect language bias (Sec. 2.1). Ground truths were manually verified only for questions few models solved. If the low accuracy is instead caused by ambiguous wording, annotation errors, or questions that are not decidable from the image (e.g., the reasoning item 'Does capital punishment appear to be legal in this scene?'), the claimed vision gap disappears. The resolution ablation (Fig. 5) does not fill this gap: it measures VLM accuracy, not whether the required information is human-accessible at the distributed resolution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces VisualOverload, a VQA benchmark consisting of 2,720 manually curated question–answer pairs over 150 high-resolution public-domain paintings, designed to stress fine-grained visual understanding in dense scenes. Questions span six categories (activity recognition, attribute recognition, counting, OCR, reasoning, scene classification), with multiple-choice and freeform answer formats, held-out ground truths, and an evaluation server. The authors evaluate 37 VLMs, report task- and difficulty-wise accuracies, and analyze error modes including counting underestimation, OCR edit-distance errors, and logical inconsistency on paired opposite binary questions. The headline finding is that the best tested model, o3, reaches only 19.6% accuracy on the hardest split and 69.5% overall, which the authors interpret as evidence of a critical gap in current VLM visual understanding.","tokens_in":21280,"tokens_out":7917,"duration_ms":66632,"significance":"If the results hold, the benchmark would provide a valuable stress test for dense-scene perception: it uses fresh, high-resolution, publicly licensed images; manual annotation; private ground truths; a broad 37-model comparison; and informative error analyses (counting tolerance curves, normalized Levenshtein distances, logical-consistency ratios). The resolution ablation and chain-of-thought experiments are useful additions. However, the significance of the headline numbers is reduced by two issues: the difficulty split is defined by the same models it is meant to evaluate, and no human baseline is provided. These issues do not invalidate the benchmark as a resource, but they do mean that the paper's central claim about a 'critical gap' requires additional support.","major_comments":[{"comment":"The hard split is defined in Section 2.1 by the average accuracy of the 37 evaluated models, with thresholds [0,20] for hard. The abstract and Section 3.2 then report that o3 achieves only 19.6% on this split. Because the split is constructed from the same model scores that are subsequently reported, the hard-split number is not an independent measure of question difficulty; it is a consequence of selecting questions on which this particular model set scored below 20%. The claim that these questions are 'hard' in a vision-specific sense requires either an intrinsic definition of difficulty or a human baseline. Please report the distribution of per-question accuracies, clarify how many questions fall into each split, and show whether o3's 19.6% differs meaningfully from the split threshold, or redefine the splits without using the evaluated models.","section":"2.1 (Difficulty splits), 3.2, Abstract"},{"comment":"No human baseline is reported, so the paper's central assertion that low VLM accuracy reveals a 'critical gap' in visual understanding is not fully supported. The quality-control checks — blind evaluation of three models, Gemini 2.5 Pro language-bias detection, and manual verification of ground truths only for questions solved by few models — are all model-based and do not establish that a competent human can answer the hard questions from the image at the distributed resolution. Moreover, several reasoning questions, e.g., 'Does capital punishment appear to be legal in this scene?' and 'I am allergic to seafood, is all of the food on the table safe for me?', require legal, cultural, or domain knowledge beyond the stated 'basic level of everyday world knowledge', which contradicts the 'knowledge-free' design goal. I recommend adding a human-subject evaluation on a stratified random sample of questions (including freeform answers, with the same extraction heuristics) and, at a minimum, softening the claims to 'current VLMs cannot solve these questions' rather than 'these questions should be solvable by basic visual understanding.'","section":"2.1 (Quality control), 3.2"},{"comment":"The overall accuracy (69.5% for o3) is dominated by the scene-classification category, which accounts for 1388 of 2720 questions (51%) and is explicitly described in Section 2.1 as not requiring fine-grained understanding. Presenting 'overall 69.5%' in the abstract alongside '19.6% on the hardest split' is misleading because the former largely reflects performance on the easiest, most global category. Please report a macro-average over the six task categories (or over the five non-scene categories) as the headline overall accuracy, and make the category composition clear in the abstract.","section":"Table 2, Abstract"}],"minor_comments":[{"comment":"The OCR normalization step replaces 'V' with 'U' and 'J' with 'I', which is a lenient equivalence that directly affects the reported OCR accuracy; please justify this choice with examples and report OCR accuracy without this normalization as a sensitivity check.","section":"2.2 (Answer extraction)"},{"comment":"Treating refusals and blank responses as 0 in Figure 2a conflates non-response with underestimation; please also present the count distribution with refusals removed or as a separate category.","section":"4 (Counting)"},{"comment":"The numerical entries for several rows are garbled in the submitted version (e.g., the InternVL3-78B row reads '78.078.0 80.534.7'); please ensure the final table is correctly typeset.","section":"Table 2"},{"comment":"The claim that the human-centered annotation process ensures 'unbiased evaluation' is too strong given that the quality-control pipeline itself relies on model outputs; please soften this wording.","section":"2.3"},{"comment":"The abstract describes the benchmark as 'slightly different', which is unnecessarily informal for a research paper; consider removing the qualifier.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and evaluation are presented carefully, with a public leaderboard and detailed error analyses. The main risk to the paper's impact is that the headline findings are partly self-referential (the difficulty split) and lack a human comparison point. Both are addressable in revision. I would also flag that the overall accuracy number should be reported as a macro-average to avoid the scene-classification majority dominating the claim. The paper fits the journal's scope as a benchmark/evaluation contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nVisualOverload is worth a look. The authors have built a real resource: 150 high-resolution public-domain paintings, 2,720 hand-written QA pairs across six task types, private ground truth, and a 37-model evaluation with error analyses. That's a solid contribution, and the images are genuinely fresh, not recycled. The per-task results (counting ~42% best, OCR ~63% best, reasoning near chance for most models) make a plausible case that current VLMs are far from robust on fine-grained dense perception.\n\nThe soft spots are real and center on the headline. The 'hardest test split' is defined as questions where average model accuracy is below 20% (Sec. 2.1). That makes the 19.6% result tautological — of course the best model sits near the ceiling of that bucket. The more informative numbers are the overall 69.5% and the task-level results, which don't depend on the split. Also, there is no human baseline anywhere in the paper. Without one, we can't rule out that many 'hard' questions are ambiguous, badly worded, or not actually decidable from the image. The quality control is thoughtful — blind model probes, language-bias removal with Gemini — but it uses the same class of models being evaluated, not humans. The paper's own example 'Does capital punishment appear to be legal in this scene?' shows the difficulty of keeping questions knowledge-free and uniquely grounded in the image.\n\nI'd also note the counting analysis and the logical-opposite consistency measure are well done. The logical inconsistency signal (some models below random on paired yes/no) is interesting even if the exact numbers would shift with a cleaner split.\n\nBottom line: this is a serious benchmark with a load-bearing presentation flaw. The fix is straightforward — add a human baseline, report the hard split as a model-calibrated partition rather than an intrinsic property, and lean on overall/task accuracy for the headline. That's a revision, not a rejection. Send it to referees.","headline":"A genuinely useful dense-scene VQA benchmark whose headline 'hard split' number is manufactured by the split's own definition; the paper deserves refereeing, with a required human baseline.","tokens_in":21688,"tokens_out":1983,"would_cite":true,"duration_ms":16285,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current vision-language models fail at simple fine-grained questions in densely packed scenes, with the best of 37 models scoring 19.6% on the hardest split.","keywords":["VisualOverload","visual question answering","dense scenes","fine-grained perception","counting accuracy","OCR in images","logical consistency","vision-language models"],"falsifier":"Check whether humans can reliably answer the hard split: give the same 430 hard questions to a few dozen human annotators without training on the images; if human accuracy is also near 20% or if annotators disagree strongly, the difficulty is in the questions themselves, not in the models' vision. Alternatively, evaluate a model with a much larger visual token budget (e.g., 40 image patches instead of 12) on the same split; if its hard-split accuracy stays below 20%, the fixed-token-budget explanation is weakened.","tokens_in":20721,"feed_emoji":"🖼️","tokens_out":6004,"duration_ms":49240,"temperature":0.7,"pith_summary":"The paper argues that current vision-language models have not solved basic fine-grained visual understanding, and that existing benchmarks overstate their ability because they mostly test global scene-level reasoning. To expose this gap, the authors introduce VisualOverload, a benchmark of 2,720 hand-written question-answer pairs about 150 high-resolution public-domain paintings filled with many figures, actions, and small details. Across 37 tested models, the strongest model scores only 19.6 percent on the hardest difficulty split and 69.5 percent overall. Error analysis shows systematic failures in counting, reading text, and giving logically consistent answers to opposite paired questions. The paper concludes that dense, detail-rich scenes remain a critical bottleneck for current vision encoders.","feed_headline":"Best vision model answers only 19.6% of hardest dense-scene questions","feed_subtitle":"A 2,720-question benchmark of crowded paintings shows models struggle with counting, OCR, and tiny details.","key_machinery":"The central object is the VisualOverload benchmark itself: 150 high-resolution scans of densely populated public-domain paintings, annotated with 2,720 manually curated questions in six categories—activity, attribute, counting, OCR, reasoning, and scene classification. Its load-bearing design choices are the private ground truth (answers withheld, scoring via an evaluation server), the pairing of every binary yes/no question with its logical opposite, and a three-level difficulty split calibrated by the average accuracy of 37 tested models. These machinery pieces let the authors measure not just accuracy but logical consistency and shortcut reliance, and they convert the benchmark's difficulty from an assertion into a measured quantity.","core_discovery":"On the paper's own terms, the central discovery is that state-of-the-art VLMs, though strong at global scene classification, collapse on simple, knowledge-free visual questions once the scene is visually overloaded. The authors define three difficulty levels by the models' own accuracy: questions scored correct by under 20% of models are 'hard', 20-90% are 'medium', and over 90% are 'easy'. Under that split, the best proprietary model achieves just 19.6% on the hard split and 69.5% overall, while the strongest open-weight model reaches 7.2% and 67.6%. The benchmark's paired logical-opposite questions further reveal that models frequently contradict themselves—answering 'yes' to 'Is it day?' and 'yes' to 'Is it night?'—with consistency dropping from 83.3% on scene questions to 60.6% on reasoning questions. The paper therefore claims that the vision encoder, which compresses images into a fixed token budget, is the limiting factor for fine-grained perception.","pith_inferences":["A testable extension would be to fix difficulty by human annotation time or by intrinsic scene properties (number of objects, text length, figure count) rather than model accuracy; if those intrinsic splits reproduce the large model gap, the headline result would become independent of the particular model cohort.","The encoder-bottleneck hypothesis predicts that giving a model more visual tokens (e.g., more image patches) should improve fine-grained performance more than scaling the language model; this is directly measurable with the released benchmark by patching InternVL3 or similar models at higher token budgets.","The pairing of logical opposites could be reused as a self-supervised probe: a model that is accurate but logically inconsistent on a pair reveals that it is likely answering from language priors or spurious correlations rather than from a coherent scene representation.","If the hard-split questions are re-run with a future model that uses adaptive computation or search over image regions, the 19.6% ceiling may rise sharply, which would localize the current failure specifically to fixed-budget vision encoders rather than to the questions themselves."],"forward_implications":["If the results hold, published VQA accuracy numbers substantially overstate real-world fine-grained perception, since models that look strong on global questions fail on detail-level questions in dense scenes.","Counting and OCR, not just reasoning, are the weakest skills: even the best counting model reaches 41.7% and the best OCR model 62.7%, so applications relying on inventory counts or reading signs in clutter are not yet safe.","Logical-consistency scoring gives a cheap extra signal beyond accuracy: a model that answers opposite paired questions inconsistently is likely exploiting shortcuts, and this measure can be added to existing benchmarks without new ground-truth annotation.","Difficulty calibrated by model performance means the hard split is a moving target: as models improve, the same split becomes easier, so the benchmark will need periodic re-splitting or an intrinsic difficulty measure to stay meaningful."],"supporting_citations":[{"why":"Introduced VQA and set the standard that the paper argues overstates ability by focusing on global understanding.","marker":"(Antol et al., 2015)"},{"why":"Showed VQA models rely on language priors, motivating the paper's grounding restrictions.","marker":"(Goyal et al., 2017)"},{"why":"Supplied the yin-yang method of pairing binary questions with logical opposites, which VisualOverload adopts to measure consistency.","marker":"(Zhang et al., 2016)"},{"why":"Demonstrated the look-and-answer bias problem and motivated the blind-performance quality control.","marker":"(Agrawal et al., 2018)"},{"why":"Documented weak counting in vision models, a failure mode the paper quantifies in dense scenes.","marker":"(Paiss et al., 2023)"},{"why":"Represented the needle-in-a-haystack high-resolution benchmark that VisualOverload contrasts with full-scene density.","marker":"(Wu & Xie, 2024)"}],"fun_headline_variants":["VLMs flunk dense-scene test: top model scores 19.6% on hard split","Crowded paintings stump vision AI: best model manages 19.6% accuracy","VisualOverload benchmark: 2,720 questions trip up top vision models","Top VLM scores 19.6% on hard questions in crowded paintings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The hardest split is defined by the models' own low accuracy (under 20%), so the headline '19.6% on hard questions' depends on the unstated premise that these particular models' difficulties reflect intrinsic properties of the scenes rather than quirks of this model cohort.","fun_headline_variants_meta":{"raw":{"variants":["VLMs flunk dense-scene test: top model scores 19.6% on hard split","Crowded paintings stump vision AI: best model manages 19.6% accuracy","VisualOverload benchmark: 2,720 questions trip up top vision models","Top VLM scores 19.6% on hard questions in crowded paintings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001233,"raw_usage":{"total_tokens":5116,"prompt_tokens":1046,"completion_tokens":4070,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":3980}},"tokens_in":662,"tokens_out":4070,"duration_ms":25908,"temperature":1.0,"reasoning_tokens":3980,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:42:29.156057+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether humans can reliably answer the hard split: give the same 430 hard questions to a few dozen human annotators without training on the images; if human accuracy is also near 20% or if annotators disagree strongly, the difficulty is in the questions themselves, not in the models' vision. Alternatively, evaluate a model with a much larger visual token budget (e.g., 40 image patches instead of 12) on the same split; if its hard-split accuracy stays below 20%, the fixed-token-budget explanation is weakened.","supporting_citations":[{"cited_title":"Lawrence Zitnick, and Devi Parikh","cited_arxiv_id":null,"evidence_quote":"Introduced VQA and set the standard that the paper argues overstates ability by focusing on global understanding."},{"cited_title":"Making the v in vqa matter: Elevating the role of image understanding in visual question answering","cited_arxiv_id":null,"evidence_quote":"Showed VQA models rely on language priors, motivating the paper's grounding restrictions."},{"cited_title":"Yin and yang: Balancing and answering binary visual questions","cited_arxiv_id":null,"evidence_quote":"Supplied the yin-yang method of pairing binary questions with logical opposites, which VisualOverload adopts to measure consistency."},{"cited_title":"Don't just assume; look and answer: Overcoming priors for visual question answering","cited_arxiv_id":null,"evidence_quote":"Demonstrated the look-and-answer bias problem and motivated the blind-performance quality control."},{"cited_title":"Teaching clip to count to ten","cited_arxiv_id":null,"evidence_quote":"Documented weak counting in vision models, a failure mode the paper quantifies in dense scenes."}],"review_version":2}