{"id":"bdc5b29c-0ba0-4e2b-8c43-4787ae2ec726","arxiv_id":"2608.10677","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 100-task expert-authored benchmark of professional charts resists saturation: the best frontier model scores 45.0%, with failures concentrated in visual perception and domain conventions.","lead":"Chartography is a new benchmark of 100 professional charts, from survival curves to wind roses, with questions written by working experts; the strongest frontier AI model scores only 45.0%. It matters because today's best models appear to solve standard chart benchmarks, yet miss values and conventions in the charts professionals actually rely on.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark's headline difficulty and failure-localization claims rest on a single LLM judge that never sees the chart, with no Chartography-specific judge-accuracy audit; if the judge is biased, all 30 configurations' scores and the 45.0% headline shift.","rationale":"The reader identified exactly the load-bearing weak spot: every pass@1 score depends on a single LLM judge whose accuracy on Chartography is assumed from MT-Bench rather than measured, and that judge is itself a scored configuration. I agree with the CONDITIONAL verdict: the benchmark artifact, provenance, and expert authoring/verification pipeline are real strengths, and the protocol is transparent and reproducible, so rejection would be disproportionate. However, the headline claims—45.0% best score, failure modes concentrated in visual perception—are graded outcomes, and grading validity is untested on this benchmark. The paper's own §4.2 acknowledges the self-judging condition, and §3.3's cited evidence [19] is about general LLM judge agreement on MT-Bench, not about equivalence adjudication for chart-derived numeric answers. A concrete audit would settle whether the concern lands. I see no additional load-bearing concern beyond the reader's: adversarial screening and small N=100 are acknowledged properties, the internal arithmetic of the leaderboard is consistent, and the reasoning-effort analysis (§5.2) is secondary to the judge issue. The most useful next step is the judge audit described above; if it passes, the CONDITIONAL can move to ACCEPT with confidence in the headline numbers.","tokens_in":9430,"tokens_out":1688,"duration_ms":16262,"concrete_test":"Run a Chartography-specific judge audit on a stratified sample of at least 200 trials (e.g., 2 trials × 100 tasks, or 5 trials × 40 tasks) spanning numeric-with-range, numeric-exact, categorical, and multi-part answers, drawn from two or three diverse configurations (e.g., GPT-5.6 Sol max, Gemini 3.5 Flash, and a weak configuration like Mistral Large 3). Have three expert verifiers (or a second independent judge that also never sees the chart) adjudicate each sampled trial with the same rubric. Compute agreement and score deltas between the Gemini 3.5 Flash judge and the expert adjudication. If agreement on pass/fail is below ~95%, or if the rescaled mean pass@1 for the best configuration moves outside 45.0 ± 3 percentage points, the headline difficulty and failure localization must be restated with corrected scores or a judge-robustness analysis.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that the best configuration reaches only 45.0% mean pass@1 and that failures concentrate in visual perception—depends on the validity of the grading protocol in §4.2. All 30 configurations are graded by a single LLM judge, Gemini 3.5 Flash, which never sees the chart. The paper justifies judge validity by citing MT-Bench agreement [19], but MT-Bench measures text-only preference following, not equivalence adjudication of numeric answers against expert-set ranges in a chart-QA setting; no Chartography-specific judge-accuracy audit is reported. Because Gemini 3.5 Flash is itself a scored configuration (§4.2), its 35.9% row is self-graded; any systematic judge bias—e.g., favoring particular answer phrasings, penalizing valid roundings, or mis-handling all-or-nothing multi-part answers—would shift the scores of the other 29 configurations as well. The adversarial difficulty screening (§3.1) also curated items against frontier models; combined with a possibly strict judge, the reported difficulty could be inflated. This is not an internal inconsistency: the internal numbers are consistent, and the protocol is transparent and reproducible. The concern is external validity of the grading step: the headline difficulty and §5.4 failure-mode localization are only as trustworthy as the judge, and judge accuracy on this benchmark is unmeasured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Chartography, a benchmark of 100 professional chart-understanding tasks that pair real-world domain-specific charts (e.g., Kaplan-Meier curves, contour maps, candlesticks, Bode plots) with questions authored by practicing professionals and verified by three additional experts. Each task includes a golden answer with expert-set acceptable ranges, and grading is all-or-nothing across answer parts. The authors evaluate 30 frontier-model configurations using 20 trials per task with no tools, reporting a best mean pass@1 of 45.0% (GPT-5.6 Sol at maximum reasoning) and scores between 9.0% and 39.5% for the rest. They also analyze failure modes, which they attribute primarily to visual perception errors, and release the dataset, images, provenance metadata, and evaluation code.","tokens_in":9705,"tokens_out":3495,"duration_ms":49438,"significance":"If the benchmarking methodology holds up, Chartography addresses a real gap: professional chart reading is under-measured by existing benchmarks that are nearing saturation on standard formats. The expert authoring and verification process, the chart-specific acceptable ranges, and the emphasis on domain conventions are valuable contributions. The transparent protocol, public release, and reproducible harness are notable strengths. However, the central empirical claims—the low aggregate scores and the localization of failures in visual perception—depend critically on the validity of a single LLM judge and on the task-selection procedure, both of which require additional scrutiny before the results can be fully trusted.","major_comments":[{"comment":"The reported scores for all 30 configurations rest on a single LLM judge, Gemini 3.5 Flash, which never sees the chart; the paper cites MT-Bench [19] for judge validity, but that evidence concerns text-only preference following rather than equivalence adjudication of numeric answers against expert-set ranges in a multimodal chart-QA setting. No Chartography-specific judge-accuracy audit is reported, so a systematic judge bias (e.g., against certain phrasings or roundings) would shift every score in Table 3 and the 45.0% headline. This is load-bearing for the paper's central claim; please add a validation study comparing judge verdicts with expert human judgements on a sample of trials, report agreement and error patterns, and ideally use multiple judges or deterministic checks for numeric range membership.","section":"§4.2 Grading"},{"comment":"The adversarial difficulty screening keeps a task only if it elicits a meaningful failure in at least one frontier model. This guarantees difficulty by construction and means the benchmark deliberately excludes tasks that current models can solve, so the statement that Chartography 'resists saturation' is partly a consequence of the selection rule rather than an empirical discovery. The paper should report how many candidate tasks were dropped during screening, describe the screening protocol in detail, and discuss the implications for what population of professional chart-reading tasks the 100-task set represents.","section":"§3.1 Task construction"},{"comment":"Gemini 3.5 Flash is both the judge and a scored configuration, so its leaderboard row (35.9% in Table 3) is self-graded. Because the judge is also used for every other row, this conflates the measurement instrument with the object of measurement. At minimum the Gemini 3.5 Flash row should be graded by a different judge, or the self-judging row should be reported separately; the paper should also disclose any potential conflict in the leaderboard description.","section":"§4.2 Grading, self-judging condition"},{"comment":"The abstract and §6 claim that failures 'concentrate in visual perception,' but §5.4 presents only qualitative categories and anecdotal examples; no quantitative distribution of failure categories across the 55% of failed trials is reported. Since this claim is central to the paper's interpretation, please add a systematic, independently coded classification of a random sample of failed trials with inter-rater agreement, or soften the claim to 'observed failures frequently fall into these categories.'","section":"§5.4 Failure modes"}],"minor_comments":[{"comment":"The caption contains a duplicated phrase: 'More charts with their with their tasks are shown in §5.5.'","section":"Figure 1 caption"},{"comment":"The 'Top model (no tools)' row for prior benchmarks cites a launch-analysis blog rather than a peer-reviewed source; please specify the exact model versions and evaluation dates, or explicitly mark those numbers as informal comparisons.","section":"Table 1"},{"comment":"The paper does not describe how the 95% confidence intervals in Table 3 are computed; please state the formula (e.g., normal approximation, Wilson interval) and whether clustering by task is accounted for.","section":"§4.3 Metric"}],"recommendation":"major_revision","confidential_remarks":"The paper is well written and the benchmark is a useful contribution, but the single-Judge grading without a domain-specific audit is a serious validity threat that must be addressed before the results can be considered reliable. The difficulty-screening circularity also deserves a more prominent caveat. The authors' transparency and willingness to release resources are commendable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chartography is a genuinely useful benchmark paper. The new thing is the combination: real professional chart formats—Kaplan-Meier curves, wind roses, Moody diagrams, 3D surface plots—with questions written by working professionals and graded against expert-set acceptable ranges, plus an explicit adversarial screen to keep difficulty high. That combination doesn't exist in prior work, and the release looks clean: 100 tasks, 20 trials per config, 30 configurations, provenance metadata, code and data public.\n\nThe headline number, 45.0% best pass@1 against 80-90% on other chart benchmarks, is striking and internally consistent. The failure-mode analysis in §5.4 is the most valuable part: the examples of correct-reasoning-wrong-reading, like the growth chart 101.5 vs 101.0, are concrete and credible.\n\nThe soft spots are real but not fatal. The stress-test is right that the whole leaderboard depends on Gemini 3.5 Flash as judge, with no Chartography-specific audit of that judge. It's a fair criticism that MT-Bench isn't the same as adjudicating numeric ranges here. But the design mitigates this more than the stress-test allows: the judge never sees the chart, so it can't substitute its own visual reading, and 51 of 100 answers have explicit ranges, making the judgment largely mechanical. Still, you'd want a human-rated sample to confirm the judge isn't systematically harsh on reasonable roundings or phrasings. The self-graded row (Gemini 3.5 Flash itself) is a minor oddity the paper admits.\n\nThe bigger caveat is the adversarial difficulty screening. Tasks that didn't produce a frontier-model failure were dropped. That's an explicit design choice, in the same spirit as GPQA, but it means the 45% ceiling describes the screened set, not a random draw of professional charts. The paper should say that more prominently. The 100-task scale is also small-ish, and chart types aren't labeled, so failure-mode claims are qualitative.\n\nWho's it for? Anyone building or evaluating multimodal models in medicine, engineering, or finance. It deserves a serious referee. The revision should add a judge-accuracy audit and nuance the difficulty claim.","headline":"A credible, well-built professional-chart benchmark with a 45% ceiling, but the headline difficulty depends on a single un-audited judge and an explicitly curated hard item set.","tokens_in":10219,"tokens_out":2309,"would_cite":true,"duration_ms":22290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces Chartography, a benchmark of 100 expert-authored professional chart tasks on which the best frontier-model configuration passes only 45.0% of trials, and argues the gap is visual perception rather than reasoning.","keywords":["professional chart understanding","chart QA benchmark","visual perception","metric grounding","multimodal LLM evaluation","expert-verified tasks","domain conventions","pass@1 evaluation"],"falsifier":"A reader could settle the central claim by running a judge-accuracy audit: take a random sample of model responses, have domain experts grade them against the golden answers using the paper's protocol, and compare those human verdicts with the automated judge's verdicts. If agreement is materially below the level reported for prior judge-validity studies, or if the judge systematically favors certain answer phrasings, then the absolute scores (including the 45.0% top score) would need revision; a different judge or an ensemble of judges would show how much the leaderboard shifts.","tokens_in":9220,"feed_emoji":"📊","tokens_out":6254,"duration_ms":53616,"temperature":0.7,"pith_summary":"This paper argues that existing chart benchmarks are saturated and miss what professionals actually do with charts, so it builds a harder test: Chartography, 100 tasks using domain-native charts such as Kaplan-Meier curves, contour maps, Sankeys, and three-dimensional surface plots, with questions written by working professionals and independently verified by three additional experts. The result is a benchmark that resists saturation: the best of 30 frontier-model configurations passes only 45.0% of trials, and most configurations land between 9.0% and 39.5%. The paper's failure analysis localizes the gap to visual perception—missed thin features, misread values on sparsely labeled axes, mishandled projected geometry, and overlooked domain conventions—rather than to reasoning depth. A sympathetic reader should take from this that current multimodal models are not yet reliable readers of professional charts, and that earlier benchmark scores overstated that readiness.","feed_headline":"Best AI scores 45% on new professional chart test","feed_subtitle":"Frontier models that ace existing benchmarks fail most expert-authored chart-reading tasks.","key_machinery":"The central object is the benchmark itself: 100 tasks pairing a professional-domain chart with a free-form question that has a prespecified golden answer and, for 51 of 100 tasks, an expert-set acceptable range. The machinery that carries the argument is the expert-calibrated grading protocol: an LLM judge that never sees the chart decides whether the model's final answer falls within the expert range or matches the golden set exactly, with multi-part answers graded all-or-nothing; scores are reported as mean pass@1 over 20 trials per task. This design ensures that grades reward reading the chart at the precision it supports, and that the judge adjudicates answer equivalence rather than re-solving the task.","core_discovery":"Chartography establishes that frontier multimodal models, which score 80–90% on existing chart benchmarks, fail a majority of professionally relevant chart-reading tasks: the strongest evaluated configuration passes only 45.0% of 2,000 graded trials, and the paper attributes the shortfall to a failure of metric grounding—anchoring a numeric or categorical answer to the correct mark, axis position, or plotted geometry—rather than to inadequate reasoning. The benchmark is deliberately constructed so that this gap is measurable: expert-authored questions with walkthroughs, triple independent verification, adversarial difficulty screening, expert-set acceptable ranges that reflect the precision a chart actually supports, and all-or-nothing grading of multi-part answers.","pith_inferences":["The paper's 'semantic recognition versus metric grounding' split likely generalizes beyond charts: any task where a model must anchor an answer to a specific position in a visual, such as maps, technical drawings, or medical images, could show the same pattern of correct identification but wrong measurement.","The no-tools protocol means the 45.0% ceiling describes bare perception; allowing zooming, cropping, or code execution under the same grading rules is a natural test of how much of the gap is a perception limit versus an interface limit.","Because the judge is a single model whose validity is borrowed from prior work rather than measured on this benchmark, the numerical scores should be read as a snapshot; a multi-judge or human-audited grading pass would tell whether the reported rank order is robust."],"forward_implications":["The same models that score above 80% on existing chart benchmarks fall below 50% on Chartography, so high marks on those benchmarks do not establish readiness for clinical, engineering, or financial chart-reading deployment.","Because 11 of 12 model families improved with elevated reasoning effort but gains were uneven, more deliberation helps when the chart is read correctly to begin with and cannot repair a misread value.","The failure analysis implies that improving visual reading alone, while holding reasoning fixed, would raise scores substantially.","With 100 expert-verified tasks, expert-calibrated ranges, and released provenance metadata, the benchmark gives future work a stable and auditable target that is not yet near saturation."],"supporting_citations":[{"why":"The chart-QA benchmark this work contrasts with, cited as nearly solved by current frontier models.","marker":"[8]"},{"why":"The arXiv-figure benchmark whose reasoning gap this work extends, with frontier scores cited around 90%.","marker":"[16]"},{"why":"The chart benchmark separating visual from textual reasoning, with frontier results cited above 80%.","marker":"[14]"},{"why":"The successor benchmark with relaxed-accuracy grading, used as a comparison point for scale and top-model score.","marker":"[9]"},{"why":"The benchmark showing models depend on printed labels rather than visual estimation, motivating Chartography's design.","marker":"[17]"},{"why":"The prior study of LLM-judge agreement that grounds the assumption that a well-specified target makes judge grading reliable.","marker":"[19]"},{"why":"The evaluation harness that implements the 20-trials-per-task protocol used for all reported scores.","marker":"[15]"}],"fun_headline_variants":["AI model fails 55% of professional chart tasks","Frontier AI flunks majority of expert chart tests","New benchmark: AI aces simple charts, fails real ones","Best AI only passes 45% on expert chart benchmark","Professional charts expose AI perception limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every reported score depends on one automated judge whose accuracy on this benchmark is assumed from prior work rather than measured here, and that judge is itself one of the scored models.","fun_headline_variants_meta":{"raw":{"variants":["AI model fails 55% of professional chart tasks","Frontier AI flunks majority of expert chart tests","New benchmark: AI aces simple charts, fails real ones","Best AI only passes 45% on expert chart benchmark","Professional charts expose AI perception limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00027,"raw_usage":{"total_tokens":1579,"prompt_tokens":850,"completion_tokens":729,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":654}},"tokens_in":466,"tokens_out":729,"duration_ms":7588,"temperature":1.0,"reasoning_tokens":654,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:35:41.100268+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could settle the central claim by running a judge-accuracy audit: take a random sample of model responses, have domain experts grade them against the golden answers using the paper's protocol, and compare those human verdicts with the automated judge's verdicts. If agreement is materially below the level reported for prior judge-validity studies, or if the judge systematically favors certain answer phrasings, then the absolute scores (including the 45.0% top score) would need revision; a different judge or an ensemble of judges would show how much the leaderboard shifts.","supporting_citations":[{"cited_title":"Xing, Hao Zhang, Joseph E","cited_arxiv_id":null,"evidence_quote":"The prior study of LLM-judge agreement that grounds the assumption that a well-specified target makes judge grading reliable."},{"cited_title":"Inspect AI: Framework for large language model evaluations","cited_arxiv_id":null,"evidence_quote":"The evaluation harness that implements the 20-trials-per-task protocol used for all reported scores."}],"review_version":1}