{"id":"1c046500-3be5-4bad-a1ea-0724ef94f5bf","arxiv_id":"2508.17398","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"DashboardQA is a new benchmark of 405 question-answer pairs over 112 interactive dashboards; the strongest tested GUI agent reaches only 38.69% accuracy.","lead":"This paper presents DashboardQA, a benchmark with 112 interactive Tableau dashboards and 405 questions that test whether AI agents can explore and reason over live dashboards. In evaluations, the best tested agent reaches only 38.69% accuracy, so current AI assistants still struggle with interactive data analysis.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DashboardQA's central difficulty claim rests on the 405 questions requiring interaction, but the abstract reports no static-screenshot or data-only baseline, leaving the benchmark's core assumption untested.","rationale":"The reader's weakest_assumption—that the 405 questions genuinely require interactive exploration—is exactly the load-bearing point. My stress-test does not introduce a new concern so much as sharpen the same one and make it actionable: the abstract contains no static-screenshot or data-only control, so the reported accuracy gap cannot be attributed to interactive reasoning. I agree with the reader that a definitive verdict cannot be rendered from the available material, and no evidence here moves that conclusion. The paper may well be sound if its methodology includes interaction-necessity checks and human baselines, but that cannot be confirmed from the corrupted full text. The proposed concrete test is a minimal, reproducible way to decide whether the concern lands: if static inputs already solve most questions, the benchmark's central difficulty claim weakens; if they do not, the benchmark is more plausible. I therefore leave the reader's UNVERDICTED verdict unchanged.","tokens_in":16331,"tokens_out":2227,"duration_ms":27458,"concrete_test":"Download the released DashboardQA repository and, for all 405 QA pairs, run a strong VLM under two non-interactive conditions: (A) the initial dashboard screenshot plus the question only, and (B) the underlying Tableau/CSV data plus the question, with no GUI actions. Compare accuracy against the reported interactive Gemini-Pro-2.5 result of 38.69%. Also have two human annotators label each question as answerable from the initial screenshot alone, from raw data alone, or requiring interaction, and report inter-annotator agreement. If condition A or B reaches within 10 percentage points of 38.69%, or if annotators judge a majority of questions solvable without interaction, the interaction-necessity assumption is falsified and the benchmark's headline difficulty interpretation should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DashboardQA measures interactive dashboard reasoning and that top agents' low accuracy (38.69% for Gemini-Pro-2.5, 22.69% for OpenAI CUA) demonstrates the difficulty of that task. This claim is load-bearing on a single unverified assumption: that a meaningful fraction of the 405 QA pairs cannot be answered from a static screenshot or from the underlying dashboard data without interaction. The abstract provides no evidence on this point—no static-image baseline, no data-only baseline, no human performance, and no annotation of how many questions require cross-view coordination, filtering, or navigation. If most questions are answerable from the initial visible view or from raw data, then the reported accuracies measure ordinary chart comprehension plus brittle tool use, and the contribution reduces to a modest benchmark-construction effort rather than a demonstration that interactive dashboard reasoning is hard. The full text supplied for this review is decoding-corrupted, so I cannot verify whether the paper's actual methodology section includes such baselines or interaction-necessity checks. That absence of verifiable evidence is not proof of a flaw, but it makes interaction necessity the least secure point in the argument. The concrete test below would settle whether this concern lands. If static-screenshot or data-only accuracy is close to the interactive agent accuracy, the benchmark's difficulty interpretation would need to be substantially revised; if it is far lower, the concern is resolved in the paper's favor.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DashboardQA, a benchmark consisting of 112 interactive dashboards from Tableau Public and 405 question-answer pairs spanning five categories (multiple-choice, factoid, hypothetical, multi-dashboard, and conversational). The authors evaluate several closed- and open-source vision-language GUI agents and report that the best agent, based on Gemini-Pro-2.5, achieves only 38.69% accuracy, while the OpenAI CUA agent reaches 22.69%. From these results, the paper concludes that interactive dashboard reasoning is a difficult task for current multimodal agents and that the benchmark can serve as a common test for progress in this area.","tokens_in":16530,"tokens_out":2185,"duration_ms":23837,"significance":"If the benchmark's validity is established, it would fill a genuine gap: prior visualization QA benchmarks focus on static charts and do not exercise the interactive, multi-view exploration that real dashboards require. A public benchmark with real dashboards and a diverse set of question types would be a useful resource for the multimodal agent community, and the reported low accuracies suggest the benchmark is not saturated. The paper's concrete claims, however, hinge on two assumptions that are not evidenced in the abstract: that a meaningful fraction of the 405 questions cannot be answered from a static screenshot or from the raw dashboard data, and that the human-defined ground truth answers are reliable and consistently gradable. Without these, the reported numbers may reflect chart comprehension plus brittle tool use rather than interactive dashboard reasoning. The authors' release of the benchmark on GitHub is a positive step for reproducibility.","major_comments":[{"comment":"The central difficulty claim is not supported by a static-screenshot or data-only baseline. The abstract reports only interactive agent accuracies; without a control condition that answers the same 405 questions from the initial view or from the underlying data, the benchmark may be measuring ordinary chart comprehension plus tool use rather than interaction-driven reasoning. Please add these baselines, or, if they already appear in the full text, point to the specific sections; this is load-bearing for the interpretation of every reported accuracy.","section":"Abstract"},{"comment":"The reported accuracies of 38.69% and 22.69% are given without error bars, confidence intervals, or statistical significance tests. With only 405 questions, the difference between agents could be within noise, and the absence of a human performance estimate leaves the benchmark's difficulty claim uncalibrated. Please report per-category accuracies, variance estimates, and a human or expert baseline, preferably with inter-annotator agreement on the ground truth answers.","section":"Abstract (evaluation reporting)"},{"comment":"The full text supplied to me is decoding-corrupted and mostly unreadable, so I cannot verify the annotation protocol, question validation, answer matching procedure, or the evaluation environment described in the methodology. Since the benchmark's utility depends on the reliability of its ground truth and on the interaction-necessity checks, please provide a readable version of the manuscript and indicate the sections that describe static-image baselines, annotation agreement, and answer scoring.","section":"Full text (provided copy)"}],"minor_comments":[{"comment":"The phrase 'first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards' should be supported by a clear comparison with prior work on GUI agents and visualization QA; currently no related benchmarks are named in the abstract.","section":"Abstract / Introduction"},{"comment":"The five question categories are listed but not defined; a one-line definition of each would help readers understand the scope of the benchmark.","section":"Abstract"},{"comment":"The GitHub link is welcome; please include a license and a statement about the Tableau Public dashboards' terms of use and attribution.","section":"Data release"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about interaction necessity is reasonable and testable: static-screenshot and data-only baselines would settle it. The provided full text is corrupted, so I cannot determine whether the manuscript already includes these analyses; if it does, the major comments are largely addressable by pointing to the existing sections. I recommend asking the authors to resubmit a clean copy and to explicitly add the baselines, variance estimates, and annotation-agreement numbers if they are not already present. The benchmark itself appears to be a potentially valuable resource, and the low reported accuracies are interesting, but the validity evidence must be visible in the text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nShort version: this is a plausible, useful new benchmark for GUI agents on interactive Tableau dashboards. The resource itself looks like a clear contribution — 112 real dashboards, 405 QA pairs across five categories, and a low current ceiling (best agent ~39%) that makes it a useful shared challenge. It's not a field-resetting conceptual advance, but it fills a real evaluation gap. If you work on vision-language agents for data analysis, you'll want to know about it.\n\nWhat it does well: it moves beyond static chart QA, which has been the focus of most prior work. The multi-dashboard and conversational categories add welcome complexity. Releasing the data and prompts is a practical plus. The abstract's framing of agent limitations — grounding, planning interaction trajectories, reasoning — is a reasonable summary of where these systems fail. The reported numbers, if accurate, show that current GUI agents are far from reliable on real dashboards.\n\nWhere it's soft: the central claim that these questions require interactive exploration is not supported by anything in the abstract. There is no static-screenshot baseline, no data-only baseline, no human performance, and no analysis of how many questions could be answered from the initial view or from the raw data. If most questions are answerable from a single screenshot, then the low accuracies reflect chart comprehension plus brittle tool use, not interactive reasoning. That is not a fatal flaw — a benchmark can still be useful — but it is a load-bearing assumption that should be tested. The stress-test note gets this right. I couldn't verify the full methodology because the supplied copy was garbled, so I can't tell whether the paper already includes such controls. If it doesn't, that is the main thing a reviewer should ask for.\n\nMinor concerns: only 405 questions is a modest size, and the abstract gives no annotation agreement or variance numbers. \"First benchmark\" claims are always hard to verify, but this one is plausible given the small existing survey work in this niche. The same team defining ground truth and evaluating models is standard practice for benchmarks, not a red flag.\n\nBottom line: this paper deserves a serious peer review, not a desk reject. The benchmark is likely to be a useful resource even if its difficulty claim needs sharpening. The authors should be required to add static-screenshot and data-only baselines, and ideally report human performance and interaction-necessity statistics. If those controls are already in the full text, the paper is close to acceptable as is. I'd bring it to the reading group; it will generate a good discussion about what makes a benchmark for interactive agents.","headline":"Useful new benchmark for GUI agents on interactive dashboards, but the abstract does not test the load-bearing claim that questions require interaction; deserves peer review with a demand for static-screenshot and data-only baselines.","tokens_in":17091,"tokens_out":2572,"would_cite":false,"duration_ms":25629,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Interactive dashboard reasoning is a distinct, largely unsolved capability for vision-language GUI agents, and DashboardQA provides a shared 112-dashboard, 405-question benchmark on which the strongest evaluated agent scores only 38.69%.","keywords":["DashboardQA","interactive dashboards","vision-language models","GUI agents","question answering","benchmark","visualization reasoning","multimodal agents"],"falsifier":"Run the same agents in a non-interactive mode that shows only the initial dashboard view and check whether accuracy matches the reported interactive numbers: if it does, many questions do not truly require interaction.","tokens_in":16130,"feed_emoji":"📊","tokens_out":7203,"duration_ms":64429,"temperature":0.7,"pith_summary":"DashboardQA is introduced as the first benchmark purpose-built to test whether vision-language GUI agents — systems that look at a screen and act on it — can answer questions by actually using interactive dashboards rather than reading static charts. It collects 112 real-world interactive dashboards and 405 question-answer pairs across five categories: multiple-choice, factoid, hypothetical, multi-dashboard, and conversational. The paper evaluates leading closed- and open-source agents and reports that even the strongest agent reaches only 38.69% accuracy, with a second prominent agent at 22.69%. The point of the paper is that interactive dashboard reasoning is a capability current models largely lack, and that DashboardQA can serve as a shared testbed for measuring progress in grounding, planning, and reasoning.","feed_headline":"Best GUI agents score just 38.7% on interactive dashboards","feed_subtitle":"The best agent answers only 38.7 percent of 405 questions on a new interactive dashboard benchmark.","key_machinery":"The load-bearing object is the DashboardQA benchmark itself: 112 interactive dashboards from a public dashboard-sharing platform paired with 405 question-answer items in five categories. Each item requires an agent to perceive the visual state of a dashboard, emit actions such as filtering, selecting, or switching views, and then answer from the resulting state. The five question types — multiple-choice, factoid, hypothetical, multi-dashboard, and conversational — are the mechanism that partitions the difficulty, and the accuracy scores across agents are the evidence that interactive dashboard reasoning, rather than static chart reading, is the unsolved part.","core_discovery":"The paper's central claim is that existing visualization question-answering benchmarks are insufficient because they ignore dashboard interactivity, and that DashboardQA fills this gap as the first benchmark explicitly designed for interactive dashboard comprehension by GUI agents. The benchmark comprises 112 interactive dashboards and 405 human-authored question-answer pairs spanning five categories, and evaluation of current agents shows consistent failure: the best evaluated system achieves 38.69% accuracy and another leading agent achieves 22.69%. The failures concentrate in three capabilities: grounding dashboard elements, planning interaction trajectories, and performing reasoning across linked views. The authors conclude that interactive dashboard reasoning is challenging for all evaluated vision-language models and that the benchmark provides a way to measure future improvement.","pith_inferences":["Editorial inference: a screenshot-only control condition would test the benchmark's core premise; if non-interactive agents match the reported scores, the difficulty comes from chart comprehension rather than from interaction.","Editorial inference: error analyses that split questions by the number of required actions could show whether low accuracy reflects planning failures or grounding failures, guiding which component to fix.","Editorial inference: the multi-dashboard and conversational categories are the most novel stress tests, since they require composing information across separate views or turns, which static chart benchmarks cannot measure at all."],"forward_implications":["If the reported accuracies hold, no evaluated GUI agent can be relied on for dashboard-driven decision support: the best agent answers fewer than four of ten questions correctly.","The benchmark gives the field a common yardstick: future agents can be compared on the same 405 items and the same five question categories.","The error pattern points to specific abilities to improve: grounding elements on screen, planning multi-step interaction trajectories, and reasoning across views.","Interactive dashboard reasoning becomes a separable benchmark task rather than a side effect of general visual question answering."],"supporting_citations":[],"fun_headline_variants":["DashboardQA: first benchmark forces GUI agents to fail badly","Interactive dashboards stump top AI agents: 38.7% top score","DashboardQA: why AI can't handle interactive dashboards yet","Best vision-language agent hits only 38.7% on DashboardQA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark assumes the 405 questions genuinely cannot be answered from a static screenshot or from the dashboard's underlying data, so that success really requires interactive exploration.","fun_headline_variants_meta":{"raw":{"variants":["DashboardQA: first benchmark forces GUI agents to fail badly","Interactive dashboards stump top AI agents: 38.7% top score","DashboardQA: why AI can't handle interactive dashboards yet","Best vision-language agent hits only 38.7% on DashboardQA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000393,"raw_usage":{"total_tokens":2068,"prompt_tokens":950,"completion_tokens":1118,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":1043}},"tokens_in":566,"tokens_out":1118,"duration_ms":9025,"temperature":1.0,"reasoning_tokens":1043,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:03:34.817276+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same agents in a non-interactive mode that shows only the initial dashboard view and check whether accuracy matches the reported interactive numbers: if it does, many questions do not truly require interaction.","supporting_citations":[],"review_version":1}