REVIEW 3 major objections 3 minor 1 cited by
DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Interactive dashboard reasoning is a distinct, largely unsolved capability for vision-language GUI agents, and DashboardQA provides a shared 112-dashboard, 405-question benchmark on which the strongest evaluated agent scores only 38.69%.
desk verdict Useful new benchmark for GUI agents on interactive dashboards, but the abstract does not test the load-bearing claim that questions require interaction; deserves peer review with a demand for static-screenshot and data-only baselines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the DashboardQA benchmark itself: 112 interactive dashboards from a public dashboard-sharing platform paired with 405 question-answer items in five categories. Each item requires an agent to perceive the visual state of a dashboard, emit actions such as filtering, selecting, or switching views, and then answer from the resulting state. The five question types — multiple-choice, factoid, hypothetical, multi-dashboard, and conversational — are the mechanism that partitions the difficulty, and the accuracy scores across agents are the evidence that interactive dashboard reasoning, rather than static chart reading, is the unsolved part.
What would settle it
Run the same agents in a non-interactive mode that shows only the initial dashboard view and check whether accuracy matches the reported interactive numbers: if it does, many questions do not truly require interaction.
Extended reading notes
Core claim
The paper's central claim is that existing visualization question-answering benchmarks are insufficient because they ignore dashboard interactivity, and that DashboardQA fills this gap as the first benchmark explicitly designed for interactive dashboard comprehension by GUI agents. The benchmark comprises 112 interactive dashboards and 405 human-authored question-answer pairs spanning five categories, and evaluation of current agents shows consistent failure: the best evaluated system achieves 38.69% accuracy and another leading agent achieves 22.69%. The failures concentrate in three capabilities: grounding dashboard elements, planning interaction trajectories, and performing reasoning across linked views. The authors conclude that interactive dashboard reasoning is challenging for all evaluated vision-language models and that the benchmark provides a way to measure future improvement.
Load-bearing premise
The benchmark assumes the 405 questions genuinely cannot be answered from a static screenshot or from the dashboard's underlying data, so that success really requires interactive exploration.
Editorial extensions
If this is right
- If the reported accuracies hold, no evaluated GUI agent can be relied on for dashboard-driven decision support: the best agent answers fewer than four of ten questions correctly.
- The benchmark gives the field a common yardstick: future agents can be compared on the same 405 items and the same five question categories.
- The error pattern points to specific abilities to improve: grounding elements on screen, planning multi-step interaction trajectories, and reasoning across views.
- Interactive dashboard reasoning becomes a separable benchmark task rather than a side effect of general visual question answering.
Reading between the lines
- Editorial inference: a screenshot-only control condition would test the benchmark's core premise; if non-interactive agents match the reported scores, the difficulty comes from chart comprehension rather than from interaction.
- Editorial inference: error analyses that split questions by the number of required actions could show whether low accuracy reflects planning failures or grounding failures, guiding which component to fix.
- Editorial inference: the multi-dashboard and conversational categories are the most novel stress tests, since they require composing information across separate views or turns, which static chart benchmarks cannot measure at all.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DashboardQA, a benchmark consisting of 112 interactive dashboards from Tableau Public and 405 question-answer pairs spanning five categories (multiple-choice, factoid, hypothetical, multi-dashboard, and conversational). The authors evaluate several closed- and open-source vision-language GUI agents and report that the best agent, based on Gemini-Pro-2.5, achieves only 38.69% accuracy, while the OpenAI CUA agent reaches 22.69%. From these results, the paper concludes that interactive dashboard reasoning is a difficult task for current multimodal agents and that the benchmark can serve as a common test for progress in this area.
Significance. If the benchmark's validity is established, it would fill a genuine gap: prior visualization QA benchmarks focus on static charts and do not exercise the interactive, multi-view exploration that real dashboards require. A public benchmark with real dashboards and a diverse set of question types would be a useful resource for the multimodal agent community, and the reported low accuracies suggest the benchmark is not saturated. The paper's concrete claims, however, hinge on two assumptions that are not evidenced in the abstract: that a meaningful fraction of the 405 questions cannot be answered from a static screenshot or from the raw dashboard data, and that the human-defined ground truth answers are reliable and consistently gradable. Without these, the reported numbers may reflect chart comprehension plus brittle tool use rather than interactive dashboard reasoning. The authors' release of the benchmark on GitHub is a positive step for reproducibility.
major comments (3)
- [Abstract] The central difficulty claim is not supported by a static-screenshot or data-only baseline. The abstract reports only interactive agent accuracies; without a control condition that answers the same 405 questions from the initial view or from the underlying data, the benchmark may be measuring ordinary chart comprehension plus tool use rather than interaction-driven reasoning. Please add these baselines, or, if they already appear in the full text, point to the specific sections; this is load-bearing for the interpretation of every reported accuracy.
- [Abstract (evaluation reporting)] The reported accuracies of 38.69% and 22.69% are given without error bars, confidence intervals, or statistical significance tests. With only 405 questions, the difference between agents could be within noise, and the absence of a human performance estimate leaves the benchmark's difficulty claim uncalibrated. Please report per-category accuracies, variance estimates, and a human or expert baseline, preferably with inter-annotator agreement on the ground truth answers.
- [Full text (provided copy)] The full text supplied to me is decoding-corrupted and mostly unreadable, so I cannot verify the annotation protocol, question validation, answer matching procedure, or the evaluation environment described in the methodology. Since the benchmark's utility depends on the reliability of its ground truth and on the interaction-necessity checks, please provide a readable version of the manuscript and indicate the sections that describe static-image baselines, annotation agreement, and answer scoring.
minor comments (3)
- [Abstract / Introduction] The phrase 'first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards' should be supported by a clear comparison with prior work on GUI agents and visualization QA; currently no related benchmarks are named in the abstract.
- [Abstract] The five question categories are listed but not defined; a one-line definition of each would help readers understand the scope of the benchmark.
- [Data release] The GitHub link is welcome; please include a license and a statement about the Tableau Public dashboards' terms of use and attribution.
Circularity Check
No circularity: DashboardQA's reported accuracies are external empirical evaluations against a newly constructed benchmark, not derivations from the benchmark's own definitions.
full rationale
I examined the abstract and the available text for any step where a claimed result is defined in terms of its inputs, fitted to a subset of data and then predicted, or justified only by a self-citation. The paper's central claims are that DashboardQA is a new benchmark of 112 interactive dashboards and 405 QA pairs, and that leading GUI agents achieve low accuracy on it (38.69% for Gemini-Pro-2.5 and 22.69% for OpenAI CUA). These are external measurements of agent behavior against a fixed test set, not consequences derived from the benchmark's definitions. The benchmark's difficulty claim rests on the empirical assumption that many questions require interactive exploration, and the abstract reports no static-screenshot or data-only baseline; however, an unverified empirical premise is a correctness or robustness concern, not a circular derivation. No equation in the supplied text is shown to reduce to its own input, and no fitted parameter is renamed as a prediction. The full text is decoding-corrupted, so methodology details cannot be fully verified, but nothing in the available evidence exhibits self-definition, fitted-input-as-prediction, or load-bearing self-citation. Score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The 112 dashboards from Tableau Public are representative of real-world interactive dashboards.
- domain assumption The 405 question-answer pairs require interactive exploration and cannot be answered from static screenshots.
- domain assumption Accuracy computed over model-generated answers is a valid measure of agent capability.
Cite this review
Pith. "Pith review of DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards." pith.science (2026). https://pith.science/paper/HTBS2WMP
@misc{pith2026250817398,
author = {Pith},
title = {Pith review of: DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards},
year = {2026},
howpublished = {\url{https://pith.science/paper/HTBS2WMP}},
note = {Machine review of arXiv:2508.17398}
}
read the original abstract
Dashboards are powerful visualization tools for data-driven decision-making, integrating multiple interactive views that allow users to explore, filter, and navigate data. Unlike static charts, dashboards support rich interactivity, which is essential for uncovering insights in real-world analytical workflows. However, existing question-answering benchmarks for data visualizations largely overlook this interactivity, focusing instead on static charts. This limitation severely constrains their ability to evaluate the capabilities of modern multimodal agents designed for GUI-based reasoning. To address this gap, we introduce DashboardQA, the first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards. The benchmark includes 112 interactive dashboards from Tableau Public and 405 question-answer pairs with interactive dashboards spanning five categories: multiple-choice, factoid, hypothetical, multi-dashboard, and conversational. By assessing a variety of leading closed- and open-source GUI agents, our analysis reveals their key limitations, particularly in grounding dashboard elements, planning interaction trajectories, and performing reasoning. Our findings indicate that interactive dashboard reasoning is a challenging task overall for all the VLMs evaluated. Even the top-performing agents struggle; for instance, the best agent based on Gemini-Pro-2.5 achieves only 38.69% accuracy, while the OpenAI CUA agent reaches just 22.69%, demonstrating the benchmark's significant difficulty. We release DashboardQA at https://github.com/vis-nlp/DashboardQA
Forward citations
Cited by 1 Pith paper
-
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
DSAgentBench is a new 275-task benchmark for end-to-end data science in real computer environments, where current AI agents, especially open-source ones, mostly fail.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
https://openai.com/index/computer-using-agent/
C omputer- U sing A gent --- openai.com. https://openai.com/index/computer-using-agent/. [Accessed 14-07-2025]
work page 2025
-
[4]
Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. 2025. https://arxiv.org/abs/2504.00906 Agent s2: A compositional generalist-specialist framework for computer use agents . Preprint, arXiv:2504.00906
arXiv 2025
-
[5]
Anthropic. 2024. https://www.anthropic.com/news/claude-3-5-sonnet Introducing the next generation of claude
work page 2024
-
[6]
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusiñol, Ernest Valveny, C. V. Jawahar, and Dimosthenis Karatzas. 2019. https://arxiv.org/abs/1905.13648 Scene text visual question answering . Preprint, arXiv:1905.13648
arXiv 2019
-
[7]
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024. https://arxiv.org/abs/2401.10935 Seeclick: Harnessing gui grounding for advanced visual gui agents . Preprint, arXiv:2401.10935
arXiv 2024
-
[8]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, and et al. Inderjit Dhillon. 2025. https://arxiv.org/abs/2507.06261 Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities . Preprint, arXiv:2507.06261
arXiv 2025
Show all 33 references
-
[9]
Gregorio Convertino, Jian Chen, Beth Yost, Y-S Ryu, and Chris North. 2003. Exploring context switching and cognition in dual-view coordinated visualizations. In Proceedings International Conference on Coordinated and Multiple Views in Exploratory Visualization-CMV 2003-, pages...
2003
-
[10]
Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, Andrea Tacchet...
2024 arXiv
-
[11]
Enamul Hoque, Vidya Setlur, Melanie Tory, and Isaac Dykeman. 2018. https://doi.org/10.1109/TVCG.2017.2744684 Applying pragmatics principles for interaction with visual analytics . IEEE Transactions on Visualization and Computer Graphics, 24(1):309--318
2018
-
[12]
Price, and Christopher Kanan
Kushal Kafle, Scott Cohen, Brian L. Price, and Christopher Kanan. 2018. https://arxiv.org/abs/1801.08163 DVQA: understanding data visualizations via question answering . CoRR, abs/1801.08163
2018 arXiv
-
[14]
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos Kadar, Adam Trischler, and Yoshua Bengio. 2018. https://arxiv.org/abs/1710.07300 Figureqa: An annotated figure dataset for visual reasoning . Preprint, arXiv:1710.07300
2018 arXiv
-
[15]
Dae Hyun Kim, Enamul Hoque, and Maneesh Agrawala. 2020. https://doi.org/10.1145/3313831.3376467 Answering questions about charts and generating visual explanations . In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI '20, page 1–13, New York, ...
2020
-
[16]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating S...
2023
-
[17]
Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua. 2025. https://arxiv.org/abs/2504.07981 Screenspot-pro: Gui grounding for professional high-resolution computer use . Preprint, arXiv:2504.07981
2025 arXiv
-
[18]
Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj, Firoz Kabir, Aaryaman Kartha, Md Tahmid Rahman Laskar, Mizanur Rahman, Shadikur Rahman, Mehrad Shahmohammadi, Megh Thakkar, Md Rizwan Parvez, Enamul Hoque, and Shafiq Joty. 2025. https://arxiv.org/abs/2504.05506 Ch...
2025 arXiv
-
[19]
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244
2022 arXiv
-
[20]
Ahmed Masry, Megh Thakkar, Aayush Bajaj, Aaryaman Kartha, Enamul Hoque, and Shafiq Joty. 2024. https://arxiv.org/abs/2407.04172 Chartgemma: Visual instruction-tuning for chart reasoning in the wild . Preprint, arXiv:2407.04172
2024 arXiv
-
[21]
V Jawahar
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V Jawahar. 2021. https://arxiv.org/abs/2104.12756 Infographicvqa . Preprint, arXiv:2104.12756
2021 arXiv
-
[22]
Khapra, and Pratyush Kumar
Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. 2020. Plotqa: Reasoning over scientific plots. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
2020
-
[23]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[24]
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, and 16 others. 2025. https://arxiv.org/abs/2501.1...
2025 arXiv
-
[25]
Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Toyama, Robert Berry, Divya Tyamagundlu, Timothy Lillicrap, and Oriana Riva. 2025. https://arxiv.org/abs/2405....
2025 arXiv
-
[26]
Tableau Software . 2025. Tableau public. https://www.tableau.com/products/public. Accessed: 2025-06-02
2025
-
[27]
Zilong Wang, Yuedong Cui, Li Zhong, Zimin Zhang, Da Yin, Bill Yuchen Lin, and Jingbo Shang. 2024 a . https://arxiv.org/abs/2407.19056 Officebench: Benchmarking language agents across multiple applications for office automation . Preprint, arXiv:2407.19056
2024 arXiv
-
[28]
Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sadhika Malladi, Alexis Chevalier, Sanjeev Arora, and Danqi Chen. 2024 b . https://arxiv.org/abs/2406.18521 Charxiv: Charting gaps in realistic chart understanding in mu...
2024 arXiv
-
[29]
Tianbao Xie, Jiaqi Deng, Xiaochuan Li, Junlin Yang, Haoyuan Wu, Jixuan Chen, Wenjing Hu, Xinyuan Wang, Yuhui Xu, Zekun Wang, Yiheng Xu, Junli Wang, Doyen Sahoo, Tao Yu, and Caiming Xiong. 2025. https://arxiv.org/abs/2505.13227 Scaling computer-use grounding via user interface ...
2025
-
[30]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh J Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, and 1 others. 2024 a . Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Informat...
2024
-
[31]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024 b . https://arxiv.org/abs/2404.07972 Oswo...
2024 arXiv
-
[32]
Yan Yang, Dongxu Li, Yutong Dai, Yuhao Yang, Ziyang Luo, Zirui Zhao, Zhiyuan Hu, Junzhe Huang, Amrita Saha, Zeyuan Chen, Ran Xu, Liyuan Pan, Caiming Xiong, and Junnan Li. 2025. https://arxiv.org/abs/2507.05791 Gta1: Gui test-time scaling agent . Preprint, arXiv:2507.05791
2025 arXiv
-
[33]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. https://arxiv.org/abs/2307.13854 Webarena: A realistic web environment for building autonomous agents . Preprint,...
2024 arXiv
-
[34]
Zifeng Zhu, Mengzhao Jia, Zhihan Zhang, Lang Li, and Meng Jiang. 2025. https://arxiv.org/abs/2410.14179 Multichartqa: Benchmarking vision-language models on multi-chart problems . Preprint, arXiv:2410.14179
2025 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.