Pith. sign in

REVIEW 3 major objections 5 minor 41 references

More than Marketing? On the Information Value of AI Benchmarks for Practitioners

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read On the basis of 19 interviews, the paper argues that AI benchmarks function as relative performance signals, but only academic researchers treat benchmark gains as sufficient evidence of progress.

desk verdict Useful qualitative study, but the paper overstates its universality and contradicts itself about whether all participants used benchmarks. read the letter →

arxiv 2412.05520 v1 pith:BVPYYB6Z submitted 2024-12-07 cs.AI

classification cs.AI
keywords AIbenchmarkspractitionerdecision-makingqualitativeinterviewstudyrelativeperformancesignalmodelevaluationtechnologyadoptionbenchmarkdesignproductandpolicydeployment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether public AI benchmark scores carry genuine information for people who must choose or deploy models, or whether they function mainly as marketing. Using semi-structured interviews with 19 practitioners across academia, product development, and policy, it tries to show that benchmarks are consistently used as relative performance signals but that their sufficiency varies by setting: researchers treat a benchmark win as progress, while product and policy practitioners regard it as only a preliminary hint that must be supplemented with task-specific and human evaluation. The authors conclude that benchmarks become more useful when they are anchored to real-world use cases, built with domain experts, transparent about scope, hard enough to resist saturation, and protective against data contamination. The value of the claim, if true, is that it redirects benchmark design away from ever-larger leaderboards and toward instruments that can actually inform deployment, procurement, and safety decisions.

What carries the argument

The analytical machinery is a two-part framing: benchmarks are treated as signals of relative, not absolute, performance, and their adoption is read through an established account of technology adoption [18, 40] whose decisive factor is perceived usefulness. The authors use this framing to explain why the same benchmark can be sufficient in research but not in product or policy: research rewards relative improvement, while deployment decisions require absolute, use-case-specific evidence.

What would settle it

Watch actual deployment decisions: collect a log of, say, 50 real model choices in product and policy settings and record whether the selected model won its public benchmark and whether any additional evaluation was performed. If a majority of decisions are made on benchmark scores alone without supplemental testing, the paper's central distinction between research and product/policy sufficiency would be contradicted.

Watch

Extended reading notes

Core claim

Across 19 interviews with researchers, product developers and managers, and policy analysts, the paper finds that benchmark scores are used almost universally as a relative signal—a way to see whether a new model or method beats a baseline. What varies is whether that relative signal is enough. In academic research, beating a benchmark is often treated as the definition of progress and is effectively required for publication. In product and policy settings, practitioners describe benchmark outperformance as insufficient for substantive decisions: a low score can block deployment, but a high score does not justify it, so teams build internal task-specific benchmarks, add human evaluation, or skip benchmarks altogether. The paper concludes that a benchmark is informative to the degree its tasks track real-world use, its goals and scope are transparent, it stays challenging without saturating, it reports trade-offs rather than a single number, and its data are protected from contamination.

Load-bearing premise

The conclusions rest on the assumption that the 19 interviewees, recruited through expert contacts, snowballing, and social-media posts, mostly male and mostly people who do use benchmarks, accurately speak for the broader populations of academic, product, and policy practitioners.

Editorial extensions

If this is right

  • If benchmark scores are primarily relative signals, then a model developer's claim that a higher score means a better product overstates what the benchmark establishes; the score only shows movement against a baseline.
  • In academic settings, the pressure to beat benchmarks will continue to define research progress and shape which problems are studied, because publication effectively requires outperforming a baseline.
  • Product and policy teams will keep investing in internal, task-specific benchmarks or manual evaluation, so leadership on public leaderboards will not by itself translate into adoption or deployment.
  • Benchmark designers who want practical uptake should anchor tasks to concrete use cases, involve domain experts, report multiple metrics reflecting trade-offs, and prevent data contamination, since those are the features practitioners said would make scores informative.
  • Even a well-designed benchmark will not eliminate the need for human evaluation in high-stakes settings; the paper's participants agreed that manual assessment remains necessary.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the relative-signal finding generalizes, a benchmark's practical value lies less in its leaderboard and more in whether its design documents identify target users and use cases; a testable extension is that benchmarks published with explicit use-case statements and human baselines will be adopted more often.
  • The paper's distinction implies a calibration strategy for the field: publish human-completion time or error baselines alongside scores, because several participants said such anchors made otherwise ambiguous scores interpretable.
  • Benchmark saturation is not just a measurement problem but an adoption problem: once a benchmark stops separating models, it stops delivering the relative signal that is its only consistently valued function, so designers should plan difficulty distributions from the start.
  • The interview method cannot fully separate what practitioners say from what they do; a natural follow-up is an observational study that logs which benchmarks actually precede deployment or procurement decisions, which would test the self-reported sufficiency gap.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports semi-structured interviews with 19 AI practitioners working in academia, product development, and policy, aiming to understand how AI benchmarks inform model-related decisions. The central finding is that participants use benchmarks primarily as relative performance signals between models, while the sufficiency of these signals for downstream decisions varies by setting: research settings treat benchmark gains as progress, whereas product and policy settings typically require additional evaluation. The paper also proposes design criteria for more effective benchmarks and interprets the results through the UTAUT framework.

Significance. If the relative-use finding holds, it is a useful corrective to the common framing of public benchmark scores as absolute quality certificates, and the cross-setting comparison between research on the one hand and product/policy on the other is a valuable contribution. The study has several strengths: a grounded-theory approach with iterative coding, a detailed code appendix (A.1), pilot testing of the interview protocol, IRB approval, and an explicit limitations section (Section 6) that acknowledges selection bias, gender skew, and the small number of benchmark-avoiders. The direct quotes provide useful evidence for how practitioners reason about benchmarks. The main claims, however, are stated too universally given the sample and are partly contradicted by the manuscript's own data; the normative recommendations in Section 5.3 are also stronger than the interview evidence can support.

major comments (3)
  1. [Section 4, opening paragraph; Section 4.2; Section 5.1] The sentence 'we found that although all participants used benchmarks for relative comparisons of models' is contradicted later in the same section: Section 4.2 states that 'while some participants developed their own benchmarks to address issues of quality and relevancy, others forewent benchmarks entirely,' and I-5 is quoted as deciding not to 'do anything programmatically today' because no existing benchmarks captured customer-relevant data. Section 5.1 similarly says 'some dismissed them entirely.' Section 6 also notes that the sample deliberately included people who consciously decided against using benchmarks. The abstract's 'across these settings' and the Finding are therefore stated too universally. The claim should be qualified, for example, to 'among participants who engaged with benchmarks' or 'most participants,' and the contradiction with Section 3.2's sample description should be reconciled.
  2. [Section 4.2.1] The informant identifiers R4 and R8 appear in this subsection ('R4 said that [our team] developed our own benchmark...' and 'R8, who developed automatic speech recognition...'), but Table 1 and all other quotes in the manuscript use the identifiers I-1 through I-19. This prevents the reader from tracing these quotes to the described informants and weakens the auditability of the qualitative evidence. Please replace all identifiers with a single consistent scheme or explicitly explain the R-prefix.
  3. [Section 5.3 and Conclusion] The recommendations are phrased as necessary properties of effective benchmarks ('effective benchmarks should provide meaningful, real-world evaluations... They must capture diverse, task-relevant capabilities...'), but the study's evidence consists of participants' perceptions and suggestions, not a validated relationship between these properties and benchmark effectiveness. The paper does not measure whether benchmarks that satisfy these criteria actually produce better decisions or are more widely adopted. Please frame these points as participant-derived implications or design hypotheses, and soften the normative 'should/must' language accordingly.
minor comments (5)
  1. [Sections 4.1.1, 4.1.2, 5.3] The headings 'Is Our Model Be/t_ter?', 'Is Be/t_ter Good Enough?', and 'What Is A Be/t_ter Benchmark?' contain the apparent typographical artifact 'Be/t_ter'; this should read 'Better.'
  2. [Section 5.1] The word 'academmic' appears in the sentence about I-15; it should be 'academic.'
  3. [Reference list, [14]] The title 'How to Rport And Benchmark Emerging Field-Effect Transistors' contains a typo; 'Rport' should be 'Report.'
  4. [Appendix A.1] Several listed codes have count 0 (e.g., '01.use.frequency.infrequently', '01.use.frequency.other', '01.use.information.popular', '01.use.trends.nochange'). Please clarify whether these zero-count codes were retained intentionally or should be removed.
  5. [Section 3.2] The description of participant categories mentions 'researchers in industry' and 'benchmark users in product development and management,' but Table 1 labels areas as Policy, Product, and Research only; aligning these terms would improve clarity.

Circularity Check

0 steps flagged · score 1.0 of 10

No circularity: interview findings are empirically grounded; the only self-citation is corroborative, not load-bearing.

full rationale

This is a qualitative interview study, not a derivation. The central claim — that practitioners use benchmarks as relative signals of model performance, with academic users treating improvement as sufficient while product and policy users demand more — is induced from coding 19 interview transcripts (Section 3.4). The supporting evidence consists of direct informant quotations and the code tallies in Appendix A.1, not from equations, fitted parameters, or an imported formal result. The paper's only self-citation is the BetterBench report [37], cited in Section 5.3 as corroboration that public benchmarks lack contamination prevention and interpretation guidance; it is not used to justify the interview findings, to supply a uniqueness theorem, or to exclude alternative interpretations, so it is not load-bearing. The UTAUT framework [40] is an external theory used as an interpretive lens, and the paper explicitly notes that UTAUT does not predict absolute levels of usage, so applying it does not smuggle in the conclusion. The internal tension between the Section 4 statement that 'all participants used benchmarks for relative comparisons' and Section 4.2's report that some participants 'forewent benchmarks entirely' is a consistency and generalizability concern about the sample, not a circularity concern; per the review rules, such correctness risks do not raise the circularity score. No passage exhibits a claim that reduces by construction to its own inputs, so no circular steps are identified.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numeric free parameters or invented entities are present. The analysis depends on four domain assumptions: self-report validity, sample representativeness, coding reliability, and applicability of the UTAUT framework. None of these are established with independent quantitative evidence.

assumptions (4)
  • domain assumption Participant self-reports accurately reflect their actual benchmark-related decisions and practices.
    All evidence comes from interviews; no observational, log, or experimental data are used. This enters in Section 3.2 and is acknowledged as a self-report limitation in Section 6.
  • domain assumption The 19-person purposive and snowball sample is representative enough to support claims about academia, product, and policy.
    The paper generalizes from this sample across three settings while acknowledging gender skew and selection bias in Section 6.
  • domain assumption Grounded-theory coding by the research team captures meaningful patterns in the transcripts.
    Multiple coders and team discussions are described in Section 3.4, but no inter-rater reliability statistics are reported, so coding consistency is assumed.
  • domain assumption UTAUT is an appropriate framework for interpreting benchmark adoption.
    The paper imports UTAUT from Davis [18] and Venkatesh et al. [40] in Section 2.4 and uses performance expectancy as a key interpretive lens; if the framework does not transfer to benchmarks, that interpretation weakens.

how reviews work

0 comments
Cite this review

Pith. "Pith review of More than Marketing? On the Information Value of AI Benchmarks for Practitioners." pith.science (2026). https://pith.science/paper/BVPYYB6Z

@misc{pith2026241205520,
  author       = {Pith},
  title        = {Pith review of: More than Marketing? On the Information Value of AI Benchmarks for Practitioners},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BVPYYB6Z}},
  note         = {Machine review of arXiv:2412.05520}
}
read the original abstract

Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not necessarily reflect the traits of interest to those who will ultimately apply AI models. In this paper, we seek to understand if and how AI benchmarks are used to inform decision-making. Based on the analyses of interviews with 19 individuals who have used, or decided against using, benchmarks in their day-to-day work, we find that across these settings, participants use benchmarks as a signal of relative performance difference between models. However, whether this signal was considered a definitive sign of model superiority, sufficient for downstream decisions, varied. In academia, public benchmarks were generally viewed as suitable measures for capturing research progress. By contrast, in both product and policy, benchmarks -- even those developed internally for specific tasks -- were often found to be inadequate for informing substantive decisions. Of the benchmarks deemed unsatisfactory, respondents reported that their goals were neither well-defined nor reflective of real-world use. Based on the study results, we conclude that effective benchmarks should provide meaningful, real-world evaluations, incorporate domain expertise, and maintain transparency in scope and goals. They must capture diverse, task-relevant capabilities, be challenging enough to avoid quick saturation, and account for trade-offs in model performance rather than relying on a single score. Additionally, proprietary data collection and contamination prevention are critical for producing reliable and actionable results. By adhering to these criteria, benchmarks can move beyond mere marketing tricks into robust evaluative frameworks.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 32 canonical work pages

  1. [1]

    Is AI Ground Truth Really True? The Dangers of Train ing and Evaluating AI Tools Based on Experts’ Know-What

    2021. Is AI Ground Truth Really True? The Dangers of Train ing and Evaluating AI Tools Based on Experts’ Know-What. MIS Quarterly 45 (2021), 1501–1525. Issue 3b. https://doi.org/10.25300/MISQ/2021/16564

  2. [2]

    William C Adams. 2015. Conducting semi-structured inte rviews. Handbook of practical program evaluation (2015), 492–505

  3. [3]

    Karen Anderson and Rodney McAdam. 2005. An empirical ana lysis of lead benchmarking and performance measurement: Guidance for qualitative research. International Journal of Quality & Reliability Management 22, 4 (2005), 354–375

  4. [4]

    Thom pson

    Mohamed Radhouene Aniba, Olivier Poch, and Julie D. Thom pson. 2010. Issues in bioinformatics benchmarking: the cas e study of multiple sequence alignment. Nucleic Acids Research 38, 21 (07 2010), 7353–7363. https://doi.org/10.1093/nar/gkq625 arXiv:https://academic.oup.com/nar/article-pdf/38/21/7353/7186841/gkq625.pdf

  5. [5]

    Anthropic. 2024. Introducing the Next Generation of Cla ude. https://www.anthropic.com/news/claude-3-family. Accessed: 2024-10-10

  6. [6]

    Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchen ko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardou lakis. 2020. A human- centered evaluation of a deep learning system deployed in cl inics for the detection of diabetic retinopathy. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–12

  7. [7]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochast ic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accounta bility, and transparency. 610–623

  8. [8]

    Nick Bostrom, Allan Dafoe, and Carrick Flynn. 2020. Publ ic policy and superintelligent AI: a vector field approach. Ethics of artificial intelligence (2020), 293–326

Show all 41 references
  1. [9]

    Nadia Boukhelifa, Anastasia Bezerianos, and Evelyne Lu tton. 2018. Evaluation of Interactive Machine Learning Systems . Springer International Publishing, Cham, 341–360. https://doi.org/10.1007/978-3-319-90403-0_17

  2. [10]

    Margarita Boyarskaya, Alexandra Olteanu, and Kate Cra wford. 2020. Overcoming failures of imagination in AI infus ed system development and deployment. arXiv preprint arXiv:2011.13416 (2020)

  3. [11]

    Ajay Brahmakshatriya and Saman Amarasinghe. 2023. D2X : An eXtensible conteXtual Debugger for Modern DSLs. In Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Op timization. 162–172. Manuscript submitted to ACM More than Marketing? On the Informa...

  4. [12]

    Peter M Chapman. 2018. Environmental Quality Benchmar ks — The Good, The Bad, and The Ugly. Environmental Science and Pollution Research 25, 4 (2018), 3043–3046

  5. [13]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henri que Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language model s trained on code. arXiv preprint arXiv:2107.03374 (2021)

  6. [14]

    Zhihui Cheng, Chin-Sheng Pang, Peiqi Wang, Son T Le, Yan qing Wu, Davood Shahrjerdi, Iuliana Radu, Max C Lemme, Lian- Mao Peng, Xiangfeng Duan, et al. 2022. How to Rport And Benchmark Emerging Field- Effect Transistors. Nature Electronics 5, 7 (2022), 416–423

  7. [15]

    Kenneth Ward Church. 2018. Emerging trends: A tribute t o Charles Wayne. Natural Language Engineering 24, 1 (2018), 155–160. https://doi.org/10.1017/S1351324917000389

  8. [16]

    Juliet Corbin and Anselm Strauss. 2015. Basics of qualitative research . Vol. 14. sage

  9. [17]

    Lee J Cronbach and Paul E Meehl. 1955. Construct validit y in psychological tests. Psychological bulletin 52, 4 (1955), 281

  10. [18]

    Fred D. Davis. 1989. Perceived Usefulness, Perceived E ase of Use, and User Acceptance of Information Technology. MIS Quarterly 13, 3 (1989), 319–340. http://www.jstor.org/stable/249008

  11. [19]

    Luciano Floridi. 2024. Three tensions in understandin g AI–Comment on Pope Francis’ message Artificial Intelligence and the Wisdom of the Heart: Towards a Fully Human Communication. A vailable at SSRN (2024)

  12. [20]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Je nnifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Ka te Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12 (2021), 86–92

  13. [21]

    Gordon, Kaitlyn Zhou, Kayur Patel, Tatsuno ri Hashimoto, and Michael S

    Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsuno ri Hashimoto, and Michael S. Bernstein. 2021. The Disagreem ent Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality. In CHI. 388:1–388:14. https://doi.org/10.1145/3411764.3445423

  14. [22]

    Declan Grabb, Max Lamparth, and Nina Vasan. 2024. Risks from Language Models for Automated Mental Healthcare: Ethi cs and Structure for Implementation. In First Conference on Language Modeling . https://openreview.net/forum?id=1pgfvZj0Rx

  15. [23]

    Helen Heath and Sarah Cowley. 2004. Developing a ground ed theory approach: a comparison of Glaser and Strauss. International journal of nursing studies 41, 2 (2004), 141–150

  16. [24]

    Abhishek Kadian, Joanne Truong, Aaron Gokaslan, Alexa nder Clegg, Erik Wijmans, Stefan Lee, Manolis Savva, Sonia C hernova, and Dhruv Batra

  17. [25]

    Shahid N Khan. 2014. Qualitative research method: Grou nded theory. International journal of business and management 9, 11 (2014), 224–233

  18. [26]

    Lewis and Albert E

    Byron C. Lewis and Albert E. Crews. 1985. The Evolution o f Benchmarking as a Computer Performance Evaluation Techni que. MIS Quarterly 9, 1 (1985), 7–16. http://www.jstor.org/stable/249270

  19. [27]

    Mark Liberman. 2010. Fred Jelinek. Computational Linguistics 36 (2010), 595–599. Issue 4

  20. [28]

    Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Q Vera L iao, Alexandra Olteanu, and Ziang Xiao. 2024. ECBD: Evidenc e-Centered Benchmark Design for NLP. arXiv preprint arXiv:2406.08723 (2024)

  21. [29]

    John Lofland, David Snow, Leon Anderson, and Lyn H Lofland . 2022. Analyzing social settings: A guide to qualitative observat ion and analysis . Waveland Press

  22. [30]

    Nestor Maslej, Loredana Fattorini, Raymond Perrault, Vanessa Parli, Anka Reuel, Erik Brynjolfsson, John Etcheme ndy, Katrina Ligett, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Rus sell Wald, and Jack Clark. 2024. The AI Index 2024 Annual Repo rt. https://aii...

  23. [31]

    Joseph E McGrath. 1995. Methodology matters: Doing res earch in the behavioral and social sciences. In Readings in human–computer interaction . Elsevier, 152–169

  24. [32]

    Michael Muller. 2014. Curiosity, creativity, and surp rise as analytic tools: Grounded theory method. In Ways of Knowing in HCI . Springer, 25–48

  25. [33]

    Anton J Nederhof. 1985. Methods of coping with social de sirability bias: A review. European journal of social psychology 15, 3 (1985), 263–280

  26. [34]

    OpenAI. 2023. GPT-4 Research and Insights. https://openai.com/index/gpt-4-research/ . Accessed: 2024-10-10

  27. [35]

    Landay, and Beverl y Harrison

    Kayur Patel, James Fogarty, James A. Landay, and Beverl y Harrison. 2008. Examining Difficulties Software Developer s Encounter in The Adoption of Statistical Machine Learning. In Proceedings of the 23rd National Conference on Artificial Int elligence - Volume 3 (Chicago, Illinoi...

  28. [36]

    Bender, A lex Hanna, and Amandalynne Paullada

    Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, A lex Hanna, and Amandalynne Paullada. 2021. AI and the Everyt hing in the Whole Wide World Benchmark. In Thirty-fifth Conference on Neural Information Processing Sy stems Datasets and Benchmarks Track (Round 2) . https://op...

  29. [37]

    Kochenderfer

    Anka Reuel*, Amelia Hardy*, Chandler Smith, Max Lamparth, and Mykel J. Kochenderfer. 2024. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices. https://betterbench.stanford.edu

  30. [38]

    Michael Saxon, Ari Holtzman, Peter West, William Yang W ang, and Naomi Saphra. 2024. Benchmarks as Microscopes: A Ca ll for Model Metrology. arXiv preprint arXiv:2407.16711 (2024)

  31. [39]

    Doga Tascilar. 2023. A Quest through Interconnected Da tasets: Research on Annotation Practices in Highly Cited Au dio Machine Learning Work and Their Utilized Datasets. (2023)

  32. [40]

    acm-jdslogo.png

    Viswanath Venkatesh, Michael G. Morris, Gordon B. Davi s, and Fred D. Davis. 2003. User Acceptance of Information Technology: Toward a Unified View. MIS Quarterly 27, 3 (2003), 425–478. http://www.jstor.org/stable/30036540 Manuscript submitted to ACM 18 Hardy et al. A Appendix ...

  33. [2020]

    Sim2real predictivity: Does evaluation in simulatio n predict real-world performance? IEEE Robotics and Automation Letters 5, 4 (2020), 6670–6677

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.