{"id":"0e51cfad-de45-43e2-8979-62385bda8f88","arxiv_id":"2412.05520","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI benchmarks act as relative progress signals for practitioners, but product and policy users often find them insufficient for substantive deployment decisions.","lead":"Interviews with 19 AI practitioners show that benchmarks are used mainly to compare models relative to each other, not to decide whether a model is good enough to deploy. The study suggests that for benchmarks to be useful in product and policy settings, they need to reflect real-world tasks and include human and domain-expert input.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal contradiction: Section 4 says all participants used benchmarks, but Section 4.2 reports some forewent benchmarks entirely; the central universal claim is unsupported as stated.","rationale":"The stress-test pass should identify the most load-bearing concern. I considered the reader's external-validity concern (19-person snowball sample, self-reports) and agree it is substantial. However, the paper contains a more direct internal inconsistency: it explicitly claims that 'all participants used benchmarks for relative comparisons of models' (Section 4) while also reporting that 'others forewent benchmarks entirely' (Section 4.2), with I-5 as a concrete example. This is not a matter of sampling or social desirability; it is a contradiction within the reported data. The central claim depends on the universality of benchmark use as a relative signal; if the sample itself includes non-users, the claim overstates the evidence. This warrants a revision of the abstract and Findings to a qualified form. The paper otherwise has merits: a transparent methodology, a 13-category coding scheme, and explicit limitations. The reader's CONDITIONAL verdict is appropriate because the paper can be accepted with revisions; my specific concern is an additional required revision. I therefore set verdict_should_be to UNCHANGED (still CONDITIONAL) and mark agreement_with_reader as disagree, since the reader's weakest assumption points to external validity rather than this internal inconsistency.","tokens_in":19244,"tokens_out":4809,"duration_ms":40418,"concrete_test":"Re-analyze the interview data to count, from the 19 transcripts, how many participants reported using benchmarks in any capacity versus consciously deciding against them. Use the code table (e.g., '01.use.frequency' has positive counts for only 7 participants; '01.use.impact.none'=2) and identify every informant who explicitly said they do not use benchmarks (e.g., I-5). If even one participant falls in the latter group, revise the Section 4 statement 'all participants used benchmarks' and the abstract's 'across these settings' to reflect the actual proportion, and rerun the central-claim inference. A simple table of benchmark-users vs. non-users per setting would settle the contradiction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—'across these settings, participants use benchmarks as a signal of relative performance difference between models'—is presented as a universal finding. Section 4 opens: 'we found that although all participants used benchmarks for relative comparisons of models, few interpreted these scores as an absolute signal.' Yet Section 4.2 states: 'While some participants developed their own benchmarks to address issues of quality and relevancy, others forewent benchmarks entirely.' Informant I-5 is quoted as deciding not to 'do anything programmatically today' because no existing benchmarks captured customer-relevant data. These statements are incompatible. If some participants deliberately did not use benchmarks, then 'all participants' is false, and the abstract's 'across these settings' overstates the pattern. The contradiction matters because the paper's main contribution is the claim that benchmark use as a relative signal is a cross-cutting phenomenon; if a non-trivial subset of the sample eschews benchmarks altogether, the phenomenon is not universal and the claim must be qualified (e.g., 'among practitioners who use benchmarks'). The reader's concern about sample representativeness is real, but this internal inconsistency is more direct: it can be settled by inspecting the paper's own data, without any external assumptions.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports semi-structured interviews with 19 AI practitioners working in academia, product development, and policy, aiming to understand how AI benchmarks inform model-related decisions. The central finding is that participants use benchmarks primarily as relative performance signals between models, while the sufficiency of these signals for downstream decisions varies by setting: research settings treat benchmark gains as progress, whereas product and policy settings typically require additional evaluation. The paper also proposes design criteria for more effective benchmarks and interprets the results through the UTAUT framework.","tokens_in":19411,"tokens_out":5032,"duration_ms":46650,"significance":"If the relative-use finding holds, it is a useful corrective to the common framing of public benchmark scores as absolute quality certificates, and the cross-setting comparison between research on the one hand and product/policy on the other is a valuable contribution. The study has several strengths: a grounded-theory approach with iterative coding, a detailed code appendix (A.1), pilot testing of the interview protocol, IRB approval, and an explicit limitations section (Section 6) that acknowledges selection bias, gender skew, and the small number of benchmark-avoiders. The direct quotes provide useful evidence for how practitioners reason about benchmarks. The main claims, however, are stated too universally given the sample and are partly contradicted by the manuscript's own data; the normative recommendations in Section 5.3 are also stronger than the interview evidence can support.","major_comments":[{"comment":"The sentence 'we found that although all participants used benchmarks for relative comparisons of models' is contradicted later in the same section: Section 4.2 states that 'while some participants developed their own benchmarks to address issues of quality and relevancy, others forewent benchmarks entirely,' and I-5 is quoted as deciding not to 'do anything programmatically today' because no existing benchmarks captured customer-relevant data. Section 5.1 similarly says 'some dismissed them entirely.' Section 6 also notes that the sample deliberately included people who consciously decided against using benchmarks. The abstract's 'across these settings' and the Finding are therefore stated too universally. The claim should be qualified, for example, to 'among participants who engaged with benchmarks' or 'most participants,' and the contradiction with Section 3.2's sample description should be reconciled.","section":"Section 4, opening paragraph; Section 4.2; Section 5.1"},{"comment":"The informant identifiers R4 and R8 appear in this subsection ('R4 said that [our team] developed our own benchmark...' and 'R8, who developed automatic speech recognition...'), but Table 1 and all other quotes in the manuscript use the identifiers I-1 through I-19. This prevents the reader from tracing these quotes to the described informants and weakens the auditability of the qualitative evidence. Please replace all identifiers with a single consistent scheme or explicitly explain the R-prefix.","section":"Section 4.2.1"},{"comment":"The recommendations are phrased as necessary properties of effective benchmarks ('effective benchmarks should provide meaningful, real-world evaluations... They must capture diverse, task-relevant capabilities...'), but the study's evidence consists of participants' perceptions and suggestions, not a validated relationship between these properties and benchmark effectiveness. The paper does not measure whether benchmarks that satisfy these criteria actually produce better decisions or are more widely adopted. Please frame these points as participant-derived implications or design hypotheses, and soften the normative 'should/must' language accordingly.","section":"Section 5.3 and Conclusion"}],"minor_comments":[{"comment":"The headings 'Is Our Model Be/t_ter?', 'Is Be/t_ter Good Enough?', and 'What Is A Be/t_ter Benchmark?' contain the apparent typographical artifact 'Be/t_ter'; this should read 'Better.'","section":"Sections 4.1.1, 4.1.2, 5.3"},{"comment":"The word 'academmic' appears in the sentence about I-15; it should be 'academic.'","section":"Section 5.1"},{"comment":"The title 'How to Rport And Benchmark Emerging Field-Effect Transistors' contains a typo; 'Rport' should be 'Report.'","section":"Reference list, [14]"},{"comment":"Several listed codes have count 0 (e.g., '01.use.frequency.infrequently', '01.use.frequency.other', '01.use.information.popular', '01.use.trends.nochange'). Please clarify whether these zero-count codes were retained intentionally or should be removed.","section":"Appendix A.1"},{"comment":"The description of participant categories mentions 'researchers in industry' and 'benchmark users in product development and management,' but Table 1 labels areas as Policy, Product, and Research only; aligning these terms would improve clarity.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — here's my read on 2412.05520. The paper does something genuinely new: it interviews 19 practitioners across academia, product, and policy about how they actually use AI benchmarks, and surfaces a distinction that hasn't been crisp in the literature: benchmarks mostly function as relative signals, not absolute ones, and whether 'better' is 'good enough' depends heavily on setting. The quotes are vivid, the grounded-theory coding is visible (521 passages, 13 code categories, counts in the appendix), and the authors are honest about the sample's limits. Section 6 acknowledges gender skew, self-report bias, and small numbers. That is real work and worth engaging.\n\nThe soft spots are real but not fatal. The most direct problem is internal: Section 4 opens by saying 'all participants used benchmarks for relative comparisons of models,' but Section 4.2 reports that some 'forewent benchmarks entirely,' and I-5 is quoted saying his company decided not 'to do anything programmatically today.' You can't have both. The abstract and the central claim need to be qualified to 'among practitioners who use benchmarks' or 'most participants.' This isn't a deep flaw in the qualitative findings, but it is a load-bearing overstatement that a referee should catch.\n\nSecond, the normative recommendations in Section 5.3—good benchmarks should do X, Y, Z—are mostly the participants' own ideals, and the paper presents them as if they follow directly from the study. They are reasonable, but the evidence for them is thin; they should be framed as themes from the interviews, not validated design guidance.\n\nThird, the inconsistent informant identifiers (R4/R8 vs I-4/I-8) is a minor traceability issue; easy to fix. The UTAUT framing is used more as rhetoric than as a tested model, but the authors admit that, so I don't hold it against them.\n\nSample size at 19 and snowball recruitment limit generalizability; the reader flagged this and it's fair, but the authors already say it. I'd rather see them engage with the contradiction above than add more caveats.\n\nWho is this for? People who design or use benchmarks, and researchers studying AI evaluation practice. It won't settle the validity debate, but it gives a usable empirical picture of how practitioners actually interpret scores.\n\nMy call: send it to peer review. It deserves a serious referee, with major revision expected. The core finding is useful and the data is real; the paper just needs to stop claiming universality it doesn't have. I'd cite it after revision.","headline":"Useful qualitative study, but the paper overstates its universality and contradicts itself about whether all participants used benchmarks.","tokens_in":19996,"tokens_out":3240,"would_cite":false,"duration_ms":27141,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On the basis of 19 interviews, the paper argues that AI benchmarks function as relative performance signals, but only academic researchers treat benchmark gains as sufficient evidence of progress.","keywords":["AI benchmarks","practitioner decision-making","qualitative interview study","relative performance signal","model evaluation","technology adoption","benchmark design","product and policy deployment"],"falsifier":"Watch actual deployment decisions: collect a log of, say, 50 real model choices in product and policy settings and record whether the selected model won its public benchmark and whether any additional evaluation was performed. If a majority of decisions are made on benchmark scores alone without supplemental testing, the paper's central distinction between research and product/policy sufficiency would be contradicted.","tokens_in":19011,"feed_emoji":"📊","tokens_out":7240,"duration_ms":70138,"temperature":0.7,"pith_summary":"The paper asks whether public AI benchmark scores carry genuine information for people who must choose or deploy models, or whether they function mainly as marketing. Using semi-structured interviews with 19 practitioners across academia, product development, and policy, it tries to show that benchmarks are consistently used as relative performance signals but that their sufficiency varies by setting: researchers treat a benchmark win as progress, while product and policy practitioners regard it as only a preliminary hint that must be supplemented with task-specific and human evaluation. The authors conclude that benchmarks become more useful when they are anchored to real-world use cases, built with domain experts, transparent about scope, hard enough to resist saturation, and protective against data contamination. The value of the claim, if true, is that it redirects benchmark design away from ever-larger leaderboards and toward instruments that can actually inform deployment, procurement, and safety decisions.","feed_headline":"AI benchmarks signal relative gains, not verdicts, interviews show","feed_subtitle":"Academics treat score gains as progress, while product and policy teams demand extra evidence before acting.","key_machinery":"The analytical machinery is a two-part framing: benchmarks are treated as signals of relative, not absolute, performance, and their adoption is read through an established account of technology adoption [18, 40] whose decisive factor is perceived usefulness. The authors use this framing to explain why the same benchmark can be sufficient in research but not in product or policy: research rewards relative improvement, while deployment decisions require absolute, use-case-specific evidence.","core_discovery":"Across 19 interviews with researchers, product developers and managers, and policy analysts, the paper finds that benchmark scores are used almost universally as a relative signal—a way to see whether a new model or method beats a baseline. What varies is whether that relative signal is enough. In academic research, beating a benchmark is often treated as the definition of progress and is effectively required for publication. In product and policy settings, practitioners describe benchmark outperformance as insufficient for substantive decisions: a low score can block deployment, but a high score does not justify it, so teams build internal task-specific benchmarks, add human evaluation, or skip benchmarks altogether. The paper concludes that a benchmark is informative to the degree its tasks track real-world use, its goals and scope are transparent, it stays challenging without saturating, it reports trade-offs rather than a single number, and its data are protected from contamination.","pith_inferences":["If the relative-signal finding generalizes, a benchmark's practical value lies less in its leaderboard and more in whether its design documents identify target users and use cases; a testable extension is that benchmarks published with explicit use-case statements and human baselines will be adopted more often.","The paper's distinction implies a calibration strategy for the field: publish human-completion time or error baselines alongside scores, because several participants said such anchors made otherwise ambiguous scores interpretable.","Benchmark saturation is not just a measurement problem but an adoption problem: once a benchmark stops separating models, it stops delivering the relative signal that is its only consistently valued function, so designers should plan difficulty distributions from the start.","The interview method cannot fully separate what practitioners say from what they do; a natural follow-up is an observational study that logs which benchmarks actually precede deployment or procurement decisions, which would test the self-reported sufficiency gap."],"forward_implications":["If benchmark scores are primarily relative signals, then a model developer's claim that a higher score means a better product overstates what the benchmark establishes; the score only shows movement against a baseline.","In academic settings, the pressure to beat benchmarks will continue to define research progress and shape which problems are studied, because publication effectively requires outperforming a baseline.","Product and policy teams will keep investing in internal, task-specific benchmarks or manual evaluation, so leadership on public leaderboards will not by itself translate into adoption or deployment.","Benchmark designers who want practical uptake should anchor tasks to concrete use cases, involve domain experts, report multiple metrics reflecting trade-offs, and prevent data contamination, since those are the features practitioners said would make scores informative.","Even a well-designed benchmark will not eliminate the need for human evaluation in high-stakes settings; the paper's participants agreed that manual assessment remains necessary."],"supporting_citations":[{"why":"Supplies the base technology-adoption account whose perceived-usefulness dimension organizes why practitioners do or do not rely on benchmarks.","marker":"[18]"},{"why":"Extends that account into the unified framework the paper applies to interpret why benchmarks fail on performance expectancy.","marker":"[40]"},{"why":"Provides the definition of benchmarks and the critique of their community adoption, grounding the paper's object of study.","marker":"[36]"},{"why":"Offers prior evidence that many public benchmarks fail to guard against contamination or explain interpretation, which the interview findings corroborate.","marker":"[37]"},{"why":"Supplies the argument that simulation or test tasks have fundamental limits in predicting real-world performance, which the paper uses to frame the benchmark-to-reality gap.","marker":"[24]"},{"why":"Illustrates the gap between expert-labeled accuracy and deployed performance with hospital AI tools, supporting the claim that relative scores do not guarantee real-world success.","marker":"[1]"},{"why":"Documents how early standardized benchmarks enabled rapid field progress by giving researchers a shared target, supporting the relative-progress finding.","marker":"[15]"}],"fun_headline_variants":["AI benchmarks: relative gains, not verdicts for deployers","Benchmark wins don't justify deployment, practitioners say","From research to product: benchmarks signal, but don't decide","Interviews: benchmarks are relative signals, not sufficient for decisions","The real value of AI benchmarks: relative, not absolute"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions rest on the assumption that the 19 interviewees, recruited through expert contacts, snowballing, and social-media posts, mostly male and mostly people who do use benchmarks, accurately speak for the broader populations of academic, product, and policy practitioners.","fun_headline_variants_meta":{"raw":{"variants":["AI benchmarks: relative gains, not verdicts for deployers","Benchmark wins don't justify deployment, practitioners say","From research to product: benchmarks signal, but don't decide","Interviews: benchmarks are relative signals, not sufficient for decisions","The real value of AI benchmarks: relative, not absolute"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000284,"raw_usage":{"total_tokens":1699,"prompt_tokens":990,"completion_tokens":709,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":626}},"tokens_in":606,"tokens_out":709,"duration_ms":6949,"temperature":1.0,"reasoning_tokens":626,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:38:08.301517+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Watch actual deployment decisions: collect a log of, say, 50 real model choices in product and policy settings and record whether the selected model won its public benchmark and whether any additional evaluation was performed. If a majority of decisions are made on benchmark scores alone without supplemental testing, the paper's central distinction between research and product/policy sufficiency would be contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the base technology-adoption account whose perceived-usefulness dimension organizes why practitioners do or do not rely on benchmarks."},{"cited_title":"Bender, A lex Hanna, and Amandalynne Paullada","cited_arxiv_id":null,"evidence_quote":"Provides the definition of benchmarks and the critique of their community adoption, grounding the paper's object of study."},{"cited_title":"Kochenderfer","cited_arxiv_id":null,"evidence_quote":"Offers prior evidence that many public benchmarks fail to guard against contamination or explain interpretation, which the interview findings corroborate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the argument that simulation or test tasks have fundamental limits in predicting real-world performance, which the paper uses to frame the benchmark-to-reality gap."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents how early standardized benchmarks enabled rapid field progress by giving researchers a shared target, supporting the relative-progress finding."}],"review_version":1}