{"id":"1ca90e76-3500-4ad8-8b6c-a63d7a528e5f","arxiv_id":"2608.11323","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Across three enterprise agent benchmarks, agent main effects are under 3% of total score variance, so leaderboard order reflects task specialization rather than a general capability advantage.","lead":"Across three public agent benchmarks, a psychometric variance decomposition finds that the agent's general capability accounts for under 3% of score variance, while the agent-by-task interaction accounts for 7-23%. If true, leaderboard ranks measure which tasks an agent handles well, not which agent is more capable, which changes how enterprises should read benchmark comparisons.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own Bayesian estimator contradicts the headline: §5.1 reports σ_a posterior median 2.19 on the logit scale, implying a capability-gap ratio near 0.48 rather than <0.001, so the 'three estimators agree' claim is unsupported.","rationale":"The reader identified the Bayesian credible interval for σ_a as a sign that the near-zero main effect could be a small-sample artifact and recommended a power analysis. That is a fair concern, but the more specific and load-bearing problem is internal: the posterior median itself, not just the interval width, contradicts Table 1. A logit-scale σ_a of 2.19 is a large random-intercept SD, and the implied capability-gap ratio with σ_a:t = 2.26 is about 0.48, whereas Table 1 reports <0.001 for TheAgentCompany. The paper's central claim depends on the three estimators agreeing, and the one estimator that is appropriate for binary outcomes disagrees sharply. This is not a case where the concern is outside consensus; it is a case where the manuscript's own reported numbers undermine the headline. I therefore recommend REJECT rather than CONDITIONAL: as written, the central empirical claim is not supported by the evidence the paper itself provides. If the authors can supply a valid conversion or revise the analysis so that the Bayesian and REML-based capability-gap ratios agree, the verdict could be revisited.","tokens_in":10520,"tokens_out":10036,"duration_ms":98691,"concrete_test":"Run the published Bayesian fitting script and compute the posterior distribution of R = σ²_a / (σ²_a + σ²_a:t) directly from the logit-scale variance components. If the posterior median and 95% credible interval of R do not contain the Table 1 gap ratios (e.g., <0.001 for TheAgentCompany and 0.187 for τ2 action_checks), then the three estimators do not agree and the central claim is not established. In the same script, print the exact conversion used to compare logit-scale variance components with the linear-model percentage-of-total-variance figures; without such a conversion, the abstract's agreement claim remains unsupported.","verdict_should_be":"REJECT","load_bearing_attack":"The central empirical anchor is that the agent main effect is below 3% of total variance and the capability-gap ratio is below 0.19, so leaderboards rank specialization rather than capability. Table 1 reports TheAgentCompany with σ²_a < 0.001% and gap ratio < 0.001. But §5.1 reports the Bayesian binomial GLMM posterior for σ_a on the logit scale with median 2.19 and 95% CrI [0.14, 3.87], alongside σ_a:t median 2.26. Using the paper's own definitions, the implied capability-gap ratio is σ²_a / (σ²_a + σ²_a:t) = 2.19² / (2.19² + 2.26²) ≈ 0.48, not below 0.001. This is not merely a small-sample power concern about detecting a tiny effect; the binary-data estimator used by the paper estimates a large agent main effect. The abstract's claim that all three estimators agree to three decimal places is therefore contradicted by the numbers in the same section. If the binomial GLMM is the appropriate estimator for binary outcomes, the headline finding fails unless a scale conversion—not provided—maps logit-scale variances to percentage-of-total-variance in a way that makes the agent main effect essentially vanish. The text instead describes the result as 'consistent with REML', which the reported medians do not support.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies four-facet Generalizability Theory to three agent-trace benchmarks (TheAgentCompany, τ²-bench, AppWorld) and an auxiliary failure-mode dataset, modeling binary success indicators as a function of agent, task, step, and error-category random effects. It reports that the agent main-effect variance is below 3% of total variance in every cell, while the agent-by-task interaction accounts for 7–23%, leading to the claim that leaderboards rank task specialization rather than uniform capability. It adds cost-aware reliability (RPD), difficulty-conditional reliability, a 50-split hold-out test, cross-dataset transfer, and MAST failure-mode analyses, and packages the diagnostics into a Deployment Decision Reliability (DDR) reporting discipline for enterprise procurement.","tokens_in":10861,"tokens_out":6931,"duration_ms":65149,"significance":"If the empirical claims survive correction, the paper is a useful methodological contribution to agent evaluation. It offers a concrete measurement-theoretic vocabulary (capability-gap ratio, difficulty-conditional Eρ², held-out reversal), a clear procurement implication, and a falsifiable negative finding about the usefulness of aggregate leaderboard scores. The open-source code release, data loaders, fit artifacts, and traceability table are substantial strengths, as is the candid limitation section. The paper also makes a novel and testable observation that training-cell reliability projections can negatively correlate with held-out reliability. However, the headline empirical claim currently rests on internally inconsistent tables and an unexplained mismatch between the frequentist and Bayesian estimators, so I cannot certify the central result in its present form.","major_comments":[{"comment":"Table 1, τ²-bench db_check row: σ²_a=0.00% with gap ratio 0.000 but Eρ²=0.899 is impossible under the definition in §3, where Eρ²=σ²_a/(σ²_a+σ²_δ). With a zero numerator the coefficient must be zero, as the nl_assertions row correctly shows. Either the variance component, the Eρ² value, or the design used to compute them is misreported; this is load-bearing because Table 1 is the sole support for the central 'σ²_a<3%' claim.","section":"Table 1"},{"comment":"The Bayesian binomial GLMM results reported in §5.1 contradict the claimed agreement among estimators. On TheAgentCompany the posterior median for σ_a is 2.19 on the logit scale (95% CrI [0.14, 3.87]) and for σ_a:t is 2.26. Squared and inserted into the paper's own capability-gap ratio, these imply approximately 2.19²/(2.19²+2.26²)≈0.48, not <0.001 as in Table 1. The sentence stating that this result is 'consistent with REML' is not supported by the reported numbers, and no scale conversion from logit-scale variance to percentage-of-total-variance is provided. The abstract's claim that all three estimators 'agree to three decimal places' is therefore unsubstantiated.","section":"§5.1"},{"comment":"§5.5 reports capability-gap ratios of 0.38 for TheAgentCompany, 0.40 and 0.35 for AppWorld, and 0.12 for τ², but Table 1 gives TheAgentCompany <0.001 and τ² values of 0.187, 0.176, 0.000, and 0.000. These cannot all be the same statistic under the definition in §3. If the §5.5 ratios use a different definition, aggregation, or set of cells, that must be stated explicitly; as written, either Table 1 or the cross-dataset transfer claim is incorrect.","section":"§5.5 and Table 1"},{"comment":"The abstract claims that σ²_a is below 3% of total variance 'in every dataset and check type', but Table 1 contains no AppWorld rows. AppWorld appears only through gap ratios in §5.5, and a gap ratio of 0.40 does not by itself bound σ²_a: if σ²_a:t is correspondingly large, σ²_a could exceed the 3% threshold. The full AppWorld variance-component table, with check-type or split breakdown, is needed to verify the headline claim.","section":"Table 1 and §4"},{"comment":"The estimation strategy is not internally coherent for binary outcomes. REML via lme4 is described as the canonical estimator for unbalanced Gaussian random-effects models, yet the observed data are binary success indicators, and Table 1's percentages appear to come from such a fit. A Gaussian linear mixed model on 0/1 outcomes produces variance components on the observed probability scale, which are not directly comparable to the logit-scale variance components of the binomial GLMM reported in §5.1 without a link-function transformation. The paper never justifies this comparison, and the discrepancy between Table 1 and the Bayesian estimates is likely rooted in this scale mismatch.","section":"§3 and Table 1"}],"minor_comments":[{"comment":"Henderson Method-I is announced as one of the three estimators in the abstract and §3, but no Method-I variance-component estimates are reported in Table 1 or anywhere else; the paper should either include these results or revise the 'three estimators' claim.","section":"§1 and §3"},{"comment":"The related-work section states that MAST/MAD contains 1,642 multi-agent traces, while §4 and §5.6 refer to 1,242 LLM-judged traces; this numerical discrepancy should be resolved or explained.","section":"§2 and §4"},{"comment":"The sentence 'All REML and CPU posteriors agree to 3–4 decimal places with the GPU posteriors' conflates a frequentist estimator (REML) with Bayesian posteriors; this needs rewriting, and the agreement claim should be supported by a comparison table rather than asserted.","section":"§5.1"},{"comment":"The reported Spearman correlation of −0.50 for three shared families is described as descriptive, which is appropriate, but the paper should also state the confidence interval or explicitly note that the estimate is compatible with a wide range of true correlations, given n=3.","section":"§5.5"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claim is currently unverifiable because of the Table 1 inconsistency and the unexplained mismatch between the Bayesian and REML variance components. I would ask the authors to supply a corrected Table 1 that includes AppWorld, a detailed scale-conversion appendix for the binomial GLMM, and a full three-estimator comparison table before considering acceptance. The open-source release and the DDR framework are valuable, and the issues appear fixable, but they are load-bearing for the headline finding."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The four-facet Generalizability Theory treatment of multi-step agent traces is new, and the paper does real work: step and error-category facets, difficulty-conditional Eρ², held-out validation, and cost-adjusted reliability are all sensible extensions that the IR/LLM-rater lineage never tried. The open code and data, the honest Limitations section, and the disclosure of the post-hoc dataset switch and late hold-out addition all suggest a careful researcher who wants to be checked. Credit where due: if the variance decomposition holds up, the finding that leaderboards mostly rank task specialization would matter for procurement decisions.\n\nThe soft spots are not minor. Table 1 reports σ²_a = 0.00% for τ2 db_check but Eρ² = 0.899. Under the paper's own formula Eρ² = σ²_a/(σ²_a + σ²_δ), a zero numerator forces Eρ² to zero. That is an internal inconsistency, not a missing footnote. Even more serious: §5.1 reports the Bayesian binomial GLMM posterior for σ_a on TheAgentCompany as median 2.19 (logit scale), with σ_a:t at 2.26. If you take the paper's own capability-gap ratio definition, that is roughly 0.48, not <0.001. The text says the Bayesian estimates \"confirm this pattern\" and agree with REML to three decimal places, but the reported medians do not support that. A scale conversion from logit variance to percentage of total variance might rescue the claim, but no conversion is given. As written, the central empirical anchor—that all three estimators agree and the agent main effect is effectively zero—is contradicted by the paper's own numbers.\n\nSmaller issues: AppWorld variance components never appear in Table 1 even though the paper reports a capability-gap ratio for it; the cross-dataset rank inversions rely on n=3; and the near-zero agent main effect on τ2 rests on three frontier agents. These are acknowledged in the Limitations, so they are weight but not decisive. The Table 1 and Bayesian mismatch are load-bearing.\n\nWho is this for? Measurement researchers and benchmark designers working on agent evaluation will get useful ideas and a cautionary example. But the paper needs a major revision: reconcile or clearly explain the Bayesian logit-scale estimates with the REML percentages, fix Table 1, and report AppWorld components. A serious referee should see it—conditional acceptance, not desk rejection—but the strongest framing should not survive without those fixes.\n\nRecommendation: send to peer review with a request for major revision.","headline":"A genuinely new G-theory application to agent traces with real empirical value, but the paper's own Bayesian estimates undercut its headline claim, and Table 1 contains an outright inconsistency that needs fixing before the strong framing can be trusted.","tokens_in":11352,"tokens_out":2403,"would_cite":false,"duration_ms":23683,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Enterprise agent leaderboards rank task specialization, not overall capability; agent identity explains under 3% of score variance.","keywords":["Generalizability Theory","agent leaderboards","variance decomposition","Deployment Decision Reliability","evaluation sizing","difficulty-conditional reliability","held-out validation","agent-task interaction"],"falsifier":"Add many more diverse agent harnesses to $\\tau^2$-bench and re-fit the Bayesian binomial GLMM used in the paper; if the posterior for the agent main effect implies $\\sigma^2_a$ above 3% of total variance with a 95% credible interval excluding zero, the capability-ceiling claim—and with it the 'leaderboards rank specialization' conclusion—is falsified.","tokens_in":10294,"feed_emoji":"📊","tokens_out":14006,"duration_ms":95055,"temperature":0.7,"pith_summary":"The paper argues that enterprise agent leaderboards are being read as capability rankings when they are actually specialization maps. It applies a four-facet variance decomposition to three open agent-trace benchmarks and finds that the agent's own main effect accounts for less than 3% of total variance in every dataset and check type, while the agent-by-task interaction accounts for 7–23%. That means the order of names on a leaderboard mostly reflects which tasks each agent handles well, not which agent is broadly more capable. If correct, this changes procurement practice: buyers should size evaluations around their own difficulty mix, cost constraints, and holdout protocol, and treat the variance-component table as a routing guide rather than a ranking. The paper packages these diagnostics into a one-page reporting discipline called Deployment Decision Reliability.","feed_headline":"Agent leaderboards rank specialization, not capability","feed_subtitle":"Across three benchmarks, agent identity is under 3% of variance; task interactions carry 7-23%.","key_machinery":"The central object is the four-facet Generalizability Theory decomposition, a crossed random-effects model $Y_{atse}=\\mu+\\nu_a+\\nu_t+\\nu_s+\\nu_e+\\nu_{at}+\\nu_{as}+\\nu_{ae}+\\nu_{ts}+\\nu_{te}+\\nu_{se}+\\varepsilon_{atse}$ that attributes variance in binary success indicators to agent, task, step, error category, and their interactions. The agent is the object of measurement, so $\\sigma^2_a$ is the signal; the capability-gap ratio $\\sigma^2_a/(\\sigma^2_a+\\sigma^2_{a:t})$ is the diagnostic that separates uniform capability (ratio near 1) from task specialization (ratio near 0). The companion Decision Study converts the variance components into the minimum number of task observations needed for a target reliability, and the paper adds three extensions: difficulty-stratified reliability per task quartile, a 50-split 70/30 held-out protocol, and a cost-adjusted Reliability-per-Dollar index.","core_discovery":"The central empirical claim is that the agent main effect $\\sigma^2_a$ is below 3% of total variance in every (dataset, check-type) cell and is exactly zero on $\\tau^2$ db_check and nl_assertions, while the agent-by-task interaction $\\sigma^2_{a:t}$ is 7.5–12.5% of total variance (7–23% across settings). Leaderboard rank order therefore reflects specialization across tasks, not a uniform capability advantage. Four corollaries sharpen the picture: aggregate generalizability $E\\rho^2$ collapses on the hardest task quartile (0.752 to 0.000 on $\\tau^2$ action_checks); training-cell $E\\rho^2$ correlates negatively with held-out $E\\rho^2$ across 50 random splits ($r=-0.90$ on $\\tau^2$); population-level diagnostics transfer across enterprise benchmarks (capability-gap ratio stable at 0.35–0.40) while per-family rankings invert (Spearman $\\hat\\rho=-0.50$, $n=3$); and on the MAST failure taxonomy, trace-level mode profiles are idiosyncratic (MAE=0.261) while cell-level profiles generalize (MAE=0.056, $r=0.83$).","pith_inferences":["Beyond the paper, if the capability-ceiling result holds beyond the current agent population, the natural unit of procurement becomes a routing table: for each task class, buy the agent that owns that cell, and treat the per-task-class $\\sigma^2_{a:t}$ table as the routing policy.","Beyond the paper, the negative training-versus-held-out correlation suggests a general optimism bias in small-population Generalizability Theory estimates; a dedicated simulation varying the number of agents and cell imbalance could quantify that bias and lead to a corrected estimator.","Beyond the paper, the variance-versus-frequency orthogonality implies that failure-mode reports should separate 'where failures occur' from 'where agents differ'; a benchmark that reports only failure activation rates will misdirect evaluation effort.","Beyond the paper, the same four-facet decomposition could be applied with rater identity as the fourth facet to test whether LLM-as-judge variance behaves like task variance in other benchmarks."],"forward_implications":["The headline score on a leaderboard should not be treated as a capability ordering: a three-point gap between two frontier agents on $\\tau^2$ action_checks is dominated by which tasks the benchmark sampled, not by which agent is more capable.","Aggregate reliability numbers are not portable to hard-task deployments: on $\\tau^2$ action_checks, $E\\rho^2$ falls from 0.752 on the full benchmark to 0.000 on the hardest quartile.","Evaluation designs that report only training-cell reliability overstate replication reliability; the held-out estimate is systematically worse, with $r=-0.90$ between projected and held-out $E\\rho^2$ on $\\tau^2$.","Cost-aware procurement can invert accuracy ranks: under Reliability-per-Dollar, o4-mini overtakes GPT-4.1 on $\\tau^2$ action_checks despite lower accuracy, so cost constraints should enter the ranking explicitly.","Benchmark vendors should report difficulty-conditional, held-out, and cost-adjusted reliability by default rather than only aggregate accuracy."],"supporting_citations":[{"why":"Supplies the Generalizability Theory variance decomposition and Decision-Study machinery the paper extends.","marker":"[Cronbach et al., 1972]"},{"why":"Provides the closed-form Method-I variance-component estimator used as the first of the three estimation checks.","marker":"[Henderson, 1953]"},{"why":"Provides the lme4 REML estimator used for the headline variance components.","marker":"[Bates et al., 2015]"},{"why":"Provides the Bambi interface for the Bayesian binomial GLMM estimator.","marker":"[Capretto et al., 2022]"},{"why":"Provides the NumPyro inference engine behind the Bayesian credible intervals.","marker":"[Phan et al., 2019]"},{"why":"Releases TheAgentCompany, one of the three primary agent-trace datasets.","marker":"[Xu et al., 2024]"},{"why":"Releases $\\tau^2$-bench, the dataset with per-check-type reward schemas and the hardest-quartile collapse evidence.","marker":"[Barres et al., 2025]"},{"why":"Releases AppWorld, the third primary dataset and the app-domain unit-test surface.","marker":"[Trivedi et al., 2024]"},{"why":"Provides the MAST taxonomy and MAD dataset used for the failure-mode generalization analysis.","marker":"[Cemri et al., 2025]"}],"fun_headline_variants":["Leaderboards rank specialization, not agent capability","Under 3% agent effect: leaderboards are task-specialization maps","Agent-by-task interaction, not the agent, drives leaderboard ranks","Why agent leaderboards actually rank task specialization","The 3% truth: agent leaderboards rank specialization, not skill"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the near-zero agent main effect is a property of the current population of frontier agents, not a small-sample artifact—on $\\tau^2$-bench only three agents are compared, and the Bayesian credible interval for the agent component is wide ($[0.14, 3.87]$ on the logit scale).","fun_headline_variants_meta":{"raw":{"variants":["Leaderboards rank specialization, not agent capability","Under 3% agent effect: leaderboards are task-specialization maps","Agent-by-task interaction, not the agent, drives leaderboard ranks","Why agent leaderboards actually rank task specialization","The 3% truth: agent leaderboards rank specialization, not skill"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001103,"raw_usage":{"total_tokens":4693,"prompt_tokens":1134,"completion_tokens":3559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":750,"completion_tokens_details":{"reasoning_tokens":3475}},"tokens_in":750,"tokens_out":3559,"duration_ms":25590,"temperature":1.0,"reasoning_tokens":3475,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:51.326155+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Add many more diverse agent harnesses to $\\tau^2$-bench and re-fit the Bayesian binomial GLMM used in the paper; if the posterior for the agent main effect implies $\\sigma^2_a$ above 3% of total variance with a 95% credible interval excluding zero, the capability-ceiling claim—and with it the 'leaderboards rank specialization' conclusion—is falsified.","supporting_citations":[],"review_version":1}