REVIEW 2 major objections 4 minor 21 references
Most of the apparent shift of AI gains toward hard tasks is a measurement artifact; a smaller real hard-task effect survives.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 00:38 UTC pith:UJHQLWEP
load-bearing objection Solid deflation result, fragile residual claim: the +0.40 logit hard-item effect is conditional on a discrimination pin that the paper never actually estimates on real data. the 2 major comments →
CurveShift: Is Agent Progress Scalar? Separating Level from Shape
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that most reported acceleration on hard tasks is not a qualitative change in capability but the mechanical consequence of a single rising ability parameter acting on a fixed logistic curve: bands near the ceiling compress their observable slope, and intermediate bands sit on the steep part, so the locus of fastest improvement migrates toward harder tasks as ability grows. After freezing ability on easy and medium items, the paper estimates a residual hard-item era effect δ: models released after September 2024 solve the hardest LiveCodeBench problems beyond what their easy and medium performance predicts. With hard-item discrimination pinned equal to the rest, δ ≈ +0.40
What carries the argument
CurveShift, an anchored two-parameter logistic (2PL) sensitivity design: ability θ is estimated only from easy and medium items and then frozen; hard items enter through per-item fixed effects; and hard-item discrimination α_hard is pinned over a grid rather than freely estimated, because a free 2PL is not stably identified (discrimination and the era DIF trade off along a near-flat ridge). The Rasch model with a single rising ability serves as the scalar null that reproduces the apparent locus migration.
Load-bearing premise
The headline residual of +0.40 logits assumes hard items are no more discriminating than easy and medium ones (α_hard = 1.00); if hard items were actually more than about a fifth more discriminating (α_hard > 1.22), the calibrated effect would cross zero and disappear.
What would settle it
Estimate hard-item discrimination α_hard from an independent source — e.g., human solver response times or a psychometric calibration on a no-scaffold benchmark where humans and models take the same items — and check whether it exceeds ≈1.22. If it does, the +0.40 logit residual vanishes under the paper's own sensitivity curve. Alternatively, a new no-scaffold math benchmark with human-completion-time difficulty and pre-2024 model coverage finding no residual hard-item gain would contradict the claim.
If this is right
- Progress reporting should separate level from shape: a scalar metric like a time-horizon doubling time cannot distinguish uniform progress from a shifting hard tail, so difficulty-stratified curves should accompany any scalar summary.
- On agentic benchmarks, model era and scaffold era are confounded; claims about post-era gains on hard tasks from such data should be read with caution, and no-scaffold measurements preferred when the claim concerns the model itself.
- The real hard-item effect is specific to short-reasoning, verifiable competitive programming tasks, not long-horizon autonomy; it does not support the narrative that the capability frontier is broadly reshaping toward hard tasks.
- The effect survives multiple robustness checks (anchor choice, leave-one-family-out, date perturbation, contamination filtering, continuous-time version), narrowing but not eliminating the residual.
Where Pith is reading between the lines
- If this separation holds, capability forecasting should move away from scalar extrapolations: the same aggregate trend can hide either uniform progress or a growing hard tail, and the two have different implications for when specific hard tasks become solvable.
- The same deflation logic could be applied to other no-scaffold domains with exogenous difficulty, such as competition mathematics; the paper notes a math attempt failed on data preconditions, which suggests a human-solve-rate anchor would be the needed next dataset.
- The +0.40 logit residual being concentrated in reasoning models and short-reasoning tasks is consistent with test-time compute helping most where outcomes are verifiable, but the paper explicitly does not test this; a direct test would vary reasoning budget within a fixed model while holding sampling constant.
- The attempt-count confound (pre-era 10 draws vs post-era 1 draw) is handled by simulation, but the finding that equal-weighting reverses sign warns that naive reanalyses of public leaderboards could easily misread the effect as zero or negative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that scalar summaries of LLM progress conflate an overall level shift with a change in the shape of the difficulty-response curve. On METR time-horizon data, it claims that a single Rasch model with rising ability largely reproduces the apparent migration of gains toward harder tasks, so most of that migration is a ceiling/discrimination artifact. On LiveCodeBench, which has no agentic scaffold and uses exogenous human difficulty labels, the paper freezes ability on easy/medium items, fits per-item fixed effects, and estimates a hard-item era effect δ while pinning rather than estimating the hard-item discrimination α_hard. Under the equal-discrimination assumption α_hard=1.00, the calibrated headline is δ≈+0.40 logits (raw +0.45), implying a hard-problem solve-rate rise from about 18% to 25%. The effect is reported to survive many robustness checks. The paper releases the LiveCodeBench Difficulty Panel and analysis code.
Significance. If the result holds, the paper makes a useful methodological and empirical contribution: it offers a concrete way to separate level from shape, shows that the widely reported hard-task acceleration is mostly a scalar artifact, and identifies a small but nonzero residual hard-item effect in competitive programming. The design is careful in several respects: ability is frozen on easy/medium items, difficulty labels are human-assigned and exogenous to the models, discrimination is pinned rather than freely fit, generated-regressor bias is calibrated by simulation on the real attempt counts, and uncertainty is addressed with a ladder of cluster bootstraps. The release of the panel and code is a positive feature for reproducibility. The main weakness is that the headline 'most conservative assumption' depends on α_hard=1.00 being the safe end of the grid, and the paper does not provide an independent, real-data estimate of hard-item discrimination to support that placement.
major comments (2)
- [Section 5, Fig. 4, §A.3] The headline δ≈+0.40 is called the 'most conservative' value because α_hard=1.00 is treated as the safe end of the grid. This is not established. The calibrated δ crosses zero at α_hard≈1.22, and the only cited support for placing the true value below 1.0 is the '0.83' estimate, which §A.3 does not actually compute on the LiveCodeBench panel. §A.3 is a simulation showing that when true α=1.0, the continuous binomial likelihood recovers 0.83; it is an estimator property, not a real-data estimate. Step 3 of §3.2 calls 0.83 'a debiased estimate from the continuous likelihood' without giving the fitting procedure or result. Given §A.1's degeneracy and §A.2's demonstration that a null 2PL absorbs DIF into inflated hard-item discrimination, the real-data provenance of 0.83 is critical. If 0.83 came from any of the models discussed in the appendix, it is not independent evidence. Without an ind
- [Section 4, Table 2] The first conclusion — that a single Rasch model with rising ability reproduces the METR locus migration — is not accompanied by the estimation details needed to check it. The text reports a fitted 'After, Rasch null' column in Table 2 and Figure 3, but does not state how the ability trajectory θ(t) is parameterized, how item difficulties are anchored, how the model is fit to the per-band success counts, or what the goodness of fit is. The claim that the null 'places the fastest-improving band at 15–60 minutes' is therefore an assertion. Since this is the basis for the abstract's statement that most apparent acceleration is a ceiling/discrimination artifact, the reader needs this information, or a reference to an appendix where it appears.
minor comments (4)
- [§A.3 vs §3.2/Table 3] The lower grid point is 0.55, called 'deflated by binarization', but §A.3's simulation reports binarization recovering α=0.49. Please reconcile these numbers in the text.
- [Fig. 4 caption] The caption states 'the calibrated estimate of hard-item discrimination on these data is 0.83'. As written this appears to be a real-data estimate, but the derivation is not given in the main text or in §A.3 (which is a simulation). Reword or provide the real-data estimation procedure.
- [Section 8] The limitation that a fully Bayesian hierarchical 2PL with a shrinkage prior on discrimination is left to future work is stated in §8, but it directly qualifies the 'most conservative' language used in §5. Consider foregrounding this caveat when the sensitivity grid is introduced.
- [§5, contamination filtering] Typo: 'thengeometry' should be 'the geometry'.
Circularity Check
Headline +0.40 logit effect is a genuine residual, but the 'most conservative' α_hard=1.00 anchor is supported by a simulation whose input is α_hard=1.00, making the conservative framing partially self-referential.
specific steps
-
other
[Section 5 (Fig. 4 caption; Table 3), Section 3.2 Step 3, Section A.3]
""The value αhard = 0.83 is a debiased estimate from the continuous likelihood" (Sec 3.2); "the calibrated estimate on these data is 0.83, below one" (Sec 5); "In our simulation with true discrimination α=1.0, binarization recovered ˆα=0.49, whereas the continuous binomial likelihood recovered ˆα=0.83" (A.3)."
The main text presents 0.83 as data evidence that hard-item discrimination is below one, justifying α_hard=1.00 as the 'most conservative' grid point and hence the +0.40 headline. But the only 0.83 in the appendix comes from a simulation whose true value is α_hard=1.00, the very assumption being supported. No real-data α_hard estimate is shown (and A.1/A.2 argue a free fit is degenerate or biased). The zero-crossing at α≈1.22 is measured on the same panel, and the documented downward bias (true 1.0 recovered as 0.83) means 'no evidence above one' is not established. The conservative anchor is therefore not externally forced; the headline's sign survives only under an assumption whose support is generated from the same null used to define the residual.
full rationale
The core CurveShift derivation is largely self-contained and not circular: difficulty labels are human-assigned and exogenous to the models (Sec 3.1), ability θ is frozen from easy/medium items so hard outcomes cannot feed back into the level (Sec 3.2), the hard-item effect δ is a fitted residual rather than a predicted value, and the positive-control simulation (A.4) plus cluster bootstraps and leave-one-family-out checks give the result independent content. The Rasch deflation of METR (Sec 4) is a standard in-sample null comparison, not a self-definitional prediction. No load-bearing self-citation appears: the Xing et al. (2026) citation in Limitations is peripheral, and the classical psychometric citations (Rasch, Birnbaum, Holland & Wainer) are external. The single concern is the support for the 'conservative' α_hard=1.00 anchor: the paper cites Section A.3 as if it provided a data-based estimate of 0.83, but A.3 only reports a simulation value under true α=1.0. That makes the claim 'we find no evidence above one' an internal assumption rather than independent evidence, and the zero-crossing at α≈1.22 lies within plausible uncertainty. This is a partial circularity in the interpretation of the headline, not a collapse of the derivation; the raw residual is not manufactured. Hence score 2.
Axiom & Free-Parameter Ledger
free parameters (5)
- Reasoning-era breakpoint (post indicator) =
2024.70 (September 2024)
- Hard-item discrimination pin alpha_hard =
0.55 / 0.83 / 1.00; headline 1.00
- Generated-regressor calibration offset and slope =
offset +0.078 at alpha=1.00; slope ~0.94
- Rasch null ability trajectory theta(t) on METR =
not reported numerically
- Per-model abilities theta_m and item fixed effects c_i =
estimated from data
axioms (6)
- domain assumption Human-assigned item difficulty labels (easy/medium/hard; completion-time bands) are exogenous and invariant across model eras.
- domain assumption LiveCodeBench involves no agentic scaffold, so model era and harness era are not collinear there.
- domain assumption A single-ability Rasch model is an adequate null for METR per-band success; uniform growth of theta generates the predicted slopes.
- standard math Continuous binomial likelihood with attempt counts as weights is the correct sampling model; binarization deflates hard-item discrimination.
- domain assumption alpha_hard = 1.00 (equal discrimination) is the conservative end of the defensible grid; true hard-item discrimination is not above approximately 1.22.
- standard math Cluster bootstrap over items/contests/models is valid; model-family level has too few clusters (11) and is underpowered.
read the original abstract
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.
Figures
Reference graph
Works this paper leans on
-
[1]
Program synthesis with large language models.arXiv preprint arXiv:2108.07732,
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models.arXiv preprint arXiv:2108.07732,
-
[5]
Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The benchmark lottery.arXiv preprint arXiv:2107.07002,
-
[6]
OckBench: Measuring the efficiency of LLM reasoning.arXiv preprint arXiv:2511.05722,
Zheng Du, Hao Kang, Song Han, Tushar Krishna, and Ligeng Zhu. OckBench: Measuring the efficiency of LLM reasoning.arXiv preprint arXiv:2511.05722,
-
[8]
A Rosetta Stone for AI benchmarks.arXiv preprint arXiv:2512.00193,
Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, and Rohin Shah. A Rosetta Stone for AI benchmarks.arXiv preprint arXiv:2512.00193,
-
[9]
Lalor, Hao Wu, and Hong Yu
John P. Lalor, Hao Wu, and Hong Yu. Building an evaluation scale using item response theory. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 648–657,
2016
-
[13]
doi: 10.18653/v1/2021.acl-long.346
Association for Computational Lin- guistics. doi: 10.18653/v1/2021.acl-long.346. URLhttps://aclanthology.org/2021.acl-long.346/. Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage?Advances in Neural Information Processing Systems, 36:55565–55581,
-
[14]
URLhttps: //doi.org/10.1145/3715754
doi: 10.1145/3715754. URLhttps: //doi.org/10.1145/3715754. Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. OSWorld: Benchmarking multimodal agents for open- ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37: 52040–52094,
-
[16]
Scaling test-time compute for LLM agents.arXiv preprint arXiv:2506.12928,
King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, Dehua Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jiaheng Liu, Yuchen Eleanor Jiang, Changwang Zhang, Chenghua Lin, Jun Wang, Ge Zhang, and Wangchunshu Zhou. Scaling test-time compute for LLM agents.arXiv preprint arXiv:2506.12928,
-
[17]
18 A.1 The free 2PL is not stably identified jointly with the era DIF coefficient The two-parameter logistic model positslogitp mi =α i(θm−β i), with discriminationα i
A Identifiability of the free 2PL and the bias of a naive residual test This appendix documents the methodological pitfalls we encountered and ruled out, so that reviewers can verify that the obvious alternatives were considered and rejected for principled reasons. 18 A.1 The free 2PL is not stably identified jointly with the era DIF coefficient The two-p...
1984
-
[19]
We used the MathArena panels for the 2025 contests (Balunovic et al.,
The natural candidate is competition mathematics without an agentic scaffold, where models answer directly and the contest origin gives a difficulty ordering. We used the MathArena panels for the 2025 contests (Balunovic et al.,
2025
-
[20]
The attempt does not fail a hypothesis test; it fails before one can be run, on two of the four conditions that make LiveCodeBench usable
(AIME, HMMT February, BRUMO, SMT, and CMIMC). The attempt does not fail a hypothesis test; it fails before one can be run, on two of the four conditions that make LiveCodeBench usable. The first missing condition is a clean exogenous difficulty. The only per-item ordering these panels carry is the problem index, which is an endogenous and noisy proxy. A g...
2025
-
[21]
23 Model Date Family Era Reasoning DSCoder-1.3b-Ins 2023.86 DeepSeek pre no Claude-3-Haiku 2024.20 Anthropic pre no GPT-4-Turbo-2024-04-09 2024.27 OpenAI pre no GPT-4O-2024-05-13 2024.37 OpenAI pre no Codestral-Latest 2024.42 Mistral pre no Qwen2-Ins-72B 2024.43 Qwen pre no Claude-3.5-Sonnet-20240620 2024.47 Anthropic pre no Mistral-Large 2024.55 Mistral ...
2023
-
[1959]
The common odds ratio is1.84(log odds+0.61), favoring post-era models, in the direction and rough magnitude of the parametric estimate
analysis matches models on their easy and medium ability, then compares pre-era and post-era performance on hard items within ability strata, with no functional form. The common odds ratio is1.84(log odds+0.61), favoring post-era models, in the direction and rough magnitude of the parametric estimate. The matching variable overlaps only in the lower abili...
2021
-
[1968]
Le, Christopher Ré, and Azalia Mirhoseini
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling.arXiv preprint arXiv:2407.21787,
-
[1980]
Data contamination through the lens of time.arXiv preprint arXiv:2310.10628,
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. Data contamination through the lens of time.arXiv preprint arXiv:2310.10628,
- [1985]
-
[2008]
Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374,
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374,
-
[2021]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168,
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168,
-
[2024]
Shanghaoran Quan, Jiaxi Yang, Bowen Yu, Bo Zheng, Dayiheng Liu, An Yang, Xuancheng Ren, Bofei Gao, YiboMiao, YunlongFeng, ZekunWang, JianYang, ZeyuCui, YangFan, YichangZhang, BinyuanHui, and Junyang Lin. CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable Elo ratings.arXiv preprint arXiv:2501.01257,
-
[2025]
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, and Kaivalya Hariharan. Agent psychometrics: Task-level performance prediction in agentic coding benchmarks. InICLR 2026 Workshop on Agents in the Wild (AIWILD),
2026
-
[2026]
ACM. doi: 10.1145/3770855.3818652. To appear. John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. SWE-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.