REVIEW 3 major objections 5 minor 1 cited by
Research evaluation with ChatGPT: Is it age, country, length, or field biased?
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper shows that ChatGPT's research-quality scores rise systematically with publication year in all 26 fields tested, and argues the scores must be normalised for field and year before being used in research evaluation.
desk verdict Solid large-scale evidence that ChatGPT scores rise with publication year in all fields, but the 'bias' label rests on an untested constant-quality assumption; still worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a journal-stable, field-balanced sample combined with regression. Up to 1,000 articles were drawn per year and per field from journals that belonged exclusively to one of 26 subject categories and had published in all five sampled years, filtering out short or missing abstracts; this produced 117,650 articles. Each title and abstract was scored once by ChatGPT 4o-mini with an expert-review quality rubric, and ordinary least squares regression was run per field with publication year, log abstract length, and ten first-author country indicators as predictors. The sampling design is what makes a year trend interpretable as a candidate bias rather than as database drift.
What would settle it
Have an expert panel blind-score a random subset of the same 117,650 articles; if expert ratings are flat across years while ChatGPT scores rise, the recency effect is a bias. Or submit the same abstracts with contemporary vocabulary swapped for neutral phrasing; if scores rise with the modern wording, the year effect is carried by textual cues rather than by true quality.
Extended reading notes
Core claim
The central discovery is a universal recency effect: for all 26 broad fields, the average ChatGPT score for 2023 was higher than for 2003, and in 101 of 104 year-to-year comparisons each later year beat the previous one. In all 26 field-level regressions the year coefficient was positive and statistically significant even after controlling for abstract length and first-author country. Field averages differ substantially, with some fields scoring above all others in every year, and ChatGPT scores correlate positively with citation counts in all fields. The paper reads the year trend as bias because the sample was drawn from the same journals across years, so it assumes article quality was roughly constant over time.
Load-bearing premise
The load-bearing premise is that the average true quality of articles in the same core journals stayed roughly constant from 2003 to 2023; the paper admits this is unproven, and if quality genuinely improved then the higher scores for newer articles would be accurate rather than biased.
Editorial extensions
If this is right
- If the paper is right, any use of ChatGPT scores in research evaluation should first divide each article's score by the average score for its field and year, mirroring standard citation-normalisation practice.
- The year bias is small but universal: it explains on average 3.6% of score variance, so it will distort comparisons of older and newer work unless corrected.
- The abstract-length association is probably not a direct ChatGPT bias; it seems to reflect weaker short-form articles and national journals, but fields with strict abstract limits should be checked for anomalies.
- Because ChatGPT scores correlate positively with citation counts in all fields, the paper strengthens the broader claim that LLM quality scores carry signal about research quality, even though the correlations are modest.
Reading between the lines
- One consequence the paper leaves implicit is that the same year trend may affect other language-model evaluators, so the normalisation advice likely extends beyond this one model and prompt.
- The size of the year effect could be underestimated or overestimated if article quality in these journals genuinely improved over two decades; a direct expert-quality benchmark on the same sample would settle whether 'bias' is the right word.
- The study's exclusion of multidisciplinary journals and of articles with short abstracts means the biases for high-profile venues and for letter-type outputs remain untested; those are exactly the outputs where peer review support is most contested.
- Because ChatGPT was not explicitly told publication years, the recency effect must be carried by textual cues correlated with time; identifying those cues could allow targeted debiasing rather than blanket normalisation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a large-scale measurement study in which 117,650 articles from 26 Scopus broad fields and five publication years (2003, 2008, 2013, 2018, 2023) were scored by ChatGPT 4o-mini using REF-style evaluation guidelines. The central descriptive findings are that ChatGPT scores increase with publication year in all 26 fields, that the increase survives controls for first-author country and title/abstract length, and that average scores differ substantially between fields. The paper also reports associations between scores and abstract length, first-author country, and citation counts. The authors interpret the year and field differences as biases and recommend normalizing ChatGPT scores by field and year before use in research evaluation.
Significance. If the year and field differences are genuine biases, the paper makes an important practical contribution: it would show that ChatGPT-based research quality scores cannot be compared across time or fields without normalization, and it would extend the known bias literature for LLM-based evaluation. The study has notable strengths: a large and carefully constructed journal-balanced sample, transparent random sampling, explicit control variables in regression analyses, and prior validation of the REF-based prompt in other work. The positive correlations with citation counts provide useful context. The main weakness is that the central inferential step—from a monotone time trend to the claim that the trend is a bias—depends on the unproven assumption that average intrinsic quality in the sampled journals was roughly constant over the period. The paper itself acknowledges this assumption is unproven, which makes the normalization recommendation conditional on an untested premise.
major comments (3)
- [Methods, Data and Conclusions] The inference that the positive year coefficient is a bias rather than a reflection of genuine quality change requires that the average intrinsic quality of articles in the sampled journals was approximately constant between 2003 and 2023. The paper explicitly states: "the quality of journals seems to be relatively stable. This is unproven." No human expert baseline is available for the same articles, so the positive year coefficient in all 26 field regressions is equally compatible with a real improvement in research quality over time. This assumption is load-bearing because it converts a descriptive trend into the recommendation that ChatGPT scores be normalized for year. I would like to see either (a) evidence that average expert-assessed quality was stable in these journals over the period, (b) a subsample with human scores on the same articles, or (c) a clear reframing of the conclusion as documenting a year association rather than a demonstrated bias.
- [Methods, Data] The minimum abstract length threshold of 786 characters discards 25% of Scopus articles. If abstract-length norms have changed over the 20-year period, the truncation removes different fractions of articles in different years and can create, amplify, or mask a year gradient. The paper does not report the fraction of excluded articles by year and field, nor does it provide a sensitivity analysis at lower thresholds. Because both the year effect and the abstract-length effect are central to the paper (RQ1 and RQ3), this design choice is not merely a detail. Please report the distribution of excluded articles and re-estimate the main regressions with a lower threshold or an explicit selection correction.
- [Regression] The regression analysis treats ChatGPT scores as an interval scale and fits publication year as a single linear term. This imposes a linear trend across the five discrete years, even though the paper's headline pattern is that each successive year scores higher than the previous one in 101 of 104 cases. Fitting year as a categorical variable (or using an ordinal model) would show whether the increase is uniform or concentrated in particular years, which matters for the recommended field-year normalization. If the trend is nonlinear, the normalization procedure may need to be adjusted accordingly.
minor comments (5)
- [Regression] The word "dependant" should be "dependent".
- [ChatGPT procedure] The citation "Thelwall, 2024ab" combines two distinct references; the text should cite "Thelwall, 2024a, 2024b" individually.
- [RQ1-4: Regression results] Figure 4's error bars are described as the minimum and maximum values from the 26 field estimates, but it is unclear whether the plotted coefficients are standardized; please clarify because year and log-length coefficients are not on the same scale as binary country indicators.
- [Methods, Data] The choice of 786 characters as the abstract-length threshold is justified heuristically; a sensitivity check around this threshold would strengthen the paper, as noted in the major comments.
- [RQ3] The discussion of the abstract-length association states that "the second and third options play a role" but does not quantify the relative contribution of journal-level quality effects versus short-form content; this remains an interpretation rather than a formal test.
Circularity Check
No circularity: the year-bias result is a direct measurement, not a derivation from fitted inputs; the unproven constant-quality assumption is an external validity limitation, not a circular step.
full rationale
This is a descriptive measurement study, not a derivation. The observed year trend is computed directly from ChatGPT output on a designed sample, and the regression includes year, abstract length, and first-author country as independent variables. No fitted parameter is renamed as a prediction, and the year coefficients are not constrained by any normalization to equal the target result. The paper explicitly concedes that the assumption of stable journal quality is unproven, but this is a limitation of external validity, not circularity: the year trend would hold even if the assumption were false. The self-citations to Thelwall and Yaghi (2024) and Thelwall (2024ab) are used as external evidence that ChatGPT scores correlate with expert quality and to justify reusing the same prompt; they are not internal inputs to the bias regressions, and no equation in the paper is equivalent by construction to another claimed output. No uniqueness theorem or author-imposed ansatz is invoked to force the conclusion. The paper's central claim, that ChatGPT scores increase with publication year, is self-contained and statistically supported; the step from trend to bias depends on an untested assumption, but that is a substantive empirical caveat rather than a circularity.
Assumptions & free parameters
free parameters (2)
- Minimum abstract length threshold =
786 characters
- Top-ten first-author country set =
ten binary variables
assumptions (4)
- domain assumption The average intrinsic quality of articles published in the same core set of journals is approximately stable across the five sampled years.
- domain assumption Scopus's 26 non-general subject categories are a valid field partition for normalization.
- domain assumption REF assessor guidelines are a usable operationalization of research quality for all included fields.
- domain assumption One ChatGPT submission per article yields scores that are stable enough for the aggregate analyses.
Cite this review
Pith. "Pith review of Research evaluation with ChatGPT: Is it age, country, length, or field biased?." pith.science (2026). https://pith.science/paper/SDQ2XJ67
@misc{pith2026241109768,
author = {Pith},
title = {Pith review of: Research evaluation with ChatGPT: Is it age, country, length, or field biased?},
year = {2026},
howpublished = {\url{https://pith.science/paper/SDQ2XJ67}},
note = {Machine review of arXiv:2411.09768}
}
read the original abstract
Some research now suggests that ChatGPT can estimate the quality of journal articles from their titles and abstracts. This has created the possibility to use ChatGPT quality scores, perhaps alongside citation-based formulae, to support peer review for research evaluation. Nevertheless, ChatGPT's internal processes are effectively opaque, despite it writing a report to support its scores, and its biases are unknown. This article investigates whether publication date and field are biasing factors. Based on submitting a monodisciplinary journal-balanced set of 117,650 articles from 26 fields published in the years 2003, 2008, 2013, 2018 and 2023 to ChatGPT 4o-mini, the results show that average scores increased over time, and this was not due to author nationality or title and abstract length changes. The results also varied substantially between fields, and first author countries. In addition, articles with longer abstracts tended to receive higher scores, but plausibly due to such articles tending to be better rather than due to ChatGPT analysing more text. Thus, for the most accurate research quality evaluation results from ChatGPT, it is important to normalise ChatGPT scores for field and year and check for anomalies caused by sets of articles with short abstracts.
Figures
Forward citations
Cited by 1 Pith paper
-
Who Gets Recommended? Investigating Gender, Race, and Country Disparities in Paper Recommendations from Large Language Models
LLM recommendations of important AI research favor recent, well-cited, team-authored papers, but do not measurably over-represent male, white, or developed-country scholars relative to a human-curated benchmark.
Reference graph
Works this paper leans on
-
[1]
Adams, J. (1998). Benchmarking international research. Nature, 396(6712), 615-618. Aksnes, D. W., Schneider, J. W., & Gunnarsson, M. (2012). Ranking national research systems by citation indicators. A comparative analysis using whole and fractionalised counting methods. Journal of Informetrics, 6(1), 36-43. Barrere, R. (2020). Indicators for the assessmen...
work page 1998
-
[2022]
https://assets.publishing.service.gov.uk/media/628cd2828fa8f55615524e8c/internati onal-comparison-uk-research-base-2022-accompanying-note.pdf Hicks, D., Wouters, P., Waltman, L., De Rijcke, S., & Rafols, I. (2015). Bibliometrics: the Leiden Manifesto for research metrics. Nature, 520(7548), 429-431. Langfeldt, L., Nedeva, M., Sörlin, S., & Thomas, D. A. (...
arXiv 2015
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.