{"id":"96892ea1-4688-4b07-bc6a-856dc5d24ae3","arxiv_id":"2506.07748","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review arguing LLM-based quality scores could surpass bibliometrics in accuracy and coverage, but with unknown biases and gaming risks that currently block real-world use.","lead":"This paper reviews how Large Language Models compare to citation-based metrics for judging research quality. It argues LLMs could be more accurate and cover more fields, but warns that unknown biases and the risk of authors overselling abstracts currently block their use in major evaluations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'more accurate than bibliometrics' claim lacks any same-sample, same-level comparison; quoted correlations with departmental REF averages are not compared against citation indicators on the same data, so the superiority claim is not yet established.","rationale":"The paper is honest and clearly labels its main empirical support as 'not conclusive' and 'suggestive', and it concedes the leakage confound. My concern is not that the authors are hiding anything; it is that the headline comparison is broader than the evidence. The 'superior to bibliometrics' claim is specifically an accuracy comparison, and accuracy claims require a common yardstick. The current yardstick changes between the LLM studies (departmental averages) and the bibliometric studies (individual expert or journal-level judgements). Since averaging reduces noise, higher raw correlations with departmental averages do not establish that a per-article indicator is more accurate. A direct head-to-head at the same aggregation level is feasible with the same REF2021 data and OpenAlex citation data, so this is a resolvable gap rather than a fatal flaw. If the head-to-head comparison were to show a consistent LLM advantage, the central claim would be much stronger; if not, the conclusion should be downgraded to 'promising but unproven'. This is why the CONDITIONAL verdict remains appropriate, and no change to the reader's verdict is needed.","tokens_in":11038,"tokens_out":9447,"duration_ms":120494,"concrete_test":"Reanalyze the REF2021 data used in Thelwall & Yaghi (2024b): for each UoA, aggregate both ChatGPT scores and normalised citation scores to the department level over the same article sets, and correlate each with the published departmental REF average scores. If ChatGPT does not consistently beat the citation indicator in these same-sample, same-level correlations, the 'more accurate' advantage in the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The empirical core of the central claim is the 'More accurate' advantage, which cites Thelwall & Yaghi (2024b): ChatGPT scores correlated with departmental average REF2021 scores across most UoAs. The claimed advantage over bibliometrics is indirect, however. The citation correlations cited for bibliometrics come from separate studies, different output sets, and mostly article-level or journal-level expert judgements, whereas the ChatGPT correlations are with departmental averages. Aggregation removes individual-rater noise, so higher correlations with departmental averages do not establish that an article-level indicator is more accurate. The paper itself flags one confound (ChatGPT 'might have leveraged public information about departmental REF quality profiles'), but even without leakage, no study in the paper compares LLM scores and citation indicators against the same human quality scores on the same outputs. Until such a head-to-head comparison is reported, the statement that LLM-based evaluations are 'already superior to bibliometrics' is a plausible hypothesis, not an evidence-based conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reviews the potential of large language models (LLMs) as research quality indicators in comparison with bibliometrics, drawing on a mixture of small-, medium-, and large-scale empirical studies, mostly by the author and co-workers. It argues that LLM-based scores are more accurate than citation-based indicators (higher correlations with human scores in most fields), have broader coverage of fields and recent years, and can in principle assess more quality dimensions. It also discusses similarities (both are indirect indicators, not direct measures) and disadvantages (unknown biases, lower transparency, limited research into use contexts, and greater risk of gaming via abstract manipulation). It concludes that LLM scores have technical potential to complement or surpass bibliometrics but that there are too many unknowns for immediate use in important contexts, recommending further research on biases, limitations, and gaming before minor supporting roles are considered.","tokens_in":11184,"tokens_out":3477,"duration_ms":39050,"significance":"The paper addresses a timely and important question for the research evaluation community. Its principal strength is the systematic mapping of the systemic and normative dimensions of LLM-based indicators, especially the discussion of gaming incentives on abstracts and journal editorial practices, which goes beyond simple accuracy comparisons. The paper is also commendably explicit about the tentativeness of the evidence, using hedged language such as 'not conclusive' and 'suggestive'. The reviewed evidence includes both published and preprint studies, and the author acknowledges the leakage confound in the large-scale REF study. However, the central comparative claim of superiority over bibliometrics rests on indirect comparisons across different samples, levels of analysis, and outcome measures, and much of the supporting evidence is author-affiliated and not independently replicated. If the claim is taken as a hypothesis needing further testing, the paper is a valuable agenda-setting review; as an evidence-based conclusion, it currently overreaches.","major_comments":[{"comment":"The claim that LLM scores are 'more accurate' than bibliometrics is not established by the cited evidence because the comparisons are not on the same footing. The key support (Thelwall & Yaghi, 2024b) correlates ChatGPT scores with departmental average REF2021 scores, whereas the bibliometric correlations cited for comparison (Thelwall et al., 2023b, 2023c) are at the level of individual articles or journals against expert scores. Correlations with departmental averages benefit from aggregation that removes individual-rater noise, so higher correlations on that basis do not demonstrate that an article-level indicator is more accurate. The paper should either report a same-sample, same-level comparison or explicitly reframe the conclusion as a hypothesis pending such a test.","section":"'Advantages: More accurate'"},{"comment":"The paper acknowledges a plausible leakage confound in the main large-scale study: 'ChatGPT might have leveraged public information about departmental REF quality profiles when scoring individual articles.' This confound is not merely a peripheral weakness; it threatens the central evidence for the 'more accurate' and 'greater coverage' claims, including the notable claim that ChatGPT is 'useless only for clinical medicine.' Because the articles were selected from high- and low-scoring departments and the scores are averaged over 30 iterations, the model could plausibly exploit the public departmental averages. The manuscript should either present a direct test that rules out this mechanism (for example, comparing scores for articles from departments with similar profiles, or using a blinded protocol) or substantially soften the superiority conclusion in light of the unresolved confound.","section":"'LLM-generated research quality indicators'"},{"comment":"The comparison that citations are 'useless' for arts and humanities while ChatGPT is 'only useless for clinical medicine' is based on different studies with different operationalizations of 'useless' (correlation thresholds, levels of aggregation, and output sets). For instance, Thelwall et al. (2023b) examine article-level citation correlations with expert scores, while Thelwall & Yaghi (2024b) examine departmental-average correlations. The claim may be true, but the evidence as presented does not support the comparative assertion. The authors should either provide a comparable analysis across fields for both indicators or present this as a preliminary observation requiring direct comparison.","section":"'Advantages: Greater coverage of science'"}],"minor_comments":[{"comment":"There are inconsistent renderings of the model name: 'ChatGPT 40-mini' appears in the medium-scale study description and 'ChatGPT 4o-mini' elsewhere; please use a consistent notation.","section":"Throughout"},{"comment":"The reference 'Thelwall & Yaghi, 2024a' is cited as evidence for the claim that LLMs can assess multiple quality dimensions, but this reference is listed as 'Submitted' rather than published; the dependence on a non-peer-reviewed manuscript should be made explicit, or the claim should be attributed to a published source.","section":"LLM-generated research quality indicators"},{"comment":"The sentence 'LLM-based quality evaluations being already superior to bibliometrics' is stronger than the hedged language elsewhere in the paper ('not conclusive', 'suggestive'); rephrasing to 'potentially superior' or 'may already be superior' would better align the conclusion with the stated limitations.","section":"Conclusion"},{"comment":"The phrase 'as of September 2024' in the opening of the Disadvantages section is inconsistent with later statements referencing February 2025; please update the temporal reference for consistency.","section":"Disadvantages"},{"comment":"Several key empirical references are non-peer-reviewed preprints (e.g., Thelwall & Yaghi, 2024b; Thelwall & Jiang, 2025; Thelwall & Kurt, 2024). While preprints are acceptable in a fast-moving field, the manuscript should note the status of these sources in the text or reference list so readers can gauge the evidence base.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is authored by a leading figure in the field and the empirical evidence relies heavily on the author's own preprints. This is not itself a flaw, but the novelty of the central claim ('already superior to bibliometrics') is modest and the claim is not yet supported by an apples-to-apples comparison. The journal may wish to consider whether the review format is appropriate for pushing a comparative conclusion beyond the available evidence; the paper is more persuasive as an agenda-setting synthesis than as a demonstration of superiority. I would also note that the author has a declared role as a Distinguished Reviewers Board member of this journal, which is disclosed but may warrant editorial awareness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a review/position piece, not a new empirical study. Thelwall synthesizes his own and others' work on using ChatGPT to score research quality and frames it against bibliometrics. The fresh angle is the systemic gaming risk: if LLM scores reward overselling, authors and journals will inflate abstracts, which could degrade the academic record. That part is worth reading.\n\nThe paper does several things well. It is plainly written, separates technical from systemic issues, and is more cautious than the abstract suggests. Thelwall explicitly says the evidence is 'not conclusive' and flags the leakage confound in the large REF study—ChatGPT might have used public department-level scores. The conclusion also walks back to a recommendation of cautious pilot use. That is responsible.\n\nThe soft spot is the central claim that LLM-based evaluations are 'already superior to bibliometrics as research quality indicators.' The stress-test note is right: the superiority assertion rests on comparing correlations from different studies, with different output sets and levels. The ChatGPT correlations are with departmental REF averages; the bibliometric correlations come from separate, mostly article-level analyses. Aggregation removes rater noise, so higher correlations at department level do not show that an article-level LLM indicator beats citations. There is no same-sample, same-head-to-head test. Also, most supporting studies are the author's own preprints, not independently replicated. He acknowledges this, but the phrasing still overstates.\n\nThe bias section is honestly framed as preliminary, and the gaming argument is plausible but speculative. These are moderate problems for a review; they don't sink the paper's usefulness as a survey of a fast-moving area.\n\nWho is this for? People in scientometrics, research evaluation, and metascience who want a compact overview of the state of play in early 2025. It deserves a serious referee: the topic is timely, the author knows the literature, and the synthesis has value even if the main comparative claim needs softening and, ideally, a call for independent head-to-head evaluations. I'd send it to review and ask for a major revision on the superiority language and evidence base.\n\nI'd bring it to a reading group as a discussion piece, and I'd cite it for the systemic gaming angle, not for the accuracy comparison.","headline":"A clearly written review of LLM-based research quality evaluation, but the 'already superior to bibliometrics' claim outruns the evidence because no study in the paper compares LLM scores and citation indicators against the same human quality scores on the same outputs.","tokens_in":11692,"tokens_out":1817,"would_cite":true,"duration_ms":24480,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that LLM-generated quality scores already rival or beat citation-based indicators as research quality signals, and that the main open questions are the AI's biases and its vulnerability to strategic abstract-writing.","keywords":["research quality evaluation","large language models","ChatGPT","bibliometrics","research quality indicators","citation counts","research evaluation","gaming"],"falsifier":"A controlled experiment could settle it: take a set of articles with known expert quality scores, submit each to an LLM with its departmental affiliation and author names removed, and compare averaged scores to the expert scores. If the correlation collapses when identifying information is stripped, the public-information leakage explanation wins; if it holds, genuine text-based quality judgement is supported. A second check would modify abstracts (inflate or deflate quality claims) while keeping the underlying research identical and see whether LLM scores move accordingly.","tokens_in":10813,"feed_emoji":"🤖","tokens_out":4404,"duration_ms":47199,"temperature":0.7,"pith_summary":"This review argues that LLM-based quality scores are already technically superior to bibliometrics for research evaluation, because they correlate more strongly with expert judgements, cover more fields and recent years, and can in principle assess rigour, originality, and significance directly. It draws on a series of studies, including a large-scale analysis of UK REF2021 results, in which ChatGPT-4o scores (averaged over many runs, from titles and abstracts) correlated positively with departmental expert scores in all but one field. The paper cautions that the evidence is not conclusive, since the AI may have used publicly available departmental quality information, and that LLM biases, opacity, and vulnerability to gaming are unresolved. If the claim holds, LLM indicators could take over the supporting role that citation metrics now play, changing what researchers and journals are incentivised to optimise.","feed_headline":"LLM scores outdo citations as quality signal, argues review","feed_subtitle":"Early evidence says AI scores track expert ratings better than citations do, though bias and gaming risks remain.","key_machinery":"The load-bearing mechanism is the 'indicator' concept from evaluative bibliometrics: a quantity that associates with research quality without purporting to measure it. For LLMs, the specific mechanism is prompt-based scoring—supplying a quality definition, usually the REF2021 four-point scale, with an article's title and abstract, asking the model for a score, and then averaging scores across many repetitions to reduce random variation. This turns the LLM into a text-based proxy for expert review, in contrast to citation counts, which are influence-based proxies.","core_discovery":"The paper's central claim is that 'LLM-based quality evaluations seem to be already superior to bibliometrics as research quality indicators, although with clear and not yet well understood biases.' The evidence comes from experiments feeding ChatGPT the titles and abstracts of published articles along with the UK REF quality definitions (rigour, originality, significance; 1* to 4*), then averaging repeated scores. Averaged ChatGPT-4o scores correlate with expert quality scores better than citation-based indicators do in most fields, the review reports, and the approach works in arts and humanities where citations are nearly useless, while failing mainly in clinical medicine. The author stresses that neither LLM scores nor citations measure quality; they are indicators, and any responsible use must weigh the known tradeoffs.","pith_inferences":["A direct test of the leakage hypothesis is feasible using preprints or anonymised abstracts, and the author's own work has not yet run that test; this is the fastest way to decide whether the claimed superiority is real.","If LLM scores do track quality through abstract claims, then gaming countermeasures could include adversarial style normalisation or requiring LLMs to justify scores from specific sentences, which would be harder to fake than overall impressions.","The same scoring mechanism could be turned into an audit instrument for human peer review, giving a stable, reproducible baseline against which reviewer bias and noise could be measured.","Because LLMs can score any output type (books, proposals, datasets), the indicator role they may take over from bibliometrics is broader than citations, and so are the systemic effects on what kinds of work researchers choose to pursue."],"forward_implications":["National research evaluation exercises like the REF could begin offering LLM scores alongside citation data as supporting evidence for panels.","Quality assessment would extend to very recent papers and to arts, humanities, and social science fields where citation indicators are weak or useless.","Researchers would face a new incentive: writing abstracts that optimise LLM-assessed originality, rigour, and significance claims, potentially encouraging overselling.","Journal editors would have an incentive to permit exaggerated abstracts if LLM-based journal indicators replace impact factors, threatening the integrity of the public record.","Any rollout would need bias audits for age, field, abstract length, gender, institution, and methodology preferences before scores are used."],"supporting_citations":[{"why":"Large-scale study correlating ChatGPT-4o scores with departmental REF2021 scores across 34 Units of Assessment, the main evidence for superiority.","marker":"(Thelwall & Yaghi, 2024b)"},{"why":"Small-scale study showing ChatGPT REF scores correlate weakly with expert scores and that averaging 15 submissions raises the correlation.","marker":"(Thelwall, 2024)"},{"why":"Follow-up showing API-based title-and-abstract submission with 30 iterations achieves a 0.67 correlation with expert scores.","marker":"(Thelwall, 2025a)"},{"why":"Establishes citation-based indicators as only weakly or moderately associated with expert quality judgements, the baseline LLMs are compared against.","marker":"(Thelwall et al., 2023b)"},{"why":"Large-scale bias scan suggesting tendencies by article age, field, and abstract length, the main known bias evidence.","marker":"(Thelwall & Kurt, 2024)"},{"why":"Defines the four-point quality scale (1* to 4*) that all the ChatGPT scoring studies ask the model to use.","marker":"(REF, 2019)"},{"why":"The Leiden Manifesto position that indicators should support rather than replace peer review, which frames the paper's responsible-use conclusion.","marker":"(Hicks et al., 2015)"}],"fun_headline_variants":["LLM scores rival citation counts in research quality checks","AI quality scores beat citations, but bias and gaming loom","ChatGPT outdoes citations on research quality, review finds","LLM indicators surpass bibliometrics for quality, with caveats","AI quality ratings outpace citation metrics, gaming risks remain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire superiority claim rests on assuming that correlations between ChatGPT scores and departmental or expert REF scores reflect the AI's genuine ability to judge quality, rather than its exploitation of publicly available information about the universities' REF profiles or stylistic patterns in abstracts.","fun_headline_variants_meta":{"raw":{"variants":["LLM scores rival citation counts in research quality checks","AI quality scores beat citations, but bias and gaming loom","ChatGPT outdoes citations on research quality, review finds","LLM indicators surpass bibliometrics for quality, with caveats","AI quality ratings outpace citation metrics, gaming risks remain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000602,"raw_usage":{"total_tokens":2812,"prompt_tokens":948,"completion_tokens":1864,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":1783}},"tokens_in":564,"tokens_out":1864,"duration_ms":15291,"temperature":1.0,"reasoning_tokens":1783,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:25:37.546951+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment could settle it: take a set of articles with known expert quality scores, submit each to an LLM with its departmental affiliation and author names removed, and compare averaged scores to the expert scores. If the correlation collapses when identifying information is stripped, the public-information leakage explanation wins; if it holds, genuine text-based quality judgement is supported. A second check would modify abstracts (inflate or deflate quality claims) while keeping the underlying research identical and see whether LLM scores move accordingly.","supporting_citations":[],"review_version":1}