Pith. sign in

REVIEW 3 major objections 5 minor 3 references

This paper claims that after ChatGPT, research novelty in Information Systems journals declined relative to prior levels for authors at non-English-dominant institutions, by 0.18 standard deviations compared with English-dominant counterpar

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-02 17:34 UTC pith:7BM6PQZG

load-bearing objection A genuinely useful, clearly-written DiD study with a plausible but unvalidated novelty measure; the central result could be style convergence rather than a decline in intellectual novelty. the 3 major comments →

arxiv 2603.22510 v3 pith:7BM6PQZG submitted 2026-03-23 cs.DL cs.AIcs.IR

Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts

classification cs.DL cs.AIcs.IR
keywords research noveltylarge language modelsChatGPTdifference-in-differencessemantic embeddingsInformation Systems journalsEnglish-dominant affiliationacademic productivity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that the post-ChatGPT productivity surge in Information Systems research came with a measurable convergence of research ideas, and that the convergence was concentrated among authors at institutions in non-English-dominant countries. Using 13,847 articles from 44 journals, it measures novelty as the semantic distance between an article's title and abstract and its nearest recently published predecessors, then compares changes before and after ChatGPT's release. The headline result is a differential decline of 0.176 standard deviations for non-English-dominant-affiliation authors, about 7 percentile points, stable across several modeling choices and absent in a placebo test. The authors argue this is consistent with LLMs making conventional writing cheaper while making exploratory thinking feel comparatively expensive, though they do not observe actual LLM use.

Core claim

The paper's central claim is that after the public release of ChatGPT, articles whose first authors are affiliated with institutions in non-English-dominant countries lost ground in relative semantic novelty compared with articles from English-dominant affiliations. In the full model the differential decline is -0.176 standard deviations (p < 0.001), around 7 percentile points, and it survives several checks: alternative neighbor counts, a 2024 treatment break, dropping COVID years, and a placebo test at 2022. The authors are careful to say this is a heterogeneous post-2022 shift rather than a causal effect of LLM adoption, because individual LLM use is not observed and no untreated group ex

What carries the argument

The measuring instrument is a document-embedding model trained on citation graphs that turns titles and abstracts into vectors; novelty is the average distance to the ten nearest articles from the previous two years, standardized within each year. It does the work of translating 'how different is this idea from what was recently published' into a comparable score. The difference-in-differences setup compares the z-scored novelty of non-English-dominant-affiliation authors against English-dominant-affiliation authors before and after the ChatGPT release, with the interaction term capturing the differential shift.

Load-bearing premise

The load-bearing premise is that the embedding distance between an article and its nearest recent predecessors measures intellectual novelty rather than writing style or topic popularity — plus the parallel-trends assumption that the two affiliation groups would have moved in step absent ChatGPT, a premise the two-year pre-period cannot strongly test.

What would settle it

Re-run the same difference-in-differences on full-text embeddings or on a style-controlled representation that strips function words and sentence templates; if the Post × Non-EDA coefficient shrinks to near zero, the metric was capturing LLM-induced prose style, not idea novelty. Alternatively, a placebo exercise in a field without writing-assistance benefits should show no differential decline if the mechanism is writing cost.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the 0.176-SD differential decline is real, the post-2022 productivity gains documented elsewhere in academia are not neutral: the extra output is, on average, closer to the field's center of gravity.
  • For non-English-dominant-affiliation researchers, the relative position in the novelty distribution fell by about 7 percentile points, meaning the right tail of most-novel papers compressed most, not a uniform quality decline.
  • The persistent result under a 2024 treatment break, COVID exclusion, and alternate neighbor counts implies the shift is not an artifact of one modeling choice.
  • If the proposed mechanism is right, editors and funders can counteract the convergence only by rewarding novelty explicitly, since LLM-assisted writing makes conventional output cheaper.
  • The result shifts the debate about LLMs and science from publication counts to the semantic positioning of the published record.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • My inference: the semantic-distance metric cannot fully separate 'idea novelty' from 'prose novelty'; if LLMs make abstracts more formulaic, the estimated decline may partly be a style effect. A style-controlled embedding or full-text analysis would settle this.
  • My inference: if the mechanism is reduced writing cost, the same differential should appear within English-dominant countries when comparing non-native vs native English speakers, which is testable with name- or education-based proxies.
  • My inference: because novelty is z-scored within year, the paper does not show that overall novelty fell; it shows relative repositioning. An aggregate decline could still exist and would require different standardization.
  • My inference: the psychological-distance explanation predicts the effect should shrink or vanish when researchers use LLMs only for editing after idea generation; an experimental design assigning LLM use to ideation vs execution stages could test this.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper asks whether the post-2022 availability of ChatGPT is associated with a differential decline in the semantic novelty of Information Systems articles whose first authors are affiliated with institutions in non-English-dominant countries (non-EDA) relative to English-dominant countries (EDA). Using 13,847 Web of Science articles from 44 A*/A journals (2020–2025), the author measures novelty as the within-year z-scored cosine distance of each article's SPECTER2 title/abstract embedding to its k = 10 nearest neighbors in a rolling t-2/t-1 window. A difference-in-differences regression with a January 2023 break and HC1 robust errors yields a Post × Non-EDA coefficient of -0.176 (p < 0.001), equivalent to about 7 percentile points. The estimate is stable across k values, an alternative 2024 break, exclusion of COVID years, and a placebo break at 2022; the balanced panel of 1,079 repeat authors gives -0.142 (p = 0.104). The paper interprets the result through construal level theory while explicitly acknowledging that LLM adoption, psychological distance, and construal levels are not directly observed.

Significance. If the measurement assumption holds, this is one of the first large-scale bibliometric demonstrations that GenAI availability coincided with a measurable convergence in published research novelty, and it is novel in focusing on non-English-dominant institutional contexts. The paper is transparent in its causal framing: the non-EDA/EDA contrast is treated as differential exposure intensity rather than a clean treatment-control contrast, and the author reports a placebo test, multiple robustness checks, and a candid statement of limitations. These are real strengths. The contribution, however, rests entirely on the validity of SPECTER2 k-NN distance as a measure of intellectual novelty rather than of abstract wording, style, or topical conformity. Because the paper itself labels this limitation 'central, not peripheral' and provides no direct construct-validity evidence, the empirical claims are currently conditional on an unvalidated outcome measure.

major comments (3)
  1. [Measuring Semantic Novelty; Limitations] The outcome is a SPECTER2 title/abstract distance, and the paper states in Limitations that 'because the paper's entire contribution rests on this novelty measure, this limitation is central, not peripheral.' No analysis establishes that changes in this distance track intellectual novelty rather than abstract wording, framing, or LLM-typical style. Since LLMs are known to change abstract wording and structure, the non-EDA convergence could be a style convergence rather than a decline in research novelty. I ask for direct construct-validity tests: (a) compare the metric to human novelty ratings on a sample; (b) show robustness to style controls, such as function-word-only embeddings, punctuation/format features, or a paraphrasing baseline; (c) use an independent novelty measure based on citations or reference combinations (e.g., CD-index-like constructs); (d) run a falsification where pre
  2. [Identification Strategy / Event Study] Parallel trends rests on only two pre-periods (2020 and 2021), so the placebo break at 2022 has low power. The balanced-panel estimate is directionally consistent but not significant (beta = -0.142, p = 0.104). The deeper problem is that any post-2022 shock differentially affecting non-EDA publications—changes in submission composition, topic popularity, editorial practices, or country-specific LLM diffusion—could generate the same interaction. A pre-LLM placebo cannot rule out confounds that only appear after 2022. Please add: (i) a within-author event study with year-by-year post coefficients to see when divergence begins and whether it is sustained; (ii) controls for topic composition and author entry/exit; (iii) an outcome placebo using a stylistic similarity measure, showing that the DiD is specific to the proposed novelty construct.
  3. [Identification Strategy / Robustness (publication lag)] Post is assigned by publication year, so many 2023 papers may have been written or accepted before ChatGPT. The 2024-break specification addresses part of this, but it leaves only two post-break years and yields a weaker coefficient (-0.137). The paper argues that publication lag likely attenuates the estimate, but this is not guaranteed if the composition of non-EDA submissions changed over 2023–2025. Please report the interaction year by year (including 2023, 2024, 2025 separately) and, if available, use submission or acceptance dates. A short discussion of how differential timing of ChatGPT adoption across countries affects the contrast would also strengthen the interpretation.
minor comments (5)
  1. [Data] Please specify the exact SPECTER2 checkpoint and how titles and abstracts were concatenated or pooled before embedding. This information is essential for reproducibility of the novelty measure.
  2. [Table 2] The 'Controls' model (Model 2) is described as including controls, but the reported coefficients are almost all near zero. Clarify which controls are in each model and why the Post coefficient is exactly 0.000 in Model 1. A note on units and scaling would help.
  3. [Robustness Checks] The COVID-drop estimate has p = 0.015. With multiple robustness tests, a brief discussion of multiple comparisons or a family-wise error perspective would be useful, though the k and break-date results are stable.
  4. [Figures 3 and 4] The captions for Figures 3 and 4 do not state whether error bars or confidence intervals are shown. If not shown, please add them or state that the figures show point estimates only.
  5. [Conclusion] Because novelty is z-scored within each year, the design identifies only relative novelty shifts; absolute aggregate declines are removed by construction. The conclusion should restate this more prominently to avoid readers inferring a global novelty decline.

Circularity Check

0 steps flagged

No significant circularity; the main result is an estimated interaction on an independently constructed outcome, and flagged limitations are construct-validity concerns rather than derivation-circularity.

full rationale

The derivation chain is not circular. Novelty is measured by SPECTER2 k-NN cosine distance z-scored within year; this is an independently constructed dependent variable, not a parameter fitted to the post/NonEDA interaction. The coefficient of interest, beta_3 = -0.176 (p < 0.001), is estimated from data; no fitted value is renamed as a prediction, and the placebo test at 2022 (beta = -0.021, p = 0.709) provides a pre-treatment check. The z-scoring does force the aggregate Post coefficient in the no-control model to zero, and the paper explicitly acknowledges this ('confirming that z-scored novelty shows no between-year trend by construction due to within-year standardization'); but the paper's contribution is the interaction term, which is not fixed by construction. The EDA/non-EDA classification and the November 2022 treatment break are external to the outcome. The Limitations section's admission that the novelty measure is 'central, not peripheral' is a construct-validity caveat about whether SPECTER2 distance captures intellectual novelty versus writing style, not a logical reduction of the result to its inputs. There are no load-bearing self-citations: the cited Kwon and Yang classification and productivity results are from different authors and are not used to define the outcome or the interaction. Therefore no circular step meets the quoted-reduction standard required for a positive finding.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

No new physical or theoretical entities are introduced. The result depends on a bundle of measurement and identification assumptions, all acknowledged in the paper, plus researcher-chosen hyperparameters that are only partially robustness-tested.

free parameters (4)
  • k (number of nearest neighbors) = 10 (robustness: 5, 15, 20)
    Chosen to define semantic novelty; central estimate is stable across values, but the metric itself changes with k.
  • Rolling prior window = t-2 to t-1
    Chosen without robustness checks; a longer window would change the reference pool and could alter the novelty scores.
  • English-dominant country set = USA, UK, Canada, Australia, NZ, Ireland
    Classification of EDA vs non-EDA affiliations follows Kwon and Yang; no alternative country sets are tested.
  • Treatment break = January 2023 (robustness: January 2024)
    Approximation of ChatGPT availability; the 2024 break partially addresses publication lag, but the pre-period then includes post-ChatGPT 2023.
axioms (4)
  • domain assumption SPECTER2 embeddings capture conceptual content, not writing style.
    Stated in Step 1; central to the validity of the novelty measure.
  • domain assumption Cosine distance to 10 nearest recent predecessors is a valid operationalization of research novelty.
    Used in Step 2; relies on prior novelty-detection literature but is not independently validated.
  • domain assumption Parallel pre-trends between EDA and non-EDA groups would hold absent LLM availability.
    Identifying assumption for DiD; only 2020 and 2021 are available for testing.
  • domain assumption Affiliation country is a valid proxy for language environment and LLM writing-assistance reliance.
    Classification used for NonEDA; admitted as a proxy in the paper.

reviewed 2026-08-02 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts." pith.science (2026). https://pith.science/paper/7BM6PQZG

@misc{pith2026260322510,
  author       = {Pith},
  title        = {Pith review of: Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7BM6PQZG}},
  note         = {Machine review of arXiv:2603.22510}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by changes in research novelty. We examine how relative abstract-level semantic novelty in Information Systems journals changed after ChatGPT became widely available and whether this change differed across institutional language contexts. We analyze 13,847 articles published from 2020 to 2025 in 44 A* and A Information Systems journals. Using SPECTER2 representations of titles and abstracts, we measure each article's semantic distance from its nearest recent predecessors and estimate a comparative pre/post model. Articles whose first authors were affiliated with institutions in non-English-dominant countries show a 0.176 standard deviation larger post-2022 decline in relative semantic novelty than articles from English-dominant affiliations, equivalent to about 7 percentile points. The pattern is similar across several alternative specifications, although the balanced-author estimate is less precise. We interpret this finding through a tension in generative AI-supported knowledge work. GenAI can widen access to prior knowledge and support new combinations, but it can also make established frames easier to reproduce. Because individual LLM use is not observed, the result identifies a heterogeneous post-2022 shift rather than an effect of LLM adoption. The study extends research on LLMs and scholarly productivity by shifting attention from publication counts to the semantic positioning of published articles and by showing that post-2022 change differs across institutional contexts.

Figures

Figures reproduced from arXiv: 2603.22510 by Ali Safari, Sahar Babaei.

Figure 1
Figure 1. Figure 1: Research model. The main path captures the association between LLM availability and semantic novelty, [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Event study coefficients for Post x Non-EDA interaction, relative to 2022. Pre-treatment coefficients (2020, 2021) are statistically insignificant. Two pre-periods provide limited but supportive evidence for the parallel trends assumption. Main Results [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Mean z-scored novelty by affiliation group and period. Non-EDA authors show higher baseline novelty pre-LLM, converging toward the EDA mean post-LLM [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Yearly mean z-scored novelty by author group. The non-EDA trajectory diverges downward after 2022, while junior-senior trajectories show no differential shift. The visual pattern is consistent with a heterogeneous post￾2022 change concentrated among non-EDA authors [PITH_FULL_IMAGE:figures/full_fig_p012_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Coefficient plot for the full model (Model 5) with 95% confidence intervals. Post x Non [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Forest plot of Post x Non-EDA coefficients across all specifications. The main effect is stable across k values and sub-samples. The placebo test crosses zero, confirming no pre-trend. The consistency of coefficients across specifications (-0.132 to -0.178) suggests the finding is not sensitive to particular modeling choices [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Distribution of z-scored novelty by affiliation group, pre vs. post LLM. EDA authors show minimal distributional shift. Non-EDA authors show a visible leftward shift (lower novelty) in the post-LLM period, with the right tail (highest-novelty papers) compressing toward the center. This distributional compression is consistent with the homogenization pattern documented by Doshi and Hauser (2024) in creative… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

3 extracted references · 1 linked inside Pith

  1. [1]

    Automation and New Tasks: How Technology Displaces and Reinstates Labor

    Acemoglu, D. and Restrepo, P. (2019). "Automation and New Tasks: How Technology Displaces and Reinstates Labor." Journal of Economic Perspectives, 33(2), 3-30. Autor, D.H., Levy, F. and Murnane, R.J. (2003). "The Skill Content of Recent Technological Change: An Empirical Exploration." Quarterly Journal of Economics, 118(4), 1279-1333. Benbya, H., Davenpor...

  2. [2023]

    Construal-Level Theory of Psychological Distance

    Trope, Y. and Liberman, N. (2010). "Construal-Level Theory of Psychological Distance." Psychological Review, 117(2), 440-463. Uzzi, B., Mukherjee, S., Stringer, M. and Jones, B. (2013). "Atypical Combinations and Scientific Impact." Science, 342(6157), 468-472. Wang, J., Veugelers, R. and Stephan, P. (2017). "Bias Against Novelty in Science: A Cautionary ...

  3. [2025]

    SciRepEval: A Multi-Format Benchmark for Scientific Document Representations

    Singh, A., D'Arcy, M., Cohan, A., Downey, D. and Feldman, S. (2023). "SciRepEval: A Multi-Format Benchmark for Scientific Document Representations." Proceedings of EMNLP

This paper was first reviewed by deepseek-v4-flash on August 2, 2026.