Pith. sign in

REVIEW 4 major objections 6 minor 213 references

Experimental Evidence on the Learning Impact of Generative AI

T0 review · 4 major / 6 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read Off-the-shelf generative AI raises students' unaided test scores by 0.27 SD immediately, and most of that gain still shows up one week later without AI.

desk verdict Clean lab RCT: off-the-shelf AI raises unaided test scores ~0.27 SD immediately and one week later, with delayed essay gains that stick for augmentation users; external validity is the real limit, not internal identification. read the letter →

arxiv 2607.08849 v1 pith:KZD7FDJ4 submitted 2026-07-09 econ.GN cs.HCq-fin.EC

classification econ.GNcs.HCq-fin.EC
keywords generativeAIstudentlearningrandomizedexperimentknowledgeretentionaugmentationversusautomationessayqualityhighereducationhumancapital
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether giving college students free access to ordinary generative AI helps or hurts real learning, not just task output. In a randomized, proctored lab experiment, undergraduates spent up to 35 minutes learning an unfamiliar technical topic and writing an analytical essay, either with AI allowed or forbidden; they then took unaided knowledge tests and wrote unaided essays immediately and about a week later. AI access raised immediate test scores by 0.27 standard deviations, and roughly three-quarters of that gain remained one week later when no one had AI. Essay quality barely moved while AI was available, but unaided essays a week later improved in style and relevance to the prompt. Those delayed writing gains were concentrated among students who used AI as a tutor (augmentation) rather than as a ghostwriter (automation); automation users' short-run quality boosts vanished once AI was removed. Two mechanisms show up in the data: treated students reallocated time away from drafting and toward reading and searching, and they reported more enjoyment. The design speaks to the chatbots students actually use, not customized tutoring systems, and holds total learning time roughly fixed.

What carries the argument

The augmentation-versus-automation classification of ChatGPT conversation logs, which separates students who use AI as a tutor (explain, clarify, feedback) from those who use it to produce draft text. That split organizes both short-run versus retained effects and students' own mental models of how AI affects learning.

What would settle it

A field experiment in ordinary coursework where students choose total study time and AI access either raises or lowers total learning once time reallocation is allowed, or where one-week retention gains disappear when topics are already familiar and stakes are real grades.

Watch

Extended reading notes

Core claim

Random assignment to off-the-shelf generative AI during a short learning-and-essay phase raises unaided knowledge-test scores by 0.27 SD immediately and by a similar amount about one week later without AI. Higher-order essay quality improves mainly after AI is removed, and those delayed gains are larger for students who use AI to explain concepts than for those who use it to generate text.

Load-bearing premise

That a fixed-time lab session with elite undergraduates on low-prior-knowledge topics, followed by one-week retention, identifies the learning effect students would get when they choose how long to study and often use AI to finish faster.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports a proctored, in-person RCT with 211 Middlebury undergraduates who learn an unfamiliar technical topic and write an analytical essay under AI-allowed or AI-forbidden conditions, then take unaided knowledge tests and write unaided essays immediately and about one week later. Random assignment raises ChatGPT use by ~67 pp. The main finding is that AI access raises unaided test scores by 0.27 SD immediately (ITT 6.7 pp on a 56.3% control mean) and by a similar 0.27 SD one week later without AI (~76% of the immediate effect). Session One essays show more AI-detected text and little quality change; Session Two unaided essays improve mainly in writing style/clarity and relevance, with larger delayed quality gains among LLM-classified “augmentation” users than “automation” users. Mechanisms include a shift of time from drafting toward reading/searching and higher reported enjoyment. The paper also reports beliefs and open-ended causal narratives about AI and learning.

Significance. If the estimates hold, this is among the cleanest experimental answers to whether off-the-shelf generative AI builds durable academic human capital rather than only short-run task performance. Strengths include: random assignment with a large first stage; multi-method compliance monitoring; unaided immediate and one-week retention assessments; dual human and AI essay grading plus objective linguistic and AI-detection measures; and a transparent literature meta-comparison. The augmentation-versus-automation heterogeneity and the time-reallocation/enjoyment mechanisms are policy-relevant and map onto students’ own mental models. The design deliberately studies unrestricted chatbots against a no-AI counterfactual, which is more informative for real student use than many scaffolded-tutor designs. External validity to ordinary coursework with endogenous study time remains the main limit, which the authors already flag.

major comments (4)
  1. Table 7, Panel D and §4.4: AI access raises any integrity violation by 12.6 pp (p=0.005). The back-of-envelope that cheating can explain ~2.2 pp of the 6.7 pp Session One test ITT (~one-third) is load-bearing for interpreting the knowledge gains as learning rather than test-taking contamination. Please report the main test-score ITT/TOT for Sessions One and Two after excluding proctor-flagged and self-reported violators (and a joint “any violation” sample), and clarify whether Session Two retention survives that restriction. If the retention effect is robust, state that prominently; if not, revise the learning interpretation accordingly.
  2. Table 8 and §5.3: The automation/augmentation split is constructed from LLM labels of treated students’ ChatGPT logs and is endogenous among users (different prompting, time use, and AI-detection rates). The differential fade-out is informative as descriptive heterogeneity, but several passages read as if use mode is a causal treatment. Please reframe Table 8 as non-causal heterogeneity among treated users, report balance of baseline ability/AI experience across use types, and avoid language that implies random assignment of automation vs augmentation.
  3. Table 6, columns 4–6 and Abstract: Session Two overall quality is +0.31 points (0.20 SD, p=0.143); only writing style and relevance are significant. The abstract’s claim that essay quality “improves in style and relevance” is accurate, but the introduction and §5.2 sometimes elevate this to broader “higher-order skills.” Please align the main text with the dimension-level pattern, report multiple-testing-adjusted inference for the five dimensions (or pre-specify primary essay outcomes), and avoid treating the imprecise overall quality index as established.
  4. Conclusion and §2: The design holds total learning time roughly fixed (~33 minutes). The paper correctly notes that ordinary AI use often saves time. Because the central policy claim is about learning impact of AI access, please add a short quantitative discussion of how large a reduction in time-on-task would be needed to offset the 0.27 SD retention gain under alternative assumptions, so readers can map the lab ITT to settings with endogenous study time.
minor comments (6)
  1. Figure 5 / Appendix B.7: State more clearly which of the 22 literature estimates are ITT vs TOT and whether all use unassisted outcomes only, so the grand mean of 0.18 SD is interpretable.
  2. Table 4: Self-assessed knowledge is essentially flat while objective scores rise. A brief discussion of why subjective knowledge does not track the test gains would help (calibration, ceiling of the 0–10 scale, or different construct).
  3. Appendix Table A1 and take-up: White students are less likely to use AI among the treated. Given the first-stage is not universal, a short note on whether TOT is driven by particular subgroups would be useful.
  4. §3.3 / double-lasso: Report the selected controls for the main test-score specifications (or an appendix table) so readers can see what residual imbalance is being adjusted.
  5. Figure 8 and Appendix C: The narrative coding is interesting but long relative to the experimental contribution; consider moving more of the causal-graph material to the appendix and keeping one summary figure in the main text.
  6. Typos/clarity: “whereasautomation” missing space in the abstract; ensure consistent Session One/Two capitalization; check that N varies slightly across tables (essay missingness) are explained once in a note.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: central claims are experimental ITTs from random assignment, not quantities forced by fitted parameters or self-definition.

full rationale

The paper’s load-bearing results are intent-to-treat (and 2SLS TOT) estimates of random assignment to off-the-shelf AI access on unaided knowledge tests and essays (Eq. 1; Tables 4–6, 8). The 0.27 SD immediate and retention effects are differences in measured outcomes between AI-allowed and AI-forbidden arms; they are not derived from a structural parameter fitted to the same outcomes, nor defined in terms of those outcomes. Augmentation vs. automation is an LLM classification of ChatGPT conversation logs used only for heterogeneity (Appendix B.6; Table 8), not an input that forces the main ITT. Mechanisms (time mix, enjoyment, integrity) and belief/narrative analyses are separately measured descriptive outcomes. Self-citations (Contractor and Reyes 2026 on campus adoption/usage) supply context only and do not underwrite identification. There is no self-definitional loop, fitted-input-as-prediction, uniqueness import, or renaming of a known result as a first-principles derivation. The design is self-contained against its own experimental benchmarks.

Assumptions & free parameters 3 free parameters · 5 assumptions · 1 invented entities

This is an empirical RCT, not a theory paper. Load-bearing premises are design and measurement choices rather than free physical constants or invented particles. The main claim rests on random assignment identifying the ITT of AI access, on tests/essays measuring learning, and on LLM labels capturing meaningful use types for heterogeneity.

free parameters (3)
  • Double-lasso control selection and strata fixed effects
    Precision and residual imbalance adjustment depend on the covariate pool and lasso selection; not a fitted structural constant, but a modeling choice that can move point estimates slightly.
  • LLM conversation classification thresholds/prompts for augmentation vs automation
    Heterogeneity results depend on Claude Opus labels of chat logs into Automation/Augmentation/Mixed/Other; different prompts or models could reassign users.
  • Essay quality aggregation (human average + AI grader average)
    Main essay outcomes average human Prolific graders and Claude scores on a 0–10 rubric; weights and grader selection affect quality estimates.
assumptions (5)
  • domain assumption Random assignment to AI-allowed vs AI-forbidden identifies the causal effect of AI access (ITT) under SUTVA and no differential attrition.
    Standard RCT identification; balance and attrition checks support it (§3.2–3.3).
  • domain assumption Unaided multiple-choice tests and analytical essays measure factual/conceptual knowledge and higher-order skills relevant to learning.
    Outcome construct validity is assumed throughout §§3.4–5.
  • domain assumption One week without resources is a meaningful retention horizon for skill accumulation claims.
    Session Two is the persistence test; longer horizons are not observed (§2.2, §5).
  • ad hoc to paper LLM labels of ChatGPT logs into augmentation vs automation recover economically meaningful use modes.
    Used for Table 8 heterogeneity; validated with prompt patterns, time use, and AI-detection rates (§5.3, App. B.6).
  • standard math Linear models with heteroskedasticity-robust SEs and double-lasso controls recover average treatment effects of interest.
    Equation (1) and 2SLS TOT (§3.3).
invented entities (1)
  • Augmentation vs automation user types (student-level) independent evidence
    purpose: Partition treated AI users to explain differential persistence of learning and essay quality.
    Operational taxonomy built from conversation logs; related to prior automation/augmentation language but newly applied as experimental moderators here.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Experimental Evidence on the Learning Impact of Generative AI." pith.science (2026). https://pith.science/paper/KZD7FDJ4

@misc{pith2026260708849,
  author       = {Pith},
  title        = {Pith review of: Experimental Evidence on the Learning Impact of Generative AI},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KZD7FDJ4}},
  note         = {Machine review of arXiv:2607.08849}
}
read the original abstract

We study how generative AI affects student learning in a randomized experiment. In proctored, in-person sessions, undergraduates learn about an unfamiliar topic and write an analytical essay with or without access to off-the-shelf generative AI, then complete unaided assessments immediately and one week later. We measure learning with knowledge tests (factual and conceptual understanding) and open-ended essays (higher-order skills). AI access raises immediate test scores by 0.27 standard deviations. These gains persist one week later. Essay quality, by contrast, changes little while students have AI access but improves in style and relevance one week later, when students write unaided. These delayed gains are larger among augmentation users-who use AI to explain concepts rather than generate text-whereas automation users' short-run quality gains vanish once AI is removed. We find evidence for two mechanisms behind the learning gains: students shift time away from drafting text and toward reading and searching for information, and they report greater learning enjoyment.

Figures

Figures reproduced from arXiv: 2607.08849 by the authors.

Figure 1
Figure 1. Experimental Sessions Timelines Panel A. Session One timeline 0 min 60 min Welcome & Instructions Baseline Test Learning Phase (35 mins max) Post-Learning Survey Post-Learning Test Panel B. Session Two timeline 0 min 45 min Welcome & Instructions Endline Test (10 Questions) Essay (20 mins max) Randomized order Exit Survey Notes: This figure shows the timeline of the two experimental sessions. Panel A shows the struc… view at source ↗
Figure 2
Figure 2. The Impact of AI Access on Generative AI Usage During the Learning Phase [PITH_FULL_IMAGE:figures/full_fig_p028_2.png] view at source ↗
Figure 3
Figure 3. Types of Generative AI Use During the Learning Phase [PITH_FULL_IMAGE:figures/full_fig_p029_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Distribution of Test Performance by Treatment Group [PITH_FULL_IMAGE:figures/full_fig_p030_4.png]
Figure 5
Figure 5. Figure 5: Effect Sizes Across AI-and-Learning Experiments [PITH_FULL_IMAGE:figures/full_fig_p031_5.png]
Figure 6
Figure 6. Figure 6: Effects of AI Access on Essay Quality and Linguistic Features [PITH_FULL_IMAGE:figures/full_fig_p032_6.png]
Figure 7
Figure 7. Figure 7: Actual and Perceived Treatment Effects on Test Performance [PITH_FULL_IMAGE:figures/full_fig_p033_7.png]
Figure 8
Figure 8. Figure 8: The Average Narrative About AI’s Effect on Learning [PITH_FULL_IMAGE:figures/full_fig_p034_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

213 extracted references · 2 canonical work pages

  1. [1]

    Review of Economic Studies , year =

    Andre, Peter and Haaland, Ingar and Roth, Christopher and Wiederholt, Mirko and Wohlfart, Johannes , title =. Review of Economic Studies , year =

  2. [2]

    2024 , institution =

    Emi, Bradley and Spero, Max , title =. 2024 , institution =. 2402.14873 , archiveprefix =

  3. [3]

    2025 , month =

    Masrour, Elyas , title =. 2025 , month =

  4. [4]

    and Spero, Max , title =

    Masrour, Elyas and Emi, Bradley N. and Spero, Max , title =. Proceedings of the 1st Workshop on Detecting AI Generated Content (GenAIDetect), COLING , pages =. 2025 , url =

  5. [5]

    International Conference on Learning Representations (ICLR) , year =

    Thai, Katherine and Emi, Bradley and Masrour, Elyas and Iyyer, Mohit , title =. International Conference on Learning Representations (ICLR) , year =

  6. [6]

    2025 , url =

    Jabarian, Brian and Imas, Alex , title =. 2025 , url =

  7. [7]

    Journal of Economic Perspectives , volume=

    Automation and New Tasks: How Technology Displaces and Reinstates Labor , author=. Journal of Economic Perspectives , volume=

  8. [8]

    Science , volume=

    What Can Machine Learning Do? Workforce Implications , author=. Science , volume=

Show all 213 references
  1. [9]

    2025 , month =

    Becker, Joel and Rush, Nate and Barnes, Beth and Rein, David , title =. 2025 , month =. 2507.09089 , archiveprefix =

  2. [10]

    2025 , month =

    Building an. 2025 , month =

  3. [11]

    and Hitzig, Zo\"e and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =

    Chatterji, Aaron and Cunningham, Thomas and Deming, David J. and Hitzig, Zo\"e and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , title =. 2025 , type =

  4. [12]

    The Generative

    Str. The Generative. 2026 , url =

  5. [13]

    VoxDevLit , volume =

    Education Technology , author =. VoxDevLit , volume =

  6. [14]

    2025 , month = aug, howpublished =

    Narayanan, Arvind , title =. 2025 , month = aug, howpublished =

  7. [15]

    2025 , doi =

    Kestin, Greg and Miller, Kelly and Klales, Anna and Milbourne, Timothy and Ponti, Gregorio , journal =. 2025 , doi =

  8. [16]

    The Impact of Generative

    Lee, Hao-Ping (Hank) and Sarkar, Advait and Tankelevitch, Lev and Drosos, Ian and Rintel, Sean and Banks, Richard and Wilson, Nicholas , booktitle =. The Impact of Generative. 2025 , publisher =

  9. [17]

    Reading Between the Lines: Modeling User Behavior and Costs in

    Mozannar, Hussein and Bansal, Gagan and Fourney, Adam and Horvitz, Eric , booktitle =. Reading Between the Lines: Modeling User Behavior and Costs in. 2024 , publisher =

  10. [18]

    2024 , eprint =

    Empirical evidence of large language model's influence on human spoken communication , author =. 2024 , eprint =

  11. [19]

    2025 , eprint =

    Ammari, Tawfiq and Chen, Meilun and Zaman, S M Mehedi and Garimella, Kiran , title =. 2025 , eprint =

  12. [20]

    The New Yorker , year =

    Hsu, Hua , title =. The New Yorker , year =

  13. [21]

    , title =

    Walsh, James D. , title =. New York Magazine , year =

  14. [22]

    2025 , url =

    Kunal Handa and Drew Bent and Alex Tamkin and Miles McCain and Esin Durmus and Michael Stern and Mike Schiraldi and Saffron Huang and Stuart Ritchie and Steven Syverud and Kamya Jagadish and Margaret Vo and Matt Bell and Deep Ganguli , title =. 2025 , url =

  15. [23]

    2026 , url =

    Maxim Massenkoff and Eva Lyubich and Peter McCrory and Ruth Appel and Ryan Heller , title =. 2026 , url =

  16. [24]

    The Quarterly Journal of Economics , volume=

    The (perceived) returns to education and the demand for schooling , author=. The Quarterly Journal of Economics , volume=. 2010 , publisher=

  17. [25]

    The Review of Economic Studies , volume=

    Determinants of college major choice: Identification using an information experiment , author=. The Review of Economic Studies , volume=. 2015 , publisher=

  18. [26]

    Generative

    Contractor, Zara and Reyes, Germ. Generative

  19. [27]

    Handbook of the Economics of Education , editor =

    Technology and Education: Computers, Software, and the Internet , author =. Handbook of the Economics of Education , editor =

  20. [28]

    Journal of Economic Literature , volume=

    Upgrading education with technology: Insights from experimental research , author=. Journal of Economic Literature , volume=. 2020 , publisher=

  21. [29]

    Journal of Memory and Language , volume=

    How many words do we read per minute? A review and meta-analysis of reading rate , author=. Journal of Memory and Language , volume=. 2019 , publisher=

  22. [30]

    Journal of Reading , volume=

    Reading rate: Theory, research, and practical implications , author=. Journal of Reading , volume=. 1992 , publisher=

  23. [31]

    Handbook 1: Cognitive domain , author=

    Taxonomy of educational objectives: The classification of educational goals. Handbook 1: Cognitive domain , author=. 1956 , publisher=

  24. [32]

    2023 , month =

    Nam, Jane , title =. 2023 , month =

  25. [33]

    and Zhang, Hao and Gonzalez, Joseph E

    Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , title =. Advances in Neural Information Processing ...

  26. [34]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages =

    Chiang, Cheng-Han and Lee, Hung-yi , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages =. 2023 , publisher =

  27. [35]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =

    Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages =. 2023 , publisher =

  28. [36]

    , title =

    Carter, Susan Payne and Greenberg, Kyle and Walker, Michael S. , title =. Economics of Education Review , year =

  29. [37]

    The Quarterly Journal of Economics , year =

    Malamud, Ofer and Pop-Eleches, Cristian , title =. The Quarterly Journal of Economics , year =

  30. [38]

    Technology and Child Development: Evidence from the

    Cristia, Julian and Ibarrar. Technology and Child Development: Evidence from the. American Economic Journal: Applied Economics , year =

  31. [39]

    and Rush, Mark and Yin, Lu , title =

    Figlio, David N. and Rush, Mark and Yin, Lu , title =. Journal of Labor Economics , year =

  32. [40]

    and Fox, Lindsay and Loeb, Susanna and Taylor, Eric S

    Bettinger, Eric P. and Fox, Lindsay and Loeb, Susanna and Taylor, Eric S. , title =. American Economic Review , year =

  33. [41]

    and Ladd, Helen F

    Vigdor, Jacob L. and Ladd, Helen F. and Martinez, Erika , title =. Economic Inquiry , year =

  34. [42]

    and Goodman, Sarena and Smith, Jonathan , title =

    Dettling, Lisa J. and Goodman, Sarena and Smith, Jonathan , title =. The Review of Economics and Statistics , year =

  35. [43]

    , title =

    Caldwell, Jane E. , title =. CBE---Life Sciences Education , year =

  36. [44]

    Education and Information Technologies , year =

    Lewin, Cathy and Somekh, Bridget and Steadman, Stephen , title =. Education and Information Technologies , year =

  37. [45]

    Journal of the European Economic Association , volume =

    Expertise , author =. Journal of the European Economic Association , volume =

  38. [46]

    Generative

    Bastani, Hamsa and Bastani, Osbert and Sungu, Alp and Ge, Haosen and Kabakc. Generative. Proceedings of the National Academy of Sciences , volume =. doi:10.1073/pnas.2422633122 , note =

  39. [47]

    Belloni, Alexandre and Chernozhukov, Victor and Hansen, Christian , year = 2014, month = may, journal =. High-

  40. [48]

    , year = 2026, journal =

    Bick, Alexander and Blandin, Adam and Deming, David J. , year = 2026, journal =. The

  41. [49]

    Generative

    Brynjolfsson, Erik and Li, Danielle and Raymond, Lindsey , year = 2025, month = may, journal =. Generative

  42. [50]

    Endoscopist

    Budzy. Endoscopist. The Lancet Gastroenterology & Hepatology , volume =

  43. [51]

    Cui, Zheyuan (Kevin) and Demirer, Mert and Jaffe, Sonia and Musolff, Leon and Peng, Sida and Salz, Tobias , year = 2026, journal =. The

  44. [52]

    Writing Code vs

    Demirer, Mert and Musolff, Leon and Yang, Liyuan , institution =. Writing Code vs. Shipping Code: Productivity Effects Across Generations of

  45. [53]

    Dell'Acqua, Fabrizio and McFowland, Edward and Mollick, Ethan R. and. Navigating the. Organization Science , doi =

  46. [54]

    and Hauser, Oliver P

    Doshi, Anil R. and Hauser, Oliver P. , journal =. Generative

  47. [55]

    Meincke, Lennart and Nave, Gideon and Terwiesch, Christian , journal =

  48. [56]

    and Kushlev, Kostadin , journal =

    Moon, Kibum and Green, Adam E. and Kushlev, Kostadin , journal =. Homogenizing Effect of Large Language Models (

  49. [57]

    The Creative Link Between Words and Ideas Is Weakening in the

    Moon, Kibum and Kushlev, Kostadin and Bank, Andrew and. The Creative Link Between Words and Ideas Is Weakening in the. doi:10.31234/osf.io/jsz58_v6 , url =

  50. [58]

    Minnesota Law Review , volume =

    Lawyering in the Age of Artificial Intelligence , author =. Minnesota Law Review , volume =

  51. [59]

    Educational Evaluation and Policy Analysis , volume =

    How Big Are Effect Sizes in International Education Studies? , author =. Educational Evaluation and Policy Analysis , volume =

  52. [60]

    Educational Researcher , volume =

    Interpreting Effect Sizes of Education Interventions , author =. Educational Researcher , volume =

  53. [61]

    and Lavy, Victor , journal =

    Angrist, Joshua D. and Lavy, Victor , journal =. Using

  54. [62]

    Kirabo and Mackevicius, Claire L

    Jackson, C. Kirabo and Mackevicius, Claire L. , journal =. What Impacts Can We Expect from School Spending Policy?

  55. [63]

    The Promise of Tutoring for

    Nickow, Andre and Oreopoulos, Philip and Quan, Vincent , journal =. The Promise of Tutoring for

  56. [64]

    and Rockoff, Jonah E

    Chetty, Raj and Friedman, John N. and Rockoff, Jonah E. , journal =. Measuring the Impacts of Teachers

  57. [65]

    Economics of Education Review , volume =

    The Economic Value of Higher Teacher Quality , author =. Economics of Education Review , volume =

  58. [66]

    Baird, Matthew and Carpanelli, Mar and Xu, Brian and Xu, Kevin , journal =. Firms'

  59. [67]

    Academy of Management Journal , volume =

    When and How Artificial Intelligence Augments Employee Creativity , author =. Academy of Management Journal , volume =

  60. [68]

    , institution =

    Dell'Acqua, Fabrizio and Ayoubi, Charles and Lifshitz, Hila and Sadun, Raffaella and Mollick, Ethan and Mollick, Lilach and Han, Yi and Goldman, Jeff and Nair, Hari and Taub, Stewart and Lakhani, Karim R. , institution =. The Cybernetic Teammate: A Field Experiment on Generative

  61. [69]

    Does Generative

    Cruces, Guillermo and. Does Generative

  62. [70]

    and Sting, Fabian J

    Lehmann, Matthias and Cornelius, Philipp B. and Sting, Fabian J. , year = 2025, month = mar, number =

  63. [71]

    Experimental

    Noy, Shakked and Zhang, Whitney , year = 2023, month = jul, journal =. Experimental

  64. [72]

    Peng, Sida and Kalliamvakou, Eirini and Cihon, Peter and Demirer, Mert , year = 2023, month = feb, number =. The. 2302.06590 , doi =

  65. [73]

    Rav. Higher. 2025 , month = feb, journal =

  66. [74]

    Perceptions and

    St. Perceptions and. Computers and Education: Artificial Intelligence , volume =

  67. [75]

    Harvard undergraduate survey on generative

    Hirabayashi, Shikoh and Jain, Rishab and Jurkovi. Harvard undergraduate survey on generative

  68. [76]

    , journal =

    Goldsmith-Pinkham, Paul and Tan, Chenhao and Zentefis, Alexander K. , journal =. Human-

  69. [77]

    Kanazawa, Kyogo and Kawaguchi, Daiji and Shigeoka, Hitoshi and Watanabe, Yasutora , journal =

  70. [78]

    Psychological Methods , volume =

    A General Approach to Causal Mediation Analysis , author =. Psychological Methods , volume =

  71. [79]

    Social Sciences & Humanities Open , volume =

    Barcaui, Andr. Social Sciences & Humanities Open , volume =

  72. [80]

    2412.16429 , archivePrefix =

  73. [81]

    and Sellen, Abigail and Rintel, Sean and Goldstein, Daniel G

    Kreijkes, Pia and Kewenig, Viktor and Kuvalja, Martina and Lee, Mina and Vitello, Sylvia and Hofman, Jake M. and Sellen, Abigail and Rintel, Sean and Goldstein, Daniel G. and Rothschild, David M. and Tankelevitch, Lev and Oates, Tim , year = 2026, journal =. Effects of

  74. [82]

    Effective Personalized

    Chung, Angel Tsai-Hsuan and Zhang, Botong and Kung, Ling-Chieh and Bastani, Hamsa and Bastani, Osbert , year = 2026, institution =. Effective Personalized

  75. [83]

    Sentence-

    Reimers, Nils and Gurevych, Iryna , booktitle =. Sentence-. 2019 , publisher =

  76. [84]

    Song, Kaitao and Tan, Xu and Qin, Tao and Lu, Jianfeng and Liu, Tie-Yan , booktitle =

  77. [85]

    Proceedings of the International AAAI Conference on Weblogs and Social Media (ICWSM) , volume=

    Vader: A parsimonious rule-based model for sentiment analysis of social media text , author=. Proceedings of the International AAAI Conference on Weblogs and Social Media (ICWSM) , volume=

  78. [86]

    From Chalkboards to Chatbots: Evaluating the Impact of Generative

    De Simone, Mart. From Chalkboards to Chatbots: Evaluating the Impact of Generative. 2025 , month = may, institution =

  79. [87]

    and Ribeiro, Ana T

    Wang, Rose E. and Ribeiro, Ana T. and Robinson, Carly D. and Loeb, Susanna and Demszky, Dorottya , title =. 2024 , institution =. 2410.03017 , archivePrefix =

  80. [88]

    2025 , type =

    Kim, Donggwan and Mitrofanov, Dmitry and Wen, Qingsong and Xu, Tianlong , title =. 2025 , type =

  81. [89]

    British Journal of Educational Technology , year =

    Xu, Xiaoqing and Qiao, Lifang and Cheng, Nuo and Liu, Hongxia and Zhao, Wei , title =. British Journal of Educational Technology , year =

  82. [90]

    Controlled Clinical Trials , year =

    DerSimonian, Rebecca and Laird, Nan , title =. Controlled Clinical Trials , year =

  83. [91]

    and Ungar, Lyle and Duckworth, Angela L

    Lira, Benjamin and Rogers, Todd and Goldstein, Daniel G. and Ungar, Lyle and Duckworth, Angela L. , title =. 2025 , institution =

  84. [92]

    and Weintrop, David and Grossman, Tovi , title =

    Kazemitabaar, Majeed and Chow, Justin and Ma, Carl Ka To and Ericson, Barbara J. and Weintrop, David and Grossman, Tovi , title =. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , year =

  85. [93]

    and Masoud, Fadi D

    Kalam, Kazi A. and Masoud, Fadi D. and Muntaser, Adam and Ranga, Raghav and Geng, Xue and Goyal, Munish , title =. Cureus , year =

  86. [94]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics , pages =

    Vanzo, Alessandro and Pal Chowdhury, Sankalan and Sachan, Mrinmaya , title =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics , pages =. 2025 , eprint =

  87. [95]

    and Goldstein, Daniel G

    Kumar, Harsh and Rothschild, David M. and Goldstein, Daniel G. and Hofman, Jake M. , title =. Artificial Intelligence in Education (AIED 2025) , series =. 2025 , doi =

  88. [96]

    2025 , institution =

    Poulidis, Stefanos and Bastani, Hamsa and Bastani, Osbert , title =. 2025 , institution =

  89. [97]

    2025 , type =

    Hausman, Naomi and Rigbi, Oren and Weisburd, Sarit , title =. 2025 , type =

  90. [98]

    Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance , journal =

    Fan, Yizhou and Tang, Lixiang and Le, Huixiao and Shen, Kejie and Tan, Shufang and Zhao, Yueying and Shen, Yuan and Li, Xinyu and Ga. Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance , journal =....

  91. [99]

    Proceedings of the Twelfth ACM Conference on Learning @ Scale (L@S '25) , year =

    Nie, Allen and Chandak, Yash and Suzara, Miroslav and Malik, Ali and Woodrow, Juliette and Peng, Matt and Sahami, Mehran and Brunskill, Emma and Piech, Chris , title =. Proceedings of the Twelfth ACM Conference on Learning @ Scale (L@S '25) , year =

  92. [100]

    Artificial Intelligence in Education

    Henkel, Owen and Horne-Robinson, Hannah and Kozhakhmetova, Nessie and Lee, Amanda , title =. Artificial Intelligence in Education. Posters and Late Breaking Results (AIED 2024) , series =. 2024 , doi =

  93. [101]

    2024 , type =

    Wiles, Emma and Krayer, Lisa and Abbadi, Mohamed and Awasthi, Urvi and Kennedy, Ryan and Mishkin, Pamela and Sack, Daniel and Candelon, Fran. 2024 , type =

  94. [102]

    and Dubey, Rachit , title =

    Liu, Grace and Christian, Brian and Dumbalska, Tsvetomira and Bakker, Michiel A. and Dubey, Rachit , title =. 2026 , eprint =

  95. [103]

    2026 , eprint =

    Caosun, Michael and Aral, Sinan , title =. 2026 , eprint =

  96. [104]

    and Healy, Paul J

    Moore, Don A. and Healy, Paul J. , title =. Psychological Review , volume =. 2008 , doi =

  97. [105]

    2026 , eprint =

    Shen, Judy Hanwen and Tamkin, Alex , title =. 2026 , eprint =

  98. [106]

    Journal of Political Economy , volume=

    Learning by Doing and Learning from Others: Human Capital and Technical Change in Agriculture , author=. Journal of Political Economy , volume=. 1995 , publisher=

  99. [107]

    American Economic Review , volume=

    Learning about a New Technology: Pineapple in Ghana , author=. American Economic Review , volume=. 2010 , publisher=

  100. [108]

    Social Dynamics of

    Bursztyn, Leonardo and Imas, Alex and Jim. Social Dynamics of. 2025 , month=

  101. [109]

    , title =

    Campos, Christopher and Singleton, John D. , title =. 2026 , institution =

  102. [110]

    and Gilbert, Sam J

    Risko, Evan F. and Gilbert, Sam J. , title =. Trends in Cognitive Sciences , volume =. 2016 , doi =

  103. [111]

    Chi, Michelene T. H. and Wylie, Ruth , title =. Educational Psychologist , volume =. 2014 , doi =

  104. [112]

    and Rilke, Rainer Michael , title =

    Fischer, Mira and Rau, Holger A. and Rilke, Rainer Michael , title =. 2025 , month = dec, institution =

  105. [113]

    2026 , month = may, institution =

    Teaching with. 2026 , month = may, institution =

  106. [114]

    and Wen, C

    Huang, S. and Wen, C. and Bai, X. and Li, S. and Wang, S. and Wang, X. and Yang, D. , title =. Journal of Medical Internet Research , volume =. 2025 , doi =

  107. [115]

    and Ouyang, J

    Gan, W. and Ouyang, J. and Li, H. and Xue, Z. and Zhang, Y. and Dong, Q. and Huang, J. and Zheng, X. and Zhang, Y. , title =. Journal of Medical Internet Research , volume =. 2024 , doi =

  108. [116]

    2025 , eprint =

    Dai, Xusheng and Wen, Zhaochun and Jiang, Jianxiao and Liu, Huiqin and Zhang, Yu , title =. 2025 , eprint =

  109. [117]

    2026 , eprint =

    Hou, Xiaoyu and Xiao, Bo and Liu, Hexu and Mueller, Shane , title =. 2026 , eprint =

  110. [118]

    and Zhang, L

    Ba, H. and Zhang, L. and Yi, Z. , title =. BMC Medical Education , volume =. 2024 , doi =

  111. [119]

    Evaluation of

    Kavadella, Argyro and. Evaluation of. JMIR Medical Education , volume =. 2024 , doi =

  112. [120]

    and Chen, L

    Wu, C. and Chen, L. and Han, M. and Li, Z. and Yang, N. and Yu, C. , title =. Medical Teacher , volume =. 2025 , doi =

  113. [121]

    and Yin, K

    Li, J. and Yin, K. and Wang, Y. and Jiang, X. and Chen, D. , title =. BMC Medical Education , volume =. 2025 , doi =

  114. [122]

    Computers and Education: Artificial Intelligence , volume =

    Bassner, Patrick and Lenk-Ostendorf, Ben and Beinstingel, Ramona and Wasner, Tobias and Krusche, Stephan , title =. Computers and Education: Artificial Intelligence , volume =. 2026 , doi =

  115. [123]

    and Restrepo, P

    Acemoglu, D. and Restrepo, P. (2019). Automation and new tasks: How technology displaces and reinstates labor. Journal of Economic Perspectives , 33(2):3--30

  116. [124]

    Ammari, T., Chen, M., Zaman, S. M. M., and Garimella, K. (2025). How students (really) use ChatGPT : Uncovering experiences among undergraduate students

  117. [125]

    Andre, P., Haaland, I., Roth, C., Wiederholt, M., and Wohlfart, J. (2026). Narratives about the macroeconomy. Review of Economic Studies . Forthcoming

  118. [126]

    Ba, H., Zhang, L., and Yi, Z. (2024). Enhancing clinical skills in pediatric trainees: A comparative study of ChatGPT -assisted and traditional teaching methods. BMC Medical Education , 24:558

  119. [127]

    Baird, M., Carpanelli, M., Xu, B., and Xu, K. (2026). Firms' GitHub Copilot adoption and labor market outcomes for software engineers. Contemporary Economic Policy

  120. [128]

    Bassner, P., Lenk-Ostendorf, B., Beinstingel, R., Wasner, T., and Krusche, S. (2026). Less stress, better scores, same learning: The dissociation of performance and learning in AI -supported programming education. Computers and Education: Artificial Intelligence , 10:100537

  121. [129]

    Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakc , \"O ., and Mariman, R. (2025). Generative AI Without Guardrails Can Harm Learning : Evidence from High School Mathematics . Proceedings of the National Academy of Sciences , 122(26):e2422633122. Correction at https://doi.or...

  122. [130]

    Becker, J., Rush, N., Barnes, B., and Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. Technical report, METR

  123. [131]

    Belloni, A., Chernozhukov, V., and Hansen, C. (2014). High- Dimensional Methods and Inference on Structural and Treatment Effects . Journal of Economic Perspectives , 28(2):29--50

  124. [132]

    P., Fox, L., Loeb, S., and Taylor, E

    Bettinger, E. P., Fox, L., Loeb, S., and Taylor, E. S. (2017). Virtual classrooms: How online college courses affect student success. American Economic Review , 107(9):2855--2875

  125. [133]

    Bick, A., Blandin, A., and Deming, D. J. (2026). The Rapid Adoption of Generative AI . Management Science

  126. [134]

    Brynjolfsson, E., Li, D., and Raymond, L. (2025). Generative AI at Work . The Quarterly Journal of Economics , 140(2):889--942

  127. [135]

    and Mitchell, T

    Brynjolfsson, E. and Mitchell, T. (2017). What can machine learning do? workforce implications. Science , 358(6370):1530--1534

  128. [136]

    Brysbaert, M. (2019). How many words do we read per minute? a review and meta-analysis of reading rate. Journal of Memory and Language , 109:104047

  129. [137]

    O., Blom, J., Buszkiewicz, M., Halvorsen, N., Hassan, C., Roma \'n czyk, T., Holme, ., Jarus, K., Fielding, S., Kunar, M., Pellise, M., Pilonis, N., Kami \'n ski, M

    Budzy \'n , K., Roma \'n czyk, M., Kitala, D., Ko odziej, P., Bugajski, M., Adami, H. O., Blom, J., Buszkiewicz, M., Halvorsen, N., Hassan, C., Roma \'n czyk, T., Holme, ., Jarus, K., Fielding, S., Kunar, M., Pellise, M., Pilonis, N., Kami \'n ski, M. F., Kalager, M., Bretthau...

  130. [138]

    and Fairlie, R

    Bulman, G. and Fairlie, R. W. (2016). Technology and education: Computers, software, and the internet. In Hanushek, E. A., Machin, S. J., and Woessmann, L., editors, Handbook of the Economics of Education , volume 5, pages 239--280. Elsevier

  131. [139]

    Caldwell, J. E. (2007). Clickers in the large classroom: Current research and best-practice tips. CBE---Life Sciences Education , 6(1):9--20

  132. [140]

    P., Greenberg, K., and Walker, M

    Carter, S. P., Greenberg, K., and Walker, M. S. (2017). The impact of computer usage on academic performance: Evidence from a randomized trial at the United States Military Academy . Economics of Education Review , 56:118--132

  133. [141]

    Carver, R. P. (1992). Reading rate: Theory, research, and practical implications. Journal of Reading , 36(2):84--95

  134. [142]

    J., Hitzig, Z., Ong, C., Shan, C

    Chatterji, A., Cunningham, T., Deming, D. J., Hitzig, Z., Ong, C., Shan, C. Y., and Wadman, K. (2025). How people use ChatGPT . Working Paper 34255, National Bureau of Economic Research

  135. [143]

    Chi, M. T. H. and Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist , 49(4):219--243

  136. [144]

    and Lee, H.-y

    Chiang, C.-H. and Lee, H.-y. (2023). Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages 15607--15631. Association for Computational Linguistics

  137. [145]

    H., Monahan, A

    Choi, J. H., Monahan, A. B., and Schwarcz, D. (2024). Lawyering in the age of artificial intelligence. Minnesota Law Review , 109(1):147--218

  138. [146]

    T.-H., Zhang, B., Kung, L.-C., Bastani, H., and Bastani, O

    Chung, A. T.-H., Zhang, B., Kung, L.-C., Bastani, H., and Bastani, O. (2026). Effective personalized AI tutors via LLM-Guided reinforcement learning. SSRN Working Paper 6423358, University of Pennsylvania

  139. [147]

    and Reyes, G

    Contractor, Z. and Reyes, G. (2026). Generative AI in higher education: Evidence from an elite college. IZA Discussion Paper Nr. 18055

  140. [148]

    Cristia, J., Ibarrar \'a n, P., Cueto, S., Santiago, A., and Sever \'i n, E. (2017). Technology and child development: Evidence from the One Laptop per Child program. American Economic Journal: Applied Economics , 9(3):295--320

  141. [149]

    H., and Lombardi, M

    Cruces, G., Fern \'a ndez Meijide , D., Galiani, S., G \'a lvez, R. H., and Lombardi, M. (2026). Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment. Technical Report 34851, National Bureau of Economic Research

  142. [150]

    K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., and Salz, T

    Cui, Z. K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., and Salz, T. (2026). The Effects of Generative AI on High-Skilled Work : Evidence from Three Field Experiments with Software Developers . Management Science

  143. [151]

    Dai, X., Wen, Z., Jiang, J., Liu, H., and Zhang, Y. (2025). How students use AI feedback matters: Experimental evidence on physics achievement and autonomy

  144. [152]

    E., Tiberti, F., Barr \'o n Rodr \'i guez, M., Manolio, F., Mosuro, W., and Dikoru, E

    De Simone, M. E., Tiberti, F., Barr \'o n Rodr \'i guez, M., Manolio, F., Mosuro, W., and Dikoru, E. J. (2025). From chalkboards to chatbots: Evaluating the impact of generative AI on learning outcomes in Nigeria . Policy Research Working Paper 11125, World Bank

  145. [153]

    Dell'Acqua, F., Ayoubi, C., Lifshitz, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S., and Lakhani, K. R. (2025). The cybernetic teammate: A field experiment on generative AI reshaping teamwork and expertise. Technical Report 33641, National B...

  146. [154]

    R., Lifshitz-Assaf , H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K

    Dell'Acqua, F., McFowland, E., Mollick, E. R., Lifshitz-Assaf , H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. (2026). Navigating the Jagged Technological Frontier : Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledg...

  147. [155]

    and Laird, N

    DerSimonian, R. and Laird, N. (1986). Meta-analysis in clinical trials. Controlled Clinical Trials , 7(3):177--188

  148. [156]

    J., Goodman, S., and Smith, J

    Dettling, L. J., Goodman, S., and Smith, J. (2018). Every little bit counts: The impact of high-speed internet on the transition to college. The Review of Economics and Statistics , 100(2):260--273

  149. [157]

    Doshi, A. R. and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances , 10(28):eadn5290

  150. [158]

    and Spero, M

    Emi, B. and Spero, M. (2024). Technical report on the Pangram AI -generated text classifier. Technical report, Pangram Labs

  151. [159]

    J., Oreopoulos, P., and Quan, V

    Escueta, M., Nickow, A. J., Oreopoulos, P., and Quan, V. (2020). Upgrading education with technology: Insights from experimental research. Journal of Economic Literature , 58(4):897--996

  152. [160]

    N., Rush, M., and Yin, L

    Figlio, D. N., Rush, M., and Yin, L. (2013). Is it live or is it internet? Experimental estimates of the effects of online instruction on student learning. Journal of Labor Economics , 31(4):763--784

  153. [161]

    A., and Rilke, R

    Fischer, M., Rau, H. A., and Rilke, R. M. (2025). AI tutoring enhances student learning without crowding out reading effort. IZA Discussion Paper 18338, IZA Institute of Labor Economics

  154. [162]

    Gan, W., Ouyang, J., Li, H., Xue, Z., Zhang, Y., Dong, Q., Huang, J., Zheng, X., and Zhang, Y. (2024). Integrating ChatGPT in orthopedic education for medical undergraduates: Randomized controlled trial. Journal of Medical Internet Research , 26:e57037

  155. [163]

    Goldsmith-Pinkham, P., Tan, C., and Zentefis, A. K. (2026). Human- AI collaboration in radiology: The case of pulmonary embolism. arXiv preprint arXiv:2601.13379

  156. [164]

    Handa, K., Bent, D., Tamkin, A., McCain, M., Durmus, E., Stern, M., Schiraldi, M., Huang, S., Ritchie, S., Syverud, S., Jagadish, K., Vo, M., Bell, M., and Ganguli, D. (2025). A nthropic education report: How university students use C laude

  157. [165]

    Hirabayashi, S., Jain, R., Jurkovi \'c , N., and Wu, G. (2024). Harvard undergraduate survey on generative AI . arXiv preprint arXiv:2406.00833

  158. [166]

    Hou, X., Xiao, B., Liu, H., and Mueller, S. (2026). The role of instructional guidance in generative AI -assisted learning: Empirical evidence from construction engineering education

  159. [167]

    Huang, S., Wen, C., Bai, X., Li, S., Wang, S., Wang, X., and Yang, D. (2025). Exploring the application capability of ChatGPT as an instructor in skills education for dental medical students: Randomized controlled trial. Journal of Medical Internet Research , 27:e68538

  160. [168]

    and Gilbert, E

    Hutto, C. and Gilbert, E. (2014). Vader: A parsimonious rule-based model for sentiment analysis of social media text. In Proceedings of the International AAAI Conference on Weblogs and Social Media (ICWSM) , volume 8, pages 216--225

  161. [169]

    and Imas, A

    Jabarian, B. and Imas, A. (2025). Artificial writing and automated detection. NBER Working Paper 34223, National Bureau of Economic Research

  162. [170]

    Jackson, C. K. and Mackevicius, C. L. (2024). What impacts can we expect from school spending policy? Evidence from evaluations in the U.S. American Economic Journal: Applied Economics , 16(1):412--446

  163. [171]

    Jensen, R. (2010). The (perceived) returns to education and the demand for schooling. The Quarterly Journal of Economics , 125(2):515--548

  164. [172]

    Jia, N., Luo, X., Fang, Z., and Liao, C. (2024). When and how artificial intelligence augments employee creativity. Academy of Management Journal , 67(1):5--32

  165. [173]

    Kanazawa, K., Kawaguchi, D., Shigeoka, H., and Watanabe, Y. (2025). AI , skill, and productivity: The case of taxi drivers. Management Science

  166. [174]

    A., Kaklamanos, E

    Kavadella, A., Dias da Silva , M. A., Kaklamanos, E. G., Stamatopoulos, V., and Giannakopoulos, K. (2024). Evaluation of ChatGPT 's real-life implementation in undergraduate dental education: Mixed methods study. JMIR Medical Education , 10:e51344

  167. [175]

    Kazemitabaar, M., Chow, J., Ma, C. K. T., Ericson, B. J., Weintrop, D., and Grossman, T. (2023). Studying the effect of AI code generators on supporting novice learners in introductory programming. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . ACM

  168. [176]

    Kestin, G., Miller, K., Klales, A., Milbourne, T., and Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports , 15:17458

  169. [177]

    Kim, D., Mitrofanov, D., Wen, Q., and Xu, T. (2025). Generative AI can improve performance and engagement without harming learning. SSRN Working Paper 5929576

  170. [178]

    Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher , 49(4):241--253

  171. [179]

    M., Sellen, A., Rintel, S., Goldstein, D

    Kreijkes, P., Kewenig, V., Kuvalja, M., Lee, M., Vitello, S., Hofman, J. M., Sellen, A., Rintel, S., Goldstein, D. G., Rothschild, D. M., Tankelevitch, L., and Oates, T. (2026). Effects of LLM use and note-taking on reading comprehension and memory: A randomised experiment in ...

  172. [180]

    M., Goldstein, D

    Kumar, H., Rothschild, D. M., Goldstein, D. G., and Hofman, J. M. (2025). Math education with large language models: Peril or promise? In Artificial Intelligence in Education (AIED 2025) , Lecture Notes in Computer Science. Springer

  173. [181]

    LearnLM : Improving Gemini for Learning

    LearnLM Team (2024). LearnLM : Improving Gemini for Learning . Google DeepMind Technical Report

  174. [182]

    Teaching with Gemini : Measuring the impact of Guided Learning on student mathematics progress in Sierra Leone

    LearnLM Team (2026). Teaching with Gemini : Measuring the impact of Guided Learning on student mathematics progress in Sierra Leone . Technical report, Google DeepMind and Fab AI

  175. [183]

    H., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., and Wilson, N

    Lee, H.-P. H., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., and Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of th...

  176. [184]

    B., and Sting, F

    Lehmann, M., Cornelius, P. B., and Sting, F. J. (2025). AI meets the classroom: When do large language models harm learning?

  177. [185]

    Lewin, C., Somekh, B., and Steadman, S. (2008). Embedding interactive whiteboards in teaching and learning: The process of change in pedagogic practice. Education and Information Technologies , 13(4):291--303

  178. [186]

    G., Ungar, L., and Duckworth, A

    Lira, B., Rogers, T., Goldstein, D. G., Ungar, L., and Duckworth, A. L. (2025). Coach not crutch: Evidence that AI can improve writing skill despite reducing effort. Technical report, University of Pennsylvania. arXiv:2502.02880

  179. [187]

    A., and Dubey, R

    Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., and Dubey, R. (2026). AI assistance reduces persistence and hurts independent performance

  180. [188]

    Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-Eval : NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages 2511--2522. Association for Computational Linguistics

  181. [189]

    and Pop-Eleches, C

    Malamud, O. and Pop-Eleches, C. (2011). Home computer use and the development of human capital. The Quarterly Journal of Economics , 126(2):987--1027

  182. [190]

    Masrour, E. (2025). Introducing Pangram 's plagiarism detection. Pangram Labs blog post

  183. [191]

    N., and Spero, M

    Masrour, E., Emi, B. N., and Spero, M. (2025). DAMAGE : Detecting adversarially modified AI generated text. In Proceedings of the 1st Workshop on Detecting AI Generated Content (GenAIDetect), COLING , pages 71--86

  184. [192]

    Meincke, L., Nave, G., and Terwiesch, C. (2025). ChatGPT decreases idea diversity in brainstorming. Nature Human Behaviour , 9(6):1107--1109

  185. [193]

    E., and Kushlev, K

    Moon, K., Green, A. E., and Kushlev, K. (2025). Homogenizing effect of large language models ( LLMs ) on creative diversity: An empirical comparison of human and ChatGPT writing. Computers in Human Behavior: Artificial Humans , 6:100207

  186. [194]

    C., Johnson, D

    Moon, K., Kushlev, K., Bank, A., Lira Luttges , B., Viskontas, I., Kaufman, J. C., Johnson, D. R., Duckworth, A., and Green, A. (2026). The creative link between words and ideas is weakening in the AI era. PsyArXiv preprint

  187. [195]

    Mozannar, H., Bansal, G., Fourney, A., and Horvitz, E. (2024). Reading between the lines: Modeling user behavior and costs in AI -assisted programming. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . ACM

  188. [196]

    Nickow, A., Oreopoulos, P., and Quan, V. (2024). The promise of tutoring for PreK --12 learning: A systematic review and meta-analysis of the experimental evidence. American Educational Research Journal , 61(1):74--107

  189. [197]

    and Zhang, W

    Noy, S. and Zhang, W. (2023). Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence . Science , 381(6654):187--192

  190. [198]

    Building an AI -ready workforce: A look at college student ChatGPT adoption in the US

    OpenAI (2025). Building an AI -ready workforce: A look at college student ChatGPT adoption in the US . Technical report, OpenAI

  191. [199]

    Peng, S., Kalliamvakou, E., Cihon, P., and Demirer, M. (2023). The Impact of AI on Developer Productivity : Evidence from GitHub Copilot

  192. [200]

    Poulidis, S., Bastani, H., and Bastani, O. (2025). Self-regulated AI use hinders long-term learning. SSRN Working Paper 5604932, University of Pennsylvania

  193. [201]

    Rav s elj, D., Ker z i c , D., Toma z evi c , N., Umek, L., Brezovar, N., et al. (2025). Higher Education Students ' Perceptions of ChatGPT : A Global Study of Early Reactions . PLOS ONE , 20(2):e0315011

  194. [202]

    and Gurevych, I

    Reimers, N. and Gurevych, I. (2019). Sentence- BERT : Sentence embeddings using Siamese BERT -networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing , pages 3982--3992. Association for Computational Linguistics

  195. [203]

    Risko, E. F. and Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences , 20(9):676--688

  196. [204]

    Shen, J. H. and Tamkin, A. (2026). How AI impacts skill formation

  197. [205]

    Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y. (2020). MPNet : Masked and permuted pre-training for language understanding. In Advances in Neural Information Processing Systems (NeurIPS) , volume 33, pages 16857--16867

  198. [206]

    o hr, C., Ou, A. W., and Malmstr \

    St \"o hr, C., Ou, A. W., and Malmstr \"o m, H. (2024). Perceptions and Usage of AI Chatbots Among Students in Higher Education Across Genders , Academic Levels and Fields of Study . Computers and Education: Artificial Intelligence , 7:100259

  199. [207]

    Str \"o mberg, D., Lei, V., and Wu, Y. (2026). The generative AI learning penalty: Evidence from Chinese secondary education. CEPR Discussion Paper 21577, Centre for Economic Policy Research

  200. [208]

    Thai, K., Emi, B., Masrour, E., and Iyyer, M. (2026). EditLens : Quantifying the extent of AI editing in text. In International Conference on Learning Representations (ICLR)

  201. [209]

    L., Ladd, H

    Vigdor, J. L., Ladd, H. F., and Martinez, E. (2014). Scaling the digital divide: Home computer technology and student achievement. Economic Inquiry , 52(3):1103--1119

  202. [210]

    and Zafar, B

    Wiswall, M. and Zafar, B. (2015). Determinants of college major choice: Identification using an information experiment. The Review of Economic Studies , 82(2):791--824

  203. [211]

    Xu, X., Qiao, L., Cheng, N., Liu, H., and Zhao, W. (2025). Enhancing self-regulated learning and learning experience in generative AI environments: The critical role of metacognitive support. British Journal of Educational Technology , 56(5):1842--1863

  204. [212]

    Yakura, H., Lopez-Lopez , E., Brinkmann, L., Serna, I., Gupta, P., Soraperra, I., and Rahwan, I. (2024). Empirical evidence of large language model's influence on human spoken communication

  205. [213]

    P., Zhang, H., Gonzalez, J

    Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023). Judging LLM -as-a-judge with MT-Bench and Chatbot Arena . In Advances in Neural Information Processing Systems , volume 36

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.