Pith. sign in

REVIEW 3 major objections 7 minor 87 references

Injecting realistic bugs into otherwise correct GenAI code steers students toward editing and verification instead of only re-prompting.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 09:17 UTC pith:JNXTBEIO

load-bearing objection Solid large-scale classroom evidence that injected near-miss bugs shift CS1 students from reprompting to local code repair inside authentic GenAI workflows. the 3 major comments →

arxiv 2607.05068 v1 pith:JNXTBEIO submitted 2026-07-06 cs.SE cs.CY

When AI Is Wrong on Purpose: How Students Respond to Buggy GenAI Code

classification cs.SE cs.CY
keywords natural language programmingcode-generating AIbug injectiondebuggingprompt-based programmingCS1code reviewverification
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Introductory programming courses that have students prompt AI for code face a practical problem: modern models often produce correct solutions on CS1 tasks, so students have little reason to inspect the output carefully. This paper shows that deliberately injecting small, runnable near-miss bugs into otherwise correct generated code changes student behavior. Across 2,636 sessions from 917 students, injected bugs more often led to direct code edits and higher next-attempt success, while natural prompt-related failures more often led students to refine their natural-language specifications. Student reflections describe the activity as useful practice in reading, reviewing, and debugging AI code, plus greater awareness that AI output still needs human verification. Together, natural failures and designed near-misses support a single workflow that trains both specification refinement and careful code repair.

Core claim

Students treat natural prompt-related failures and deliberately injected near-miss bugs as different kinds of problems: natural bugs mainly drive specification refinement through re-prompting, while injected bugs mainly drive localized code edits and higher immediate success, making verification and repair an explicit part of prompt-based GenAI programming.

What carries the argument

A bug-injection pipeline that intercepts GenAI-generated code: if the code fails hidden tests it is shown as a natural bug; if it passes, a validated near-miss variant with a subtle runnable fault is shown instead, forcing students to inspect and repair rather than accept the output.

Load-bearing premise

That differences in student action and success can be attributed mainly to bug source rather than to the design fact that injected bugs are created only after the model first produced fully correct code, making them systematically closer to a working solution.

What would settle it

An A/B classroom study that randomly enables or disables injection (or matches natural and injected bugs for distance-to-correctness) and still finds the same action and success gap between bug sources would support the claim; if the gap disappears when nearness is controlled, the claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper studies how CS1 students respond to buggy GenAI code inside a prompt-based programming platform that mixes naturally failing model outputs with deliberately injected, validated near-miss bugs. Using 2,636 filtered sessions from 917 students (6,071 bug-fixing turns), the authors operationalize turn-level labels (n:first/i:first/n:any/i:any), first-action choices (prompt/edit/noop), and immediate success (edit-and-run success vs. counterfactual pre-injection success after a prompt). They report that injected near-misses are associated with more edit-first behavior, smaller edits, fewer edit-and-run cycles, and higher immediate success, while natural bugs are associated with more prompt-first behavior and specification-oriented strategies (signature updates, task extension, reframing). Post-lab Likert and thematically coded free-text reflections emphasize code understanding, debugging, and awareness of GenAI limits. The authors argue that combining natural failures with designed near-misses can support both specification refinement and verification/repair in the same workflow.

Significance. This is a timely, well-scoped contribution to GenAI-in-CS1 research. The core empirical result—an observational association between failure context (specification-mismatch vs. near-miss) and students’ next repair move—is directly useful for platform and activity design. Strengths include classroom scale, a clear middleware pipeline with audit logs, independence-aware testing restricted to first buggy turns, dual quantitative/qualitative analysis with strong IRR (Krippendorff α ≈ 0.87 for prompts; mean α ≈ 0.82 for reflections), and an explicit non-causal framing of natural vs. injected comparisons. The work productively bridges prompt-problems literature with controlled-failure debugging practice and offers concrete design implications (diagnose-before-action scaffolds; near-misses as verification scaffolds). Even without a randomized no-injection arm or objective learning gains, the descriptive behavioral map is a solid, citable result for the community.

major comments (3)
  1. [Abstract; §3.2; §3.5; §4.2 RQ1 summary] §3.2 pipeline + §3.5/RQ1 interpretation: Injected bugs exist only after a candidate already passes hidden tests, so they are near-misses by construction. The paper repeatedly states that comparisons are associations, not causal effects of “bug source” (§3.5; RQ1 summary; §5.3). That is correct and load-bearing. However, the abstract, strongest claims, and several result sentences still read as if “injected vs. natural” itself drives edits/success. Please consistently rephrase the operative contrast as near-miss / localized fault vs. specification-or-logic mismatch (of which injection is one engineered source), so the non-causal reading cannot be missed by abstract-only readers.
  2. [§3.5; Fig. 6b; §4.2] §3.5 immediate-success definition for prompt actions: Success after a prompt is defined counterfactually via the pre-injection audit of the next assistant output. This is methodologically careful, but it is easy to misread the high prompt-success rates after injected bugs (e.g., ~87–88% in Fig. 6b) as “students fixed the task by reprompting and received a working solution.” Under the pipeline, a subsequent correct generation is typically re-injected. The manuscript should state more prominently in §4.2 and the figure caption what “immediate success” means for prompt turns, and avoid language that implies students completed the task via prompting alone after an injected bug.
  3. [Table 1; Fig. 6; §4.1–4.2] Table 1 / pooled analyses: Injection rate and session success vary sharply by problem (Array II: ~51% injected, 99% success; Binary: ~7% injected, ~51% success). Pooled n:any/i:any fractions and success rates are therefore compositionally sensitive to easier, high-injection tasks. Per-problem tests are reported, which is good, but the main narrative and abstract lean on pooled contrasts. Please either (a) weight or stratify pooled summaries more carefully, or (b) lead with the consistent per-problem directional pattern and treat pooled numbers as secondary descriptive summaries.
minor comments (7)
  1. [Fig. 1; §3.1] Fig. 1 caption and §3.1: The walkthrough is excellent; a one-line note that the student is not told whether a failure is natural or injected would help readers who skip §3.3.
  2. [Table 2] Table 2: “Nb. bugs” after normalization is useful; briefly state whether frequency ranking is by occurrence count or distinct students, so readers can judge representativeness of the two examples.
  3. [Fig. 7e] Fig. 7e: Report sample sizes already appear in row labels; adding absolute counts in cells (or a companion table) would make rare strategies easier to interpret, especially for n:first/i:first.
  4. [§3.4] §3.4 filtering: 177/2813 sessions (6.3%) removed for non-alternation or missing code. A short sensitivity note (e.g., whether removed sessions were harder problems) would strengthen robustness claims.
  5. [§4.4; Fig. 8c] RQ3 coding: Concept learning and Bug types have lower α (0.61, 0.59). The paper already treats them cautiously; consider demoting them in the main frequency plot or merging with neighboring themes to avoid over-interpreting low-reliability labels.
  6. [Abstract; §1; §5] Minor prose: “2,636studentsessions” and similar missing spaces appear in the abstract/intro; fix spacing and a few long sentences in §5.1–5.2 for readability.
  7. [§3.3; §5.3] Ethics/credit design: Students needed only 2/5 tasks for credit and Binary had low completion. Mention possible self-selection into harder tasks when discussing generalizability in §5.3.

Circularity Check

0 steps flagged

No significant circularity: observational associations from classroom logs, not predictions forced by definition or self-citation.

full rationale

This is an empirical CS-education study. The load-bearing claims are turn-level associations (edit vs. prompt rates, immediate success, coded prompt strategies) measured from 2,636 sessions and 6,071 bug-fixing turns, plus post-task reflections. Nothing is defined in terms of the quantity it is said to predict; there is no fitted parameter re-labeled as a prediction; and no uniqueness theorem or ansatz is imported from prior author work to force the result. The design fact that injected bugs are created only after a candidate already passes hidden tests (pipeline §3.2) makes them near-misses by construction and partly explains higher immediate success—but the authors explicitly interpret natural-vs-injected comparisons as differences in response behavior rather than causal effects of bug source (§3.5, RQ1 summary, Limitations). That is a stated design confound and a limit on causal learning claims, not circular derivation. Self-citations to the Prompt Programming platform and related tools are infrastructure references, not load-bearing uniqueness arguments. The reported patterns (more edit-first and higher success after injected near-misses; more specification-oriented prompting after natural failures) are grounded in logs and qualitative coding and do not reduce to their inputs by definition.

Axiom & Free-Parameter Ledger

1 free parameters · 4 axioms · 1 invented entities

Empirical classroom study; load-bearing premises are design and measurement choices rather than free physical constants. The main free parameter is the number of candidate buggy variants generated per injection (set to 5). Core axioms are the operational definitions of natural vs. injected bugs and the first-action coding rules. Invented entities are the two bug-source labels and the slip-through category, introduced for analysis.

free parameters (1)
  • number of candidate buggy variants per injection = 5
    Hard-coded to 5; first compiling-and-failing candidate is served. Affects injection success rate and therefore the observed mix of i vs. slip-through turns.
axioms (4)
  • ad hoc to paper A GenAI response that fails hidden tests after a student prompt is a natural bug; a response that would have passed but is replaced by a validated near-miss is an injected bug.
    Operational definition introduced in Section 3.2 and used for all subsequent labeling.
  • ad hoc to paper Student next-step response is the first logged action (prompt before any edit, or edit before any prompt); turns with neither are noop.
    Defines the primary outcome variables for RQ1 (Section 3.4).
  • ad hoc to paper Immediate success after a prompt is defined counterfactually by the pre-injection audit of the next model output.
    Necessary because injected bugs would otherwise make every subsequent generation appear to fail (Section 3.5).
  • domain assumption Students in a graded CS1 lab who are warned that the model makes mistakes will treat the activity as authentic verification practice.
    Background assumption of the classroom deployment (Section 3.3).
invented entities (1)
  • natural bug (n) vs. injected bug (i) labels, plus slip-through no independent evidence
    purpose: Partition every GenAI response so that student behavior can be compared across failure sources.
    Analytic categories created by the middleware; no independent existence outside the study design.

pith-pipeline@v1.1.0-grok45 · 28590 in / 2784 out tokens · 26075 ms · 2026-07-11T09:17:03.658986+00:00 · methodology

0 comments
read the original abstract

As Generative AI (GenAI) becomes increasingly central to software development, CS education is integrating prompt-centered workflows where students describe intended program behavior in natural language to elicit code. However, professional practice requires careful review and verification of GenAI-generated code that may appear correct while containing subtle faults. This creates a challenge for CS1-level activities, where current models often solve tasks correctly and reduce students' incentive to closely inspect generated outputs. We investigate how prompt-centered programming activities can be adapted to better foster these practices. Specifically, we explore an approach where realistic, runnable bugs are injected into otherwise correct solutions, thus requiring students to read and repair generated outputs. We analyzed 2,636 sessions from 917 students, and examined behavior across instances of naturally occurring prompt-related failures and deliberately injected bugs within each session. Our findings show that students responded differently across bug sources. Deliberately injected bugs more often led to direct code edits and higher next-attempt success, suggesting localized repair of near-miss solutions. Prompt-related failures instead more often led students to refine prompts by clarifying constraints, updating function signatures, adding edge cases, or reframing the task. Student reflections reinforce the emphasis on review and repair, describing useful practice in code understanding, code review, and debugging, as well as a more careful verification mindset and greater awareness of GenAI limitations. Ultimately, prompt-related failures and injected bugs together support a pedagogically useful GenAI workflow, where students practice both specification refinement through prompts and debugging through code editing.

Figures

Figures reproduced from arXiv: 2607.05068 by Adish Singla, Ahana Ghosh, Alkis Gotovos, James Prather, Juho Leinonen, Jyotika Mahapatra, Kaitlin Riegel, Paul Denny, Victor-Alexandru P\u{a}durean.

Figure 1
Figure 1. Figure 1: The interface of Prompt Programming [46], exemplified on Array IV. The interaction proceeds as follows: (1) the student sends an initial prompt that mis-specifies the task (reversing the array segment from start to end); (2) the GenAI assistant responds with a reversing implementation, which is incorrect for the intended sorting task; (3) the student reprompts, correcting the request to sort the segment in… view at source ↗
Figure 2
Figure 2. Figure 2: Bug injection pipeline in our study. After a student submits a prompt, the GenAI assistant generates code that is [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Classroom deployment. (a) presents the two stage data collection procedure in the study. In the first stage, data is [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Examples of student sessions, along with the terminology used throughout analysis. A session consists of a sequence [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Data analysis procedures. Overview of the data analysis procedures across research questions. (top) For RQ1 and [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: RQ1 quantitative results. (a) shows the first-action distributions after a buggy GenAI response (shown for four [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: RQ2 mixed-method results. (a) shows the mean absolute Levenshtein distance between the initial buggy GenAI code [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: RQ3 reflections on learning with buggy GenAI code. (a) shows perceived ease of locating and fixing bugs and [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

87 extracted references · 2 linked inside Pith

  1. [1]

    Vincent Aleven, Elmar Stahl, Silke Schworm, Frank Fischer, and Raven Wallace

  2. [2]

    Review of Educational Research(2003)

    Help Seeking and Help Design in Interactive Learning Environments. Review of Educational Research(2003)

  3. [3]

    Leonardo Banh, Florian Holldack, and Gero Strobel. 2025. Copiloting the Future: How Generative AI Transforms Software Engineering.Information and Software Technology(2025)

  4. [4]

    Elizabeth L Bjork and Robert A Bjork. 2011. Making Things Hard on Yourself, But in a Good Way: Creating Desirable Difficulties to Enhance Learning.Psychology and the Real World: Essays Illustrating Fundamental Contributions to Society(2011)

  5. [5]

    Robert A. Bjork. 1994. Memory and Metamemory Considerations in the Training of Human Beings. InMetacognition: Knowing about Knowing. MIT Press

  6. [6]

    Virginia Braun and Victoria Clarke. 2006. Using Thematic Analysis in Psychology. Qualitative Research in Psychology(2006)

  7. [7]

    Nikitha Donekal Chandrashekar, Sehrish Basir Nizamani, Margaret Ellis, and Naren Ramakrishnan. 2026. Demystify, Use, Reflect: Preparing Students To Be Informed LLM-users. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  8. [8]

    Michelene TH Chi, Nicholas De Leeuw, Mei-Hung Chiu, and Christian LaVancher

  9. [9]

    Eliciting Self-explanations Improves Understanding.Cognitive science (1994)

  10. [10]

    Heeryung Choi, Tung Phung, Mengyan Wu, Adish Singla, and Christopher Brooks. 2025. Reflection-Satisfaction Tradeoff: Investigating Impact of Reflec- tion on Student Engagement with AI-Generated Programming Hints.CoRR abs/2512.04630 (2025)

  11. [11]

    Victoria Clarke and Virginia Braun. 2017. Thematic Analysis.The Journal of Positive Psychology(2017)

  12. [12]

    Gergely Márk Csányi, István Üveges, Dorina Lakatos, Dóra Ripszám, Kornélia Kozák, Dániel Nagy, and János Pál Vadász. 2025. Sentence-Level Rhetorical Role Labeling in Judicial Decisions.Big Data and Cognitive Computing(2025)

  13. [13]

    Cunningham

    Mehmet Arif Demirtas, Max Fowler, Nicole Hu, and Kathryn I. Cunningham. 2024. Validating, Refining, and Identifying Programming Plans Using Learning Curve Analysis on Code Writing Data. InProceedings of the Conference on International Computing Education Research (ICER)

  14. [14]

    Paul Denny et al . 2024. Computing Education in the Era of Generative AI. Commun. ACM(2024)

  15. [15]

    Heffernan, Tanja Käser, Steven Moore, Anna N

    Paul Denny, Sumit Gulwani, Neil T. Heffernan, Tanja Käser, Steven Moore, Anna N. Rafferty, and Adish Singla. 2024. Generative AI for Education (GAIED): Advances, Opportunities, and Challenges.CoRRabs/2402.01580 (2024)

  16. [16]

    Becker, and Brent N

    Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2023. Promptly: Using Prompt Problems to Teach Learners How to Effectively Utilize AI Code Generators.CoRR abs/2307.16364 (2023)

  17. [17]

    Becker, and Brent N

    Paul Denny, Juho Leinonen, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, and Brent N. Reeves. 2024. Prompt Problems: A New Programming Exercise for the Generative AI Era. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  18. [18]

    Stephen H. Edwards. 2004. Using Software Testing to Move Students from Trial- and-Error to Reflection-in-Action. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  19. [19]

    Anastasia Efklides. 2011. Interactions of Metacognition With Motivation and Affect in Self-Regulated Learning: The MASRL Model.Educational Psychologist (2011)

  20. [20]

    Sue Fitzgerald, Gary Lewandowski, Renée McCauley, Laurie Murphy, Beth Simon, Lynda Thomas, and Carol Zander. 2008. Debugging: Finding, Fixing and Flailing, a Multi-institutional Study of Novice Debuggers.Computer Science Education18 (2008)

  21. [21]

    Sharon Nelson-Le Gall. 1981. Help-seeking: An Understudied Problem-solving Skill in Children.Developmental Review(1981)

  22. [22]

    John Hattie and Helen Timperley. 2007. The Power of Feedback.Review of Educational Research(2007). ICER 2026 Vol. 1, August 11–14, 2026, Uppsala, Sweden Victor-Alexandru Pădurean et al

  23. [23]

    Xinying Hou, Ruiwei Xiao, Runlong Ye, Michael Liut, and John C. Stamper

  24. [24]

    InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

    Exploring Student Choice and the Use of Multimodal Generative AI in Programming Learning. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  25. [25]

    Manu Kapur. 2008. Productive Failure.Cognition and Instruction(2008)

  26. [26]

    Katz and John R

    Irvin R. Katz and John R. Anderson. 1987. Debugging: An Analysis of Bug- Location Strategies.Human-Computer Interaction(1987)

  27. [27]

    Ericson, David Weintrop, and Tovi Grossman

    Majeed Kazemitabaar, Justin Chow, Carl Ka To Ma, Barbara J. Ericson, David Weintrop, and Tovi Grossman. 2023. Studying the Effect of AI Code Generators on Supporting Novice Learners in Introductory Programming. InProceedings of the Conference on Human Factors in Computing Systems (CHI)

  28. [28]

    Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman

    Majeed Kazemitabaar, Xinying Hou, Austin Z. Henley, Barbara Jane Ericson, David Weintrop, and Tovi Grossman. 2023. How Novices Use LLM-based Code Generators to Solve CS1 Coding Tasks in a Self-Paced Learning Environment. In Proceedings of the Koli Calling International Conference on Computing Education Research (Koli Calling)

  29. [29]

    Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. InProceedings of the Conference on Human Factors in Computing Systems (CHI)

  30. [30]

    Nina Keith and Michael Frese. 2008. Effectiveness of Error Management Training: A Meta-analysis.Journal of Applied Psychology(2008)

  31. [31]

    Smith, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil

    Chris Kerslake, Paul Denny, David H. Smith, James Prather, Juho Leinonen, Andrew Luxton-Reilly, and Stephen MacNeil. 2024. Integrating Natural Language Prompting Tasks in Introductory Programming Courses. InProceedings of the Virtual Global Computing Education Conference (SIGCSE Virtual)

  32. [32]

    Ko and Brad A

    Amy J. Ko and Brad A. Myers. 2008. Debugging Reinvented: Asking and Answer- ing Why and Why Not Questions About Program Behavior. InProceedings of the International Conference on Software Engineering (ICSE)

  33. [33]

    Nachiket Kotalwar, Alkis Gotovos, and Adish Singla. 2024. Hints-in-browser: Benchmarking Language Models for Programming Feedback Generation. In Annual Conference on Neural Information Processing Systems (NeurIPS)

  34. [34]

    Klaus Krippendorff. 2011. Computing Krippendorff’s Alpha-Reliability

  35. [35]

    2018.Content Analysis: An Introduction to Its Methodology

    Klaus Krippendorff. 2018.Content Analysis: An Introduction to Its Methodology. SAGE Publications

  36. [36]

    Reeves, Paul Denny, James Prather, and Brett A

    Juho Leinonen, Arto Hellas, Sami Sarsa, Brent N. Reeves, Paul Denny, James Prather, and Brett A. Becker. 2023. Using Large Language Models to Enhance Programming Error Messages. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  37. [37]

    Evanfiya Logacheva, Arto Hellas, James Prather, Sami Sarsa, and Juho Leinonen

  38. [38]

    InProceedings of the Conference on International Computing Education Research (ICER)

    Evaluating Contextually Personalized Programming Exercises Created with Generative AI. InProceedings of the Conference on International Computing Education Research (ICER)

  39. [39]

    Koedinger, and Sherry Tongshuang Wu

    Qianou Ma, Hua Shen, Kenneth R. Koedinger, and Sherry Tongshuang Wu. 2024. How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging. InProceedings of the Artificial Intelligence in Education (AIED)

  40. [40]

    Becker, Michel Wermelinger, and Karen Reid

    Stephen MacNeil, Juho Leinonen, Paul Denny, Natalie Kiesler, Arto Hellas, James Prather, Brett A. Becker, Michel Wermelinger, and Karen Reid. 2024. Discussing the Changing Landscape of Generative AI in Computing Education. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  41. [41]

    Renée McCauley, Sue Fitzgerald, Gary Lewandowski, Laurie Murphy, Beth Simon, Lynda Thomas, and Carol Zander. 2008. Debugging: A Review of the Literature from an Educational Perspective.Computer Science Education(2008)

  42. [42]

    Tilman Michaeli and Ralf Romeike. 2019. Improving Debugging Skills in the Class- room: The Effects of Teaching a Systematic Debugging Process. InProceedings of the Workshop in Primary and Secondary Computing Education (WiPSCE)

  43. [43]

    Ran Mo, Dongyu Wang, Wenjing Zhan, Yingjie Jiang, Yepeng Wang, Yuqi Zhao, Zengyang Li, and Yutao Ma. 2025. Assessing and Analyzing the Correctness of GitHub Copilot’s Code Suggestions.ACM Transactions on Software Engineering and Methodology(2025)

  44. [44]

    Laurie Murphy, Gary Lewandowski, Renée McCauley, Beth Simon, Lynda Thomas, and Carol Zander. 2008. Debugging: the Good, the Bad, and the Quirky – A Qualitative Analysis of Novices’ Strategies. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  45. [45]

    Neuendorf

    Kimberly A. Neuendorf. 2002.The Content Analysis Guidebook. SAGE Publica- tions

  46. [46]

    Manh Hung Nguyen, Victor-Alexandru Padurean, Alkis Gotovos, Sebastian Tschi- atschek, and Adish Singla. 2025. Synthesizing High-Quality Programming Tasks with LLM-Based Expert and Student Agents. InProceedings of the Artificial Intel- ligence in Education (AIED)

  47. [47]

    Julian Oertel, Jil Klünder, and Regina Hebig. 2025. Don’t Settle for the First! How Many GitHub Copilot Solutions Should You Check?Information and Software Technology(2025)

  48. [48]

    Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo, and Benjamin Kok Siew Gan. 2025. Evaluating ChatGPT to Answer Multi-Modal Exercises in Computer Science Education. InProceedings of the Innovation and Technology in Computer Science Education Conference (ITiCSE)

  49. [49]

    Stack Overflow. 2024. Stack Overflow Developer Survey: AI section. https: //survey.stackoverflow.co/2024/ai

  50. [50]

    Victor-Alexandru Padurean, Paul Denny, Alkis Gotovos, and Adish Singla. 2025. Prompt Programming: A Platform for Dialogue-based Computational Problem Solving with Generative AI Models. InProceedings of the Innovation and Technol- ogy in Computer Science Education Conference (ITiCSE)

  51. [51]

    Victor-Alexandru Padurean, Paul Denny, and Adish Singla. 2025. BugSpotter: Au- tomated Generation of Code Debugging Exercises. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  52. [52]

    Passonneau

    Rebecca J. Passonneau. 2006. Measuring Agreement on Set-valued Items (MASI) for Semantic and Pragmatic Annotation. InProceedings of the Conference on Language Resources and Evaluation (LREC)

  53. [53]

    Tung Phung, José Cambronero, Sumit Gulwani, Tobias Kohn, Rupak Majumdar, Adish Singla, and Gustavo Soares. 2023. Generating High-Precision Feedback for Programming Syntax Errors using Large Language Models. InProceedings of the International Conference on Educational Data Mining (EDM)

  54. [54]

    Tung Phung, Victor-Alexandru Padurean, Anjali Singh, Christopher Brooks, José Cambronero, Sumit Gulwani, Adish Singla, and Gustavo Soares. 2024. Automating Human Tutor-Style Programming Feedback: Leveraging GPT-4 Tutor Model for Hint Generation and GPT-3.5 Student Model for Hint Validation. InProceedings of the International Learning Analytics and Knowled...

  55. [55]

    Chanathip Pornprasit and Chakkrit Tantithamthavorn. 2024. Fine-Tuning and Prompt Engineering for Large Language Models-Based Code Review Automation. Information and Software Technology(2024)

  56. [56]

    James Prather et al. 2023. The Robots Are Here: Navigating the Generative AI Revolution in Computing Education. InProceedings of the Working Group Reports on Innovation and Technology in Computer Science Education (ITiCSE-WGR)

  57. [57]

    James Prather et al. 2024. Beyond the Hype: A Comprehensive Review of Current Trends in Generative AI Research, Teaching Practices, and Tools. InProceedings of the Working Group Reports on Innovation and Technology in Computer Science Education (ITiCSE-WGR)

  58. [58]

    It’s Weird That it Knows What I Want

    James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett B. Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2024. "It’s Weird That it Knows What I Want": Usability and Interactions with Copilot for Novice Programmers.ACM Transactions on Computer-Human Interaction(2024)

  59. [59]

    Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S

    James Prather, Brent N. Reeves, Juho Leinonen, Stephen MacNeil, Arisoa S. Ran- drianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, and Ben Briggs. 2024. The Widening Gap: The Benefits and Harms of Generative AI for Novice Pro- grammers. InProceedings of the Conference on International Computing Education Research (ICER)

  60. [60]

    Christian Rahe and Walid Maalej. 2025. How Do Programming Students Use Generative AI?Proceedings of the ACM on Software Engineering(2025)

  61. [61]

    Reeves, James Prather, Paul Denny, Juho Leinonen, Stephen MacNeil, Andrew Luxton-Reilly, Sebastian Mateos Nicolajsen, and Claus Brabrand

    Brent N. Reeves, James Prather, Paul Denny, Juho Leinonen, Stephen MacNeil, Andrew Luxton-Reilly, Sebastian Mateos Nicolajsen, and Claus Brabrand. 2025. Prompts First, Precision Later: Reviving the Vision of Natural Language Program- ming for Computing Education. InProceedings of the Koli Calling International Conference on Computing Education Research (K...

  62. [62]

    Brian J. Reiser. 2004. Scaffolding Complex Learning: The Mechanisms of Struc- turing and Problematizing Student Work.Journal of the Learning Sciences(2004)

  63. [63]

    Jake Renzella, Alexandra Vassar, Lorenzo Lee Solano, and Andrew Taylor. 2025. Compiler-Integrated, Conversational AI for Debugging CS1 Programs. InPro- ceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  64. [64]

    Gordon, Carina Negreanu, Christian Poelitz, Sruti Srini- vasa Ragavan, and Ben Zorn

    Advait Sarkar, Andrew D. Gordon, Carina Negreanu, Christian Poelitz, Sruti Srini- vasa Ragavan, and Ben Zorn. 2022. What is it like to program with artificial intelligence?. InProceedings of the Conference of the Psychology of Programming Interest Group (PPIG)

  65. [65]

    Sami Sarsa, Paul Denny, Arto Hellas, and Juho Leinonen. 2022. Automatic Gen- eration of Programming Exercises and Code Explanations Using Large Language Models. InProceedings of the Conference on International Computing Education Research (ICER)

  66. [66]

    Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, and David S

    Jaakko J. Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, and David S. Doermann. 2024. Future of Software Development with Generative AI.Automated Software Engineering(2024)

  67. [67]

    Andreas Scholl and Natalie Kiesler. 2025. SCRIPT - Supportive Chatbot for Resolving Introductory Programming Tasks. InProceedings of the Innovation and Technology in Computer Science Education Conference (ITiCSE)

  68. [68]

    Valerie J Shute. 2008. Focus on Formative Feedback.Review of Educational Research(2008)

  69. [69]

    Debra Steele-Johnson and Zachary T Kalinoski. 2014. Error Framing Effects on Performance: Cognitive, Motivational, and Affective Pathways.The Journal of Psychology(2014)

  70. [70]

    Steve Stemler. 2000. An Overview of Content Analysis.Practical Assessment, Research, and Evaluation(2000)

  71. [71]

    John Sweller. 1988. Cognitive Load During Problem Solving: Effects on Learning. Cognitive science(1988)

  72. [72]

    Desmarais, and Giuliano Antoniol

    Florian Tambon, Arghavan Moradi-Dakhel, Amin Nikanjam, Foutse Khomh, Michel C. Desmarais, and Giuliano Antoniol. 2025. Bugs in Large Language When AI Is Wrong on Purpose: How Students Respond to Buggy GenAI Code ICER 2026 Vol. 1, August 11–14, 2026, Uppsala, Sweden Models Generated Code: An Empirical Study.Empirical Software Engineering (2025)

  73. [73]

    Cordeiro

    Norbert Tihanyi, Tamás Bisztray, Mohamed Amine Ferrag, Ridhi Jain, and Lucas C. Cordeiro. 2025. How Secure Is AI-Generated Code? A Large-Scale Comparison of Large Language Models.Empirical Software Engineering(2025)

  74. [74]

    Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter

    Annapurna Vadaparty, Daniel Zingaro, David H. Smith IV, Mounika Padala, Christine Alvarado, Jamie Gorson Benario, and Leo Porter. 2024. CS1-LLM: Integrating LLMs into CS1 Instruction. InProceedings of the Innovation and Technology in Computer Science Education Conference (ITiCSE)

  75. [75]

    Glassman

    Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. InProceedings of the Conference on Human Factors in Computing Systems (CHI)

  76. [76]

    Olga Viberg, Jacqueline Wong, Yael Feldman-Maggor, Nora Dunder, and Car- rie Demmans Epp. 2025. Chatting with Code: Exploring LLMs as Learning Partners in Programming Education. InProceedings of the Artificial Intelligence in Education (AIED)

  77. [77]

    Mitchell, and Chris Piech

    Sierra Wang, John C. Mitchell, and Chris Piech. 2024. A Large Scale RCT on Effective Error Messages in CS1. InProceedings of the Technical Symposium on Computer Science Education (SIGCSE)

  78. [78]

    Winne and Allyson F

    Philip H. Winne and Allyson F. Hadwin. 1998. Studying as Self-Regulated Learn- ing. InMetacognition in Educational Theory and Practice. Lawrence Erlbaum Associates, 277–304

  79. [79]

    Stephanie Yang, Miles Baird, Eleanor O’Rourke, Karen Brennan, and Bertrand Schneider. 2024. Decoding Debugging Instruction: A Systematic Literature Re- view of Debugging Interventions.ACM Transactions on Computing Education (2024)

  80. [80]

    Stephanie Yang, Hanzhang Zhao, Yudian Xu, Karen Brennan, and Bertrand Schnei- der. 2024. Debugging with an AI Tutor: Investigating Novice Help-seeking Be- haviors and Perceived Learning. InProceedings of the Conference on International Computing Education Research (ICER)

Showing first 80 references.