Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A short, structured lab session can shift students' reported comfort and willingness to use generative AI without increasing self-reported use on graded work.

desk verdict Useful mixed-methods evaluation of a GenAI literacy lab, but the causal claim outruns the pre/post design and the reported 'large' effects rest on an inflated effect-size metric. read the letter →

arxiv 2505.00100 v2 pith:NQYWK5KT submitted 2025-04-30 cs.CY cs.AIcs.ET

classification cs.CYcs.AIcs.ET
keywords GenerativeAI(GenAI)coreskilldevelopmentpedagogicalframeworkAI-Labquantitativequalitativemixedmethodscomputingeducationstudentperceptionsself-reportedusage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper evaluates the AI-Lab, a short scaffolded intervention in which students learn prompting, critique generated outputs in class, and reflect on a homework assignment using generative AI. Across two semesters in three computer science courses and one first-year engineering course, paired pre/post surveys (831 perception responses, 826 usage responses) and six focus groups tracked how students' reported openness, comfort, frequency of use, and prompting strategies changed. The central claim is that even a short intervention can make students more comfortable and open to using generative AI for conceptual, debugging, and homework help, and can shift how they prompt and evaluate outputs, without increasing self-reported frequency of use on graded homework or projects. The authors argue this matters because it addresses the faculty concern that teaching AI use encourages academic dishonesty, while giving students a more deliberate, learning-oriented way to engage the tools.

What carries the argument

The AI-Lab framework is a four-stage pedagogical sequence: instructor preparation (selecting a topic likely to expose AI errors), a prelab with preparatory materials and a baseline survey, an in-class lab in which the instructor demonstrates GenAI output and students critique and correct it collaboratively, and a post-lab homework requiring documented attempts to steer the AI plus a follow-up survey. This staged structure is the mechanism that is claimed to convert naive experimentation into deliberate, critical use. The quantitative evaluation rests on paired Wilcoxon signed-rank tests with rank-biserial effect sizes, and the qualitative analysis on thematic coding of focus-group transcripts.

What would settle it

A reader could settle the causal claim by running the same pre/post surveys in the same courses with a randomly assigned no-intervention control group, or by comparing semesters in which the intervention was offered to those in which it was not while holding the syllabus-policy requirement constant; if control students show the same comfort and openness gains, the AI-Lab is not the active ingredient.

Watch

Extended reading notes

Core claim

The central discovery, stated as the authors would state it, is that the AI-Lab intervention changes the quality of students' engagement with generative AI more than the quantity. Students' self-reported openness to using GenAI for conceptual questions and homework help increased, and comfort increased for conceptual, debugging, and homework scenarios, with rank-biserial correlations in the 'large' range. Frequency of use for homework and programming projects did not change significantly, while frequency of use for debugging increased. In focus groups, students described moving from trial-and-error prompting to iterative, context-rich prompting, becoming skeptical of confidently stated but incorrect outputs, and articulating explicit boundaries about when not to use AI. The paper interprets this as evidence that structured scaffolding can promote mindful, reflective AI use rather than overreliance.

Load-bearing premise

The load-bearing premise is that the pre-to-post changes in students' self-reported attitudes are caused by the AI-Lab intervention itself, even though the same period included a university policy change on AI syllabus statements, growing student familiarity with GenAI, and no control group.

Editorial extensions

If this is right

  • If the central claim holds, computing educators can introduce GenAI tools through a short structured activity without measurable increases in self-reported use on graded assignments, addressing a common faculty worry about cheating.
  • Students' prompting strategies and skepticism of AI outputs are teachable within a single lab session, suggesting that the 'AI literacy' that matters is less about raw tool exposure and more about critique and reflection.
  • The finding that self-reported use for debugging increased while homework use stayed flat implies that scaffolding can channel GenAI use toward lower-stakes, skill-building activities.
  • The larger shift seen in the first-year engineering course, compared with CS courses, suggests the intervention may help students with less prior tool exposure more, motivating adoption outside CS majors.
  • Reported desire to use GenAI mostly shifted toward more use in CP, DSA-CS, and ENGR but toward less use in DSA-DSAI, hinting that effects depend on students' starting point.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An inference the authors leave implicit: because the study has no control group and the university began requiring AI syllabus statements between the two semesters, the reported shifts may partly reflect growing institutional and societal familiarity with GenAI rather than the AI-Lab alone; the paper's own semester comparison documents a large baseline difference.
  • The qualitative themes suggest a testable extension: behavioral trace data (e.g., actual prompt logs before and after the intervention) could determine whether 'iterative prompting' and 'skepticism' translate into measurable changes in tool interaction rather than only self-report.
  • A plausible next experiment would separate the intervention's components — the in-class critique, the homework reflection, and the prelab materials — to identify which stage drives the comfort and openness gains.
  • The semester difference also implies that the intervention's effects may be time-dependent, so replicating it now versus a year ago may yield different baselines; future evaluations should treat institutional AI policy as a covariate.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a mixed-methods evaluation of the AI-Lab, a scaffolded intervention for teaching undergraduate CS and engineering students to use generative AI deliberately. Across two semesters and three CS courses plus one engineering course, the authors collected paired pre/post surveys (N≈830) and six post-intervention focus groups. With Wilcoxon signed-rank tests they find statistically significant increases in several self-reported openness, comfort, and debugging-frequency items, while homework/project use frequency stays flat, and they interpret the corresponding rank-biserial correlations (0.88–0.94) as large effects. The qualitative themes describe students adopting more iterative prompting, greater skepticism of outputs, and clearer boundaries around integrity.

Significance. If the causal claim were supported, this would be a useful contribution to the emerging literature on GenAI pedagogy: it is one of the larger explicitly evaluated interventions, it pairs quantitative and qualitative data, and it explicitly documents a semester-level policy shift that confounds the design (Section 4.2). The honesty about the confound is a strength, and the large paired sample is a real asset. However, the central claim that a short intervention can shift attitudes is not established by the design, and the reported 'large effect sizes' are not supported by the raw distributions. The paper has value as a pilot or exploratory evaluation, but the current framing overstates what the evidence can show.

major comments (4)
  1. [§4.2, Fig. 5] The pre/post design has no non-intervention comparison group, and the paper itself documents a major confound: between Spring and Fall 2024, Purdue moved from recommending to requiring AI statements in course syllabi, and Figure 5 shows that self-reported baseline homework-use frequency differs sharply by semester (low-use categories: 28.49% in S24 vs 52.30% in F24). The combined-semester Wilcoxon tests therefore cannot separate the intervention effect from concurrent policy change, maturation, or growing familiarity with GenAI over the semester. The causal language in the abstract ('can shift') and in Section 4.5 ('can promote') is not supported by the evidence. To keep the causal claim, the authors should provide a semester-specific analysis with an explicit modeling of the policy change, or reframe the claim as 'students reported shifts' and openly state that the design cannot rule out confounders.
  2. [§2.2.2, §3.2.1, Tables 3–8, Table 13] The claim of 'large effect sizes' is not supported by the raw distributions. For instance, openness to conceptual questions moves from 50.6% to 53.2% at the top category (Table 3), and comfort with conceptual questions moves from 35.3% to 38.0% at the top category (Table 6). With several hundred students, nearly all responses are unchanged, so a rank-biserial correlation of 0.90–0.94 is driven by the few discordant pairs that almost all go in one direction; it is not evidence that most students changed substantially. The authors should report (i) the proportion of students who increased, decreased, and stayed the same for each item, and (ii) a more interpretable effect size such as Cliff's delta with a confidence interval, and should temper the 'substantial and meaningful changes' language in Section 3.2.1.
  3. [Table 13 and list in §3.2] Table 13 reports the openness-to-debugging item with p=0.6592 and no effect size, yet the list of statistically significant questions in Section 3.2 includes 'How open are you to using GenAI to get help with debugging?' as item (2). This is a direct contradiction that must be resolved; it appears that the significance list is wrong and the table is correct, but as written the paper's own results are internally inconsistent.
  4. [§3.3, §4.4] The qualitative focus-group themes are retrospective self-reports collected by the same researchers who implemented the intervention, and the statement in §4.4 that 'These reflections were not influenced by pressure from the intervention or researchers' is unsupported. Demand characteristics are a plausible source of the observed reports (e.g., students echoing the intervention's stated goals of 'mindful usage'). The focus groups cannot provide a counterfactual, so they do not rescue the causal interpretation. This limitation should be acknowledged explicitly.
minor comments (5)
  1. [§1.2] RQ2 is phrased as a yes/no question ('Are Students Open to Using GenAI...?'), but the study analyzes shifts in Likert-scale openness/comfort. The research question should be aligned with the outcome measured.
  2. [§2.2.2] The effect-size interpretation thresholds (0.30–0.50 medium, >0.50 large) are presented without citation; give a reference or justify the cutoffs.
  3. [§2.3] The number of focus-group participants and the total number of groups (six) are mentioned, but the number of participants per group is not reported; include this for transparency.
  4. [§3.4] The discussion of peer-usage perceptions is confusing: the text says 'only around 50% of students predicted correctly' but then reports that 84.54% believed ≥50% of peers used GenAI, while 75.42% of students themselves used it on occasion or more. Please clarify the claim and the arithmetic.
  5. [§4.1, Table 14] The statement that ENGR showed 'a higher intervention impact' is based on the count of significant p-values; this is not a valid comparison of effect size, especially with a small (n=53) sample. Use an effect-size comparison or drop the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the empirical evaluation uses new pre/post data and focus groups; self-citations are contextual and not load-bearing.

full rationale

This paper is an empirical mixed-methods evaluation rather than a derivation, so most circularity patterns do not apply. No parameter is fitted and then renamed as a prediction; no equation is derived from an assumption equivalent to the conclusion. The outcomes (openness, comfort, frequency) are direct Likert self-reports analyzed with paired Wilcoxon signed-rank tests, and the analysis does not construct those outcomes from the intervention definition. The AI-Lab framework is attributed to the authors' prior work [8,3], but the current paper explicitly frames itself as providing the first rigorous empirical evaluation of the framework, and its findings rest on newly collected paired surveys and focus groups rather than on the cited prior papers. The self-citations are definitional and contextual, describing what the intervention is and why it was developed; they do not import a uniqueness theorem, forbid competing explanations, or substitute for the new data. The paper's own Section 4.2 identifies a real confound: Purdue moved from recommending to requiring AI syllabus statements between semesters, and baseline homework-use frequency differs sharply by semester. That is a threat to causal inference, not a circular reduction; a confounded empirical claim is still an empirical claim whose inputs are independent of the outputs. Similarly, the retrospective item asking students how their desire to use GenAI changed because of the intervention (Figure 4) raises demand-characteristics concerns, but it is an outcome measure rather than a fitted input. Therefore no specific circular step can be quoted, and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No fitted parameters or invented entities; this is an empirical education study. The key assumptions are the validity of self-reports, implementation fidelity, and the pre/post matching integrity.

assumptions (3)
  • domain assumption Self-reported survey responses accurately reflect students' actual GenAI usage and attitudes.
    The central finding rests entirely on pre/post self-reports; the paper itself acknowledges the need to triangulate with behavioral traces.
  • domain assumption The AI-Lab intervention was implemented with fidelity across courses and semesters.
    Data are pooled across four courses and two semesters with no fidelity checks or instructor-level implementation reports.
  • domain assumption There is no differential attrition or matching error between pre and post surveys.
    Only students with both surveys are included, but the matching mechanism is not described.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation." pith.science (2026). https://pith.science/paper/NQYWK5KT

@misc{pith2026250500100,
  author       = {Pith},
  title        = {Pith review of: A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NQYWK5KT}},
  note         = {Machine review of arXiv:2505.00100}
}
read the original abstract

Background and Context. Generative AI (GenAI) tools are increasingly used in programming courses, but we have limited evidence about how brief instruction can foster responsible, learning-oriented use. Objectives. We evaluate "AI-Lab", a scaffolded GenAI literacy intervention, asking how students' self-reported GenAI usage and their openness and comfort using GenAI for conceptual, debugging, and homework tasks change after participation. Methods. Across two semesters in three CS courses and one first-year engineering course at a U.S. university, we deployed the "AI-Lab" (pre-lab orientation, in-class critique of GenAI outputs, and a required homework reflection), collecting paired pre/post surveys (Perception N=831; Usage N=826) and six post-intervention focus groups; primary inferential analyses used the three CS courses (N=778 and 773, respectively). We analyzed survey shifts with paired non-parametric tests and focus groups via thematic analysis. Findings. Openness increased for conceptual questions and homework help, and comfort increased for conceptual, debugging, and homework scenarios; self-reported frequency of GenAI use for homework and projects remained stable, while self-reported use for debugging increased. Focus group participants described adopting more iterative prompting strategies, becoming more skeptical of correctness, and articulating clearer boundaries around integrity and dependence. Implications. A short, structured intervention can shift students' reported comfort with and willingness to use GenAI and influence the strategies they describe for engaging with it without increasing overall self-reported use on graded work. These results motivate future work triangulating surveys with behavioral traces and learning measures.

Figures

Figures reproduced from arXiv: 2505.00100 by the authors.

Figure 1
Figure 1. AI-Lab Framework for integrating GenAI in programming courses [3]. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Openness to GenAI being allowed in a collegiate class, for CS and ENGR students, from the post-perception survey. There was [PITH_FULL_IMAGE:figures/full_fig_p013_2.png] view at source ↗
Figure 3
Figure 3. “No Idea” responses excluded. They compose the following amounts: Combined: 121/826 (14.65%), CS only: 111/773 (14.36%), [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Self-reported change in desire to use GenAI [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: A stark difference between Fall 2024 and Spring 2024, in percentage of students who used GenAI moderately or very frequently [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Failure of Plagiarism Detection in Competitive Programming

    cs.CY 2025-05 conditional novelty 3.0 of 10

    Code similarity detectors miss obfuscated or AI-generated submissions in competitive programming, so the author recommends combining automated screening, manual review, and oral interviews.

Reference graph

Works this paper leans on

21 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [1]

    Generative Ai. 2025. Developing Your Approach to Generative AI. https://www.scholarlyteacher.com/post/developing-your-approach-to- generative-ai

  2. [2]

    Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos

    Brett A. Becker, Paul Denny, James Finnie-Ansley, Andrew Luxton-Reilly, James Prather, and Eddie Antonio Santos. 2023. Programming Is Hard - Or at Least It Used to Be: Educational Opportunities and Challenges of AI Code Generation. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V. 1 (Toronto ON, Canada) (SIGCSE 2023). Ass...

  3. [3]

    Andres Bejarano, Ethan Dickey, and Rhianna Setsma. 2025. Implementing the AI-Lab Framework: Enhancing Introductory Programming Education for CS Majors. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2 (Pittsburgh, PA, USA) (SIGCSETS 2025). Association for Computing Machinery, New York, NY, USA, 1383–1384. https://doi.o...

  4. [4]

    Virginia Braun and Victoria Clarke and. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101. https://doi.org/10.1191/1478088706qp063oa arXiv:https://www.tandfonline.com/doi/pdf/10.1191/1478088706qp063oa

  5. [5]

    Christopher Bull and Ahmed Kharrufa. 2024. Generative Artificial Intelligence Assistants in Software Development Education: A Vision for Integrating Generative Artificial Intelligence Into Educational Practice, Not Instinctively Defending Against It. IEEE Software 41, 2 (2024), 52–59. https://doi.org/10.1109/MS.2023.3300574

  6. [6]

    Curry, Ingrid M

    Leslie A. Curry, Ingrid M. Nembhard, and Elizabeth H. Bradley. 2009. Qualitative and Mixed Methods Provide Unique Con- tributions to Outcomes Research. Circulation 119, 10 (2009), 1442–1452. https://doi.org/10.1161/CIRCULATIONAHA.107.742775 arXiv:https://www.ahajournals.org/doi/pdf/10.1161/CIRCULATIONAHA.107.742775

  7. [7]

    Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N

    Paul Denny, James Prather, Brett A. Becker, James Finnie-Ansley, Arto Hellas, Juho Leinonen, Andrew Luxton-Reilly, Brent N. Reeves, Eddie Antonio Santos, and Sami Sarsa. 2024. Computing Education in the Era of Generative AI. Commun. ACM 67, 2 (Jan. 2024), 56–67. https://doi.org/10.1145/ 3624720

  8. [8]

    Ethan Dickey, Andres Bejarano, and Chirayu Garg. 2024. AI-Lab: A Framework for Introducing Generative Artificial Intelligence Tools in Computer Programming Courses. SN Computer Science 5, 6 (2024), 720. https://doi.org/10.1007/s42979-024-03074-y

Show all 21 references
  1. [9]

    José Antonio Donaire, Mònica Puntí Brun, Konstantina Zerva, Raquel Camprubí Subirana, and Núria Galí Espelt. 2025. De la tiza al chip: el uso de la inteligencia artificial en las aulas

  2. [10]

    Becker, Andrew Luxton-Reilly, and James Prather

    James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather. 2022. The Robots Are Coming: Exploring the Implications of OpenAI Codex on Introductory Programming. In Proceedings of the 24th Australasian Computing Education Conference (Virtual Event...

  3. [11]

    Henrique Freitas, Mírian Oliveira, Milton Jenkins, and Oveta Popjoy. 1998. The Focus Group, a qualitative research method. Journal of Education 1, 1 (1998), 1–22

  4. [12]

    Alan Harrison. 2025. What role should Generative AI play in education? - Teach Computing. https://teachcomputing.org/blog/ai-in-education

  5. [14]

    Smith, Juho Leinonen, Stephen MacNeil, Andrew Luxton-Reilly, and Brett A

    Chris Kerslake, Paul Denny, David H. Smith, Juho Leinonen, Stephen MacNeil, Andrew Luxton-Reilly, and Brett A. Becker. 2025. Exploring Student Reactions to LLM-Generated Feedback on Explain in Plain English Problems. In Proceedings of the 56th ACM Technical Symposium on Comput...

  6. [15]

    Ban It Till We Understand It

    Sam Lau and Philip J. Guo. 2023. From "Ban It Till We Understand It" to "Resistance is Futile": How University Programming Instructors Plan to Adapt as More Students Use AI Code Generation and Explanation Tools such as ChatGPT and GitHub Copilot. In Proceedings of the 2023 ACM...

  7. [16]

    Evanfiya Logacheva, Arto Hellas, James Prather, Sami Sarsa, and Juho Leinonen. 2024. Evaluating Contextually Personalized Programming Exercises Created with Generative AI. In Proceedings of the 2024 ACM Conference on International Computing Education Research - Volume 1 (Melbo...

  8. [17]

    Desmarais, and Zhen Ming (Jack) Jiang

    Arghavan Moradi Dakhel, Vahid Majdinasab, Amin Nikanjam, Foutse Khomh, Michel C. Desmarais, and Zhen Ming (Jack) Jiang. 2023. GitHub Copilot AI pair programmer: Asset or Liability? Journal of Systems and Software 203 (2023), 111734. https://doi.org/10.1016/j.jss.2023.111734

  9. [18]

    Ipek Ozkaya and Douglas Schmidt. 2024. Generative AI and Software Engineering Education. Carnegie Mellon University, Software Engineering Institute’s Insights (blog). https://doi.org/10.58012/AHC6-GK69 Accessed: 2025-Mar-19

  10. [19]

    Santos, and Marcos Zampieri

    Nishat Raihan, Mohammed Latif Siddiq, Joanna C.S. Santos, and Marcos Zampieri. 2025. Large Language Models in Computer Science Education: A Systematic Literature Review. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1 (Pittsburgh, PA, USA)...

  11. [20]

    Johnny Saldaña. 2013. The coding manual for qualitative researchers (2 ed.). SAGE Publications

  12. [21]

    Sue Sentance and Jane Waite. 2022. Perspectives on AI and data science education. InUnderstanding Computing Education (Vol 3): AI, data science, and young people, ser. Proceedings of the Raspberry Pi Foundation Research Seminars . University of Cambridge and Raspberry Pi Found...

  13. [22]

    Ahmed Tlili. 2024. Can artificial intelligence (AI) help in computer science education? A meta-analysis approach. Revista Española de Pedagogía 82, 289 (2024), 10. Manuscript submitted to ACM

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.