{"id":"25f7ef6d-fc3e-4f19-bff8-73e3c33ff9ca","arxiv_id":"2607.15518","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Computational physics courses should be re-centered on verification and oral defense because AI-generated artifacts no longer certify student understanding.","lead":"A physics educator argues that when AI can write and run simulation code, the skill students most need is verifying and explaining results, not writing code. He proposes a teaching and testing plan built around AI-free coding quizzes and oral defenses of AI-assisted projects.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"White-box dose is the load-bearing unvalidated premise: if one Euler encounter does not build verification skill, the oral-defense gate certifies ritual, not understanding, and the assessment consequence is unsupported.","rationale":"The reader's weakest-assumption analysis correctly identifies the white-box dose as the load-bearing empirical premise. The paper is a well-argued pedagogical framework, not an empirical study, and its central claims are presented as testable propositions rather than established results. The author is commendably explicit about the lack of data: Sec. VIII Q1 names the dose and the verification-curriculum calibration as 'the framework's central untested assumptions,' and Sec. VII acknowledges that the pilot-skills decay literature points in the opposite direction. My independent reading confirms this is the weakest link, because both the verification-centric curriculum and the oral-defense assessment design presuppose that a minimal white-box encounter produces durable conceptual understanding. If that premise fails, the framework's strongest claim—that the artifact no longer certifies the student—remains true as a diagnosis, but the proposed remedy (verification training plus oral defense) would not restore trustworthy certification; it would merely move the performance problem from artifact production to defense ritual. The paper's own caveats and the lack of empirical evidence justify the CONDITIONAL verdict, not a stronger rejection. There are secondary concerns (rubric validation, inter-rater reliability, scalability) that the paper acknowledges and defers to future data, but none is as fundamental as the white-box dose. I therefore agree with the reader's assessment and recommend no change to the verdict.","tokens_in":14351,"tokens_out":4450,"duration_ms":48922,"concrete_test":"Run a pre-registered comparison in the Fall 2026 full-protocol cohort: randomly assign half of the students to receive the white-box Euler phase (hand-coded integrator plus AI-free quiz) and half to skip it, holding all other instruction constant. At the end of the semester, measure both groups' diagnostic accuracy on error-injected black-box simulations (Sec. V B) and their pass rates on the verification-gate dimension of the oral-defense rubric (Sec. VI C). If the white-box group shows no significant advantage on either measure, the framework's central pedagogical assumption fails and the assessment claims must be downshifted to proposal status.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The framework's central claim—that verification, not code authorship, is the load-bearing skill and that the oral defense can certify it—depends on an unvalidated pedagogical premise: a single brief white-box phase per method class (Sec. V A: 'students build the fifteen-line Euler integrator, watch it fail at large dt, and must implement it on an AI-free quiz') is sufficient to leave the 'conceptual residue' needed to later verify agent-generated simulations. The paper itself calls this 'the framework's central untested assumption' (Sec. VIII, Q1) and concedes the pilot analogy cuts against it: manual flying skills decayed under training doses far larger than a brief white-box phase (Sec. VII, citing [56]). If one encounter is too few, students will not internalize the failure modes and mental models required to catch agent errors; the verification repertoire of Sec. IV will be performed as ritual—run because the rubric demands it—without the sensemaking that gives it force. Then the oral defense's verification gate (Sec. VI C) would certify rehearsed behavior rather than understanding, undermining the paper's strongest claim about what assessment can measure. The author's own one-week workshop experience (Sec. V C) involved an expert with pre-existing deep conceptual pillars; it does not generalize to novices. This is not an internal inconsistency—the author honestly labels it a testable bet—but it is the load-bearing link in the argument, and no data yet support it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a tool-invariant framework for teaching and assessing computational methods in physics, organized around five pillars (inputs/outputs, conceptual understanding, terminology, sensemaking, and tool operation). It argues that as agentic AI assumes code-writing and execution, verification—not code authorship—becomes the load-bearing skill, and that unsupervised submitted artifacts no longer certify student understanding. The author proposes a two-instrument assessment model: AI-free in-class coding quizzes for the white-box phase, and comment-stripped oral defenses with a verification-gated rubric for AI-assisted project work. The paper explicitly positions itself as a framework and a structure for adaptation, honestly labels its central assumptions as untested, and lists open questions, most importantly whether a brief white-box phase suffices to build the conceptual residue needed for later verification of agent-generated simulations.","tokens_in":14612,"tokens_out":6292,"duration_ms":71522,"significance":"If the framework holds, it provides a timely and actionable response to the collapse of unsupervised computational assessment, and it reframes the computational physics curriculum around validation authority and disciplinary discourse rather than artifact production. The paper is strong in its conceptual organization, its synthesis of the computational-thinking and assessment literatures, and its explicit, falsifiable open questions. It is also unusually transparent: it names its central untested assumptions, reports the anecdotal status of its evidence, and discloses its own AI-assisted authorship. The main risk is that the practical recommendations will be read as validated when they are, by the author's own account, a testable design awaiting cohort data.","major_comments":[{"comment":"The white-box dose is the load-bearing empirical premise: the claim that a single hand-coded Euler encounter plus weekly quizzes leaves enough conceptual residue to verify agent-generated simulations is untested. The author's citation [56] shows manual-flying skills decay under training doses far larger than this phase; the reply that weekly quizzes re-exercise the residue is plausible but not evidence, since the quizzes assess coding fluency and injected-error diagnosis, not transfer to bespoke agent artifacts. Because the oral-defense verification gate depends on this transfer, the manuscript should either present preliminary pre/post evidence or explicitly mark the practical recommendations as conditional on a hypothesis that the first cohorts will test.","section":"Secs. V A, VII, VIII Q1"},{"comment":"The evidence base for the defense protocol is one semester in two upper-level courses (n=8), without comment-stripping and without the formal rubric. Despite this, the text states 'ten minutes suffices' and 'genuine and performed command separate within minutes' as design facts. These are anecdotal judgments, not measurements. Because the verification gate is the core assessment innovation, the protocol's validity requires at least inter-rater reliability, gate-failure rates, and differential-impact data (the author lists these in Sec. VIII Q1 but does not provisionally answer them). Please separate 'experience so far' from 'validated practice' and soften the prescriptive language accordingly.","section":"Secs. VI B and VI C"},{"comment":"The framework claims that error-injection exercises and provided-code quizzes train the code-reading fluency that the defense walkthrough demands, and that verification is 'the teachable core.' Yet open question 2 concedes that how much code reading directing agents requires is unresolved. The assertion in Sec. V B that the exercises 'provide the supervisory practice that the agentic workflow no longer provides by itself' is therefore stronger than the evidence. A revision should state the assumed mechanism linking error-injection practice to real-agent verification and specify the measurement that would confirm or refute it.","section":"Secs. V B, VI B, VIII Q2"}],"minor_comments":[{"comment":"The headings contain typographical artifacts: 'COMPUT A TION HAS AL W A YS BEEN DELEGA TED' and 'V erify'. These should be corrected to 'COMPUTATION HAS ALWAYS BEEN DELEGATED' and 'Verify'.","section":"Secs. II and III B"},{"comment":"The phrase 'the deeper reply is constructive alignment itself' is awkward; 'deeper point' would be clearer. As written, it is unclear whether 'reply' refers to the answer to the adversarial objection or to the assessment mechanism.","section":"Sec. VI C"},{"comment":"The final row 'What is new' has no entries in the antecedent columns; clarify whether this is a merged row or a deliberate empty row. The supplementary tables S1 and S2 are referenced but not included in the main text; state where they are available.","section":"Table I"},{"comment":"The statement 'the ten-minute session length is measured—a semester of ... stands behind it' overstates what a semester of informal experience provides. Use 'estimated from experience' rather than 'measured'.","section":"Sec. VI B"},{"comment":"The sentence 'an end-of-semester survey drew exactly one response, itself a small lesson in instrumentation' is effective and honest, but the subsequent 'response split sharply' should be flagged even more clearly as one instructor's impression, since the only systematic instrument produced n=1.","section":"Sec. V D"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest about its own limitations, which is to its credit. My main worry is that the abstract and conclusion will be read as validated recommendations by instructors who skim. The revision should make the conditional status of the white-box dose and the defense protocol impossible to miss. I do not see grounds for rejection: the framework fills a genuine gap and the open questions are correctly specified. My choice of major_revision rather than minor_revision reflects the fact that the load-bearing assumption is not merely underemphasized but explicitly untested, and the prescriptive language needs to be brought into alignment with the evidence status."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Leading take: this is a well-argued, honest framework paper. The central claim—under agentic AI, verification, not code authorship, becomes the load-bearing skill, and unsupervised artifacts no longer certify students—is coherent and grounded in the literature. The author is explicit about the limits: Sec. V A and Sec. VIII Q1 name the white-box dose and verification calibration as the central untested assumptions; Sec. VI B reports only one semester of informal defenses without comment-stripping or a formal rubric. There is no internal contradiction and no circular argument.\n\nWhat is actually new: the validation-authority principle (the bespoke-versus-socially-validated distinction), the tool-invariant re-weighting of the five pillars, and the verification-gated oral defense with comment-stripping. The five pillars themselves are honestly repackaged from computational-thinking frameworks, with Table I giving due credit. The verification repertoire in Sec. IV is a useful classroom operationalization, and the paper handles the equity objection without hand-waving.\n\nThe soft spot is the load-bearing premise: a single white-box encounter per method class is assumed to leave enough conceptual residue for later verification of agent-generated work. The author himself concedes the pilot-manual-skills analogy cuts against this, and the first-person workshop experience involved an expert with already-deep pillars, not a novice. If one encounter is too few, the oral defense risks certifying ritual rather than understanding. That said, the author labels this a testable bet, and the proposed Fall 2026 data will address it. The assessment claims are also stronger than the current evidence: ten-minute defenses separating genuine from performed command is based on one instructor's small-class experience. These are limitations, but they are named limitations, and for a framework/proposal paper they are acceptable.\n\nThe citation pattern looks solid. The author's PICUP self-citations are relevant to his role as PI, and he cites opposing evidence (deskilling, constructionism, Artigue) honestly. The AI-use disclosure at the end is transparent and consistent with the paper's own principles.\n\nThis paper is for anyone designing computational physics courses or assessments in the AI era, and for PER researchers studying the measurement problem. It deserves a serious referee; my recommendation is to engage with it, and to pressure-test the white-box dose and rubric calibration during review.","headline":"Solid, honest framework paper that makes a credible case for verification-centered teaching and oral-defense assessment under agentic AI; the load-bearing white-box dose is untested but explicitly flagged as a bet.","tokens_in":15119,"tokens_out":2982,"would_cite":true,"duration_ms":33677,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["01.40.Ha"],"model":"deepseek-v4-flash","headline":"What students must know stays constant across tool changes; only the weight on verification moves, so assessment must move from submissions to live defense.","keywords":["tool-invariant framework","agentic AI","computational physics education","verification","oral assessment","white-box/black-box","sensemaking","assessment collapse"],"falsifier":"A controlled experiment comparing two cohorts on the same simulation task—one that completes the hand-coded Euler white-box exercise, one that jumps straight to agent-generated code—would settle the central claim: if both groups detect injected errors at equal rates, the framework's core assumption fails.","tokens_in":14161,"feed_emoji":"🎓","tokens_out":3531,"duration_ms":38093,"temperature":0.7,"pith_summary":"This paper argues that what students must learn to use computational methods has stayed constant across tool transitions: inputs and outputs, conceptual understanding, terminology, sensemaking, and tool operation. Agentic AI is the latest and largest offloading yet, shifting the weight almost entirely onto verification, because AI-generated simulations are bespoke artifacts that no community has validated. The traditional homework artifact—code, plots, report—no longer certifies the student, since AI can produce it on demand. The paper's response is a two-instrument assessment: AI-free in-class coding quizzes for a brief white-box phase, paired with ten-minute oral defenses of comment-stripped, AI-assisted work, gated on a verification dimension. The product of a computational physics course becomes the student's ability to explain and defend computational artifacts in the language of physics.","feed_headline":"Verification, not coding, is the new core skill","feed_subtitle":"When AI generates artifacts on demand, only supervised assessment—AI-free quizzes and live defenses—can certify students.","key_machinery":"The five-pillar framework: competent use of any computational method requires (1) specifying inputs and outputs, (2) conceptual understanding of the method, (3) precise terminology, (4) sensemaking to judge whether outputs are right, and (5) operating the current tool. The argument runs through a Specify–Predict–Delegate–Verify–Interpret loop and the 'principle of validation authority'—delegation is safe only where the delegator retains the ability to validate outputs. That principle distinguishes opaque but socially validated library code from opaque, bespoke, agent-generated simulations, and it is what makes verification newly load-bearing.","core_discovery":"The paper claims that agentic AI is the latest step in a centuries-long migration of mechanical work from human to tool, and that what students must know has not changed: inputs and outputs, conceptual understanding, terminology, sensemaking, and tool operation. What changes is weight: with AI writing code, sensemaking and verification become load-bearing, and the unsupervised artifact no longer certifies the student. Because each AI-generated simulation is a bespoke artifact with a population of one, the verification burden that library ecosystems once amortized across the community now lands on each student, for every artifact, every time. The constructive consequence is a two-instrument a","pith_inferences":["The 'proxy collapse' generalizes beyond physics: any discipline that relies on artifacts—code, essays, designs—faces the same assessment problem once AI can generate those artifacts on demand.","The validation-authority principle suggests that as AI improves, the verification burden may not shrink as much as expected, because bespoke artifacts never acquire the community validation that libraries enjoy.","A concrete testable extension is whether prediction-first exercises measurably improve detection of injected errors; if not, the verification curriculum risks becoming ritual rather than genuine calibration.","The framework implies class size becomes an academic-integrity variable: institutions that cannot staff oral defenses may issue systematically less trustworthy credentials, an equity concern the paper makes explicit."],"forward_implications":["If the framework holds, computational physics courses will shift from teaching code authorship to teaching specification, prediction, and verification as the core skills.","Unsupervised computational assignments can no longer certify student understanding; courses must move to supervised formats such as proctored quizzes and oral defenses.","A brief white-box phase—one hand-coded Euler integrator per method—is enough to leave the conceptual residue needed for later verification and diagnosis of agent errors.","A ten-minute oral defense with a verification gate can reliably distinguish genuine understanding from performed command, and is feasible in the small classes where the subject lives.","AI-free coding quizzes remain the only valid measure of coding fluency, which the paper regards as a permanent but small learning goal."],"fun_headline_variants":["AI writes code, students verify","Verification replaces coding as core skill","The artifact doesn't certify; the defense does","Stop teaching code, start teaching verification","From coding to verifying in the age of AI"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"A single brief hand-coded encounter with each method leaves enough conceptual residue for students to verify agent-generated simulations later","fun_headline_variants_meta":{"raw":{"variants":["AI writes code, students verify","Verification replaces coding as core skill","The artifact doesn't certify; the defense does","Stop teaching code, start teaching verification","From coding to verifying in the age of AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000328,"raw_usage":{"total_tokens":1665,"prompt_tokens":734,"completion_tokens":931,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":867}},"tokens_in":478,"tokens_out":931,"duration_ms":8472,"temperature":1.0,"reasoning_tokens":867,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T23:04:42.145514+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment comparing two cohorts on the same simulation task—one that completes the hand-coded Euler white-box exercise, one that jumps straight to agent-generated code—would settle the central claim: if both groups detect injected errors at equal rates, the framework's core assumption fails.","supporting_citations":[],"review_version":1}