{"id":"e10f6aad-2b3a-4461-b5b6-bac78af05a0b","arxiv_id":"2606.29754","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Proposes treating LLMs as stochastic systems for student experimentation in statistics courses using Veridical Data Science and PCS principles, with four example activities across educational levels.","lead":"The paper argues that statistics curricula should treat LLMs as objects of inquiry where students design experiments to examine output variability, bias, and prompt sensitivity. A smart generalist might read it to understand one proposed way to build critical thinking about AI tools into data science education.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's identification of the skill-transfer assumption as weakest matches the only plausible point of vulnerability for a normative curricular proposal. Because the manuscript supplies no data or formal argument that would require that assumption to be true in a stronger sense, the UNVERDICTED classification remains appropriate; the concrete_test above is simply a natural next verification step rather than a refutation of any existing claim.","tokens_in":1603,"tokens_out":327,"duration_ms":26678,"concrete_test":"Implement one of the four proposed activities (e.g., the introductory 'ask it twice' exercise) in a single statistics class section, administer a brief pre/post quiz on core concepts of variability and prompt sensitivity, and compare against a control section using conventional examples; retention of correct statistical language after four weeks would provide direct evidence on whether the approach transfers without introducing confusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a curricular recommendation rather than an empirical assertion: that structured student experiments on LLM output distributions can serve as objects of inquiry to teach variability, bias, and stability. For a position/discussion piece this recommendation does not rest on a testable factual premise that can be falsified within the manuscript itself. The reader's weakest assumption correctly flags the risk that stochastic model behavior could confuse rather than clarify statistical concepts, yet that risk is already acknowledged in the proposal's framing and does not create an internal inconsistency or hidden assumption that would invalidate the suggestion to try the activities. No further load-bearing technical flaw is apparent from the described scope.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript is a discussion piece arguing that statistics curricula should treat large language models (LLMs) not only as tools but as objects of inquiry. Students would design small experiments to probe variability, bias, and prompt sensitivity in LLM outputs, organized around the Veridical Data Science framework and Predictability-Computability-Stability (PCS) principles. Four curricular examples are proposed, ranging from introductory 'ask it twice' activities to graduate-level PCS stability audits of LLM-based analysis workflows.","tokens_in":1688,"tokens_out":283,"duration_ms":28973,"significance":"If adopted, the proposal could help statistics education adapt to AI tools by reinforcing core concepts of uncertainty and reproducibility through direct engagement with stochastic systems. The VDS/PCS grounding provides a coherent structure for critical inquiry that aligns with existing statistical pedagogy and may foster transferable reasoning skills about model behavior.","major_comments":[],"minor_comments":[{"comment":"The abstract and introduction would benefit from a brief explicit list or table summarizing the four curricular examples and their intended educational levels to improve readability for readers scanning the proposal.","section":null},{"comment":"Section describing the graduate-level PCS stability audits should clarify how the stability metric is operationalized when applied to LLM workflows, as the current framing leaves the concrete implementation steps somewhat open.","section":null}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary and significance assessment of the manuscript, as well as the recommendation for minor revision. The referee's description accurately captures the core proposal of treating LLMs as stochastic systems for student experimentation within the Veridical Data Science and PCS framework.","responses":[],"tokens_in":1094,"tokens_out":73,"duration_ms":34190,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main suggestion is to have statistics students treat LLMs as objects of study by running small experiments on their output distributions to learn about variability, bias, and stability. It outlines activities at different levels based on the Veridical Data Science and PCS frameworks.\n\nThis builds sensibly on prior work by the author. The examples, like asking the model the same question twice in intro classes or auditing stability in advanced ones, give a clear way to integrate the idea without needing new theory. It correctly identifies that LLMs are stochastic and that probing them can illustrate statistical principles.\n\nThe soft spot is the complete lack of any trial or assessment. We have no information on whether these activities actually improve learning or if the unpredictability of LLMs creates more confusion than insight. The proposal flags this issue but leaves it unaddressed.\n\nThis piece is for statistics educators who are updating their courses to deal with AI tools. Someone looking for practical ideas on how to discuss LLMs in class might find the structure helpful.\n\nIt should go to peer review for a discussion or education-focused outlet. The recommendation is worth considering even without data, as long as reviewers understand it's a starting point for conversation.","headline":"A curricular proposal for probing LLMs in stats classes that makes sense on paper but lacks any supporting evidence.","tokens_in":2162,"tokens_out":311,"would_cite":false,"duration_ms":35444,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Statistics curricula should treat LLMs as objects of inquiry where students design experiments to analyze output variability and bias.","keywords":["large language models","statistics education","veridical data science","PCS principles","stochastic systems","prompt sensitivity","curriculum examples","data science education"],"falsifier":"A controlled comparison in which students who complete LLM probing activities show no measurable improvement in understanding variability, bias, or prompt effects compared with students taught the same concepts through conventional examples.","tokens_in":2480,"feed_emoji":"","tokens_out":588,"duration_ms":33662,"temperature":0.7,"pith_summary":"The paper argues that LLMs function as interactive stochastic systems whose behaviors require direct statistical investigation rather than passive use. Students learn core concepts by creating small experiments that measure how outputs vary, respond to different prompts, and exhibit bias, then analyzing those distributions. This method follows the Veridical Data Science framework and PCS principles to structure activities that scale from introductory repetition checks to graduate-level workflow audits. Four concrete curricular examples illustrate how the approach fits different education levels.","feed_headline":"Statistics students should probe LLMs as stochastic systems","feed_subtitle":"Small experiments on output variability and prompt effects teach core skills for analyzing bias and distributions.","key_machinery":"Veridical Data Science framework together with Predictability-Computability-Stability (PCS) principles, used to structure student experiments that treat LLM outputs as data for statistical analysis.","core_discovery":"Large language models are interactive stochastic systems whose most consequential behaviors remain only partially understood, therefore statistics curricula should treat them as objects of inquiry: students probe variability, bias, and prompt sensitivity by designing small experiments and analyzing distributions of outputs, organized through the Veridical Data Science framework and PCS principles across educational levels with four proposed curricular examples.","pith_inferences":["The method could extend to comparing output distributions across different LLMs to illustrate model selection principles.","Integration with traditional simulation exercises might help isolate what students learn specifically from real LLM stochasticity.","Departments adopting this approach may need to develop shared prompt libraries so experiments remain reproducible across classes."],"forward_implications":["Introductory students begin with activities such as asking an LLM the same question twice to observe response variability.","Graduate students conduct PCS stability audits on entire LLM-assisted analysis workflows.","Students at all levels practice treating sequences of LLM outputs as empirical distributions that require statistical summarization.","Curricula gain explicit attention to how prompt changes alter output distributions and introduce bias."],"fun_headline_variants":["Test LLMs as stochastic systems in statistics classes","Analyze LLM distributions via student-designed experiments","Probe LLM bias using PCS in educational workflows","Engage stats students in Veridical LLM output studies","Curricula examples for auditing LLM stability with PCS"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Student-designed experiments on LLM outputs will reliably build transferable statistical reasoning skills even though the models' randomness may create confusion or misleading intuitions.","fun_headline_variants_meta":{"raw":{"variants":["Test LLMs as stochastic systems in statistics classes","Analyze LLM distributions via student-designed experiments","Probe LLM bias using PCS in educational workflows","Engage stats students in Veridical LLM output studies","Curricula examples for auditing LLM stability with PCS"]},"model":"grok-4.3","cost_usd":0.005818,"raw_usage":{"total_tokens":2622,"prompt_tokens":536,"num_sources_used":0,"completion_tokens":68,"cost_in_usd_ticks":58178000,"prompt_tokens_details":{"text_tokens":536,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2018,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":536,"tokens_out":68,"duration_ms":29903,"temperature":1.0,"reasoning_tokens":2018,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T04:15:16.369336+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled comparison in which students who complete LLM probing activities show no measurable improvement in understanding variability, bias, or prompt effects compared with students taught the same concepts through conventional examples.","supporting_citations":[],"review_version":1}