{"id":"50509f01-2b99-4fe0-8d7b-c0d90b421a20","arxiv_id":"2509.01914","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI-simulated tutoring dialogues are more explanation-driven and less questioning than human ones, producing a simple 'explain-acknowledge' loop instead of a 'question-answer-feedback' loop.","lead":"This study compares 49 real one-on-one math tutoring dialogues with AI-simulated versions of the same tutoring sessions, using IRF coding and Epistemic Network Analysis. It finds that human tutors use more questioning and feedback while AI tutors lean on extended explanations, supporting the view that current LLMs produce pedagogically shallower interactions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'AI explanation-simplistic response' vs human 'question-factual response' contrast depends on a BERT coder validated only on human dialogues; no reliability metrics are reported for AI-simulated text, so the pattern could be a coding artifact.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the BERT coder's validity on AI-simulated dialogues is the gatekeeper for virtually every quantitative result in the paper. I agree this is the single most important threat to the central claim. All of the headline findings—the subtype proportion differences in Table 2 and the ENA network contrast in Figures 3–4—are computed from codes assigned by a model that was fine-tuned exclusively on human-coded transcripts. Without reliability evidence for the AI corpus, the 'fundamental divergence' could be a measurement artifact rather than a property of the dialogues. The paper's other limitations (one LLM, one simulation framework, undergraduate tutors, GPT-polishing of human transcripts) are real but mostly affect generalizability; the coding reliability gap affects internal validity. The reader's CONDITIONAL verdict is appropriate, and this concern should be treated as a prerequisite for full acceptance rather than a reason to reject outright, because the study design is otherwise reasonable and the reported human-coding Kappa is good. A concrete re-coding study on a subset of AI dialogues would settle the question directly.","tokens_in":10207,"tokens_out":4906,"duration_ms":59697,"concrete_test":"Have two trained human coders independently apply the IRF coding scheme to a random subset (e.g., 10–12 dialogues) of the AI-simulated corpus, blind to the BERT codes. Compute (a) human-human Cohen's Kappa, (b) BERT-vs-human Cohen's Kappa per code, and (c) a confusion matrix focusing on F-E, R-SR, I-Q, and R-FR. Then recompute the ENA centroid separation and the strongest network edges using only the human codes for that subset. If BERT-vs-human Kappa on AI dialogues is close to the reported human-human Kappa (0.824) and the F-E↔R-SR edge remains strongest in the human-coded AI network, the concern is resolved. If Kappa is substantially lower or the edge pattern shifts, the headline divergence is at least partly a coding artifact and the conclusions must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.4 reports human-human Cohen's Kappa of 0.824 for the human dialogues, but the AI-simulated dialogues were coded by a BERT model fine-tuned on human-coded data, with only an unspecified 'manual verification.' No inter-coder reliability, confusion matrix, or error analysis is provided for the out-of-domain AI text. Since the paper's central ENA finding—that human dialogues center on an I-Q/R-FR loop while AI dialogues center on an F-E/R-SR loop—is computed entirely from these BERT codes, a systematic bias in how the BERT coder labels AI utterances would propagate directly into the headline result. AI-generated dialogue has systematic linguistic differences from polished human transcripts (e.g., longer, more formulaic explanations; shorter, more predictable student replies), and a classifier trained exclusively on human-coded examples can mislabel such text in a consistent direction. For instance, if the coder over-assigns F-E to AI teacher turns and R-SR to AI student turns, the signature 'explanation-simplistic response' loop and the significant X-axis centroid separation in Figure 3 would be inflated or even entirely artifact. The paired t-tests in Table 2 inherit the same measurement risk, because the behavioral proportions being compared are generated by the same coder. This is not a question of whether the authors are honest—it is a question of whether the measurement instrument is valid for the second corpus it was applied to, and the paper provides no quantitative evidence that it is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares 49 authentic human one-on-one math tutoring dialogues (fifth-grade students; university volunteer tutors; transcripts polished with GPT) with 49 AI-simulated dialogues generated using GPT-4o and a three-agent SocraticLM framework, with the same tutoring questions and a 'core tutoring approach' distilled from each human dialogue as inputs. Both corpora were coded with a nine-code IRF scheme: human dialogues were double-coded by human annotators (Cohen's κ = 0.824), while AI dialogues were coded by a BERT classifier fine-tuned on the human-coded data, with an unspecified 'manual verification.' The authors then ran paired t-tests on behavioral proportions and Epistemic Network Analysis (ENA). Results show that human dialogues have longer utterances and higher proportions of I-Q, R-FR, and F-F, whereas AI dialogues have higher R-SR, R-RR, and F-E. ENA reveals a large X-axis centroid separation (t(84.35)=9.33, p<0.001, d=1.97), with human networks centered on an I-Q↔R-FR connection and AI networks on an F-E↔R-SR connection. The paper concludes that human tutoring is cognitively guided, while AI tutoring is essentially an information-transfer loop. The authors acknowledge limitations regarding the tutor sample, simulated student diversity, and the entanglement of LLM and agent design.","tokens_in":10520,"tokens_out":5830,"duration_ms":68691,"significance":"If the measurement pipeline is valid for both corpora, this is a valuable empirical contribution: it uses a paired design that controls for problem content, applies a well-established IRF framework, and demonstrates ENA as a visualization and inference tool for comparing human and AI educational dialogue. The large effect sizes and the clear separation of interactional patterns make the finding easy to communicate to the AIED community. The paper is also honest about several limitations (student tutors, one LLM/agent framework, simplified AI student). However, the central result rests on two unvalidated preprocessing/coding steps: the BERT coder's reliability on AI-generated text and the GPT-based polishing of human transcripts. These concerns are not peripheral; they directly affect the headline comparisons in Table 2 and Figures 3–4. The manuscript is therefore a promising but not yet fully supported contribution; the required additional validation is feasible within the scope of a revision.","major_comments":[{"comment":"The reliability evidence reported for human coding does not transfer to the AI-simulated corpus. Cohen's κ = 0.824 applies to two human coders on human dialogues; the AI dialogues were then coded by a BERT model fine-tuned on human-coded examples, and the 'manual verification' is not quantified. Because AI-generated text is systematically different from human transcripts, the coder can mislabel specific codes in a consistent direction — for example, over-assigning F-E to AI teacher explanations or R-SR to AI student replies. Table 2 and the ENA networks in Figures 3–4 are computed entirely from these codes, so the central 'explanation–simplistic response' loop may in part be a coding artifact. Please report a human-coded reliability sample for AI dialogues: per-code precision/recall or a confusion matrix, plus Kappa, and state how disagreements with the 'manual verification' were resolve","section":"§3.4 Data analysis"},{"comment":"All human transcripts were 'refined sentence by sentence using GPT-based text polishing' before coding. This preprocessing can inflate measured utterance lengths and rewrite or remove short, disfluent student turns such as 'mm,' 'okay,' and hesitations — exactly the kinds of utterances coded as R-SR. The human-vs-AI differences in utterance length and in R-SR/F-F proportions could therefore be partly introduced by the polishing step rather than by authentic interaction. Please quantify this effect: e.g., compare original ASR transcripts with the polished versions on a subset, or code the raw turns most likely to be affected, and report the polishing model/version and the editing criteria.","section":"§3.1 Participants and experiment procedure"},{"comment":"The paper does not specify how the 'core tutoring approach' was extracted from each human dialogue, how many AI dialogues were generated (presumably 49, but this is not stated), or the generation parameters (model version, temperature, sampling, number of runs, prompt template). If each AI dialogue is paired with a human dialogue via the extracted 'core tutoring approach,' the paired t-tests in Table 2 need this pairing made explicit. More generally, the abstract and conclusion claim a 'fundamental divergence' for 'AI tutoring' as a whole, but the evidence comes from one LLM (GPT-4o), one three-agent framework, and one prompt template. The Discussion does acknowledge this entanglement, but the claims in the abstract should be scoped to the tested configuration. Please report the full generation protocol so that readers can assess reproducibility and generality.","section":"§3.2 Simulation Data Generation"}],"minor_comments":[{"comment":"Ten paired t-tests are run without a multiple-comparison correction; report adjusted p-values or an FDR procedure, and include effect sizes (e.g., Cohen's d) for each behavioral subtype. Also clarify whether the pairs are (human dialogue i, AI dialogue generated from i).","section":"Table 2"},{"comment":"Report the ENA parameters used in the ENA Web Toolkit: window size, co-occurrence threshold, normalization, and rotation method. State how many 'lines' or units were used in the t-test for centroid separation (the df = 84.35 suggests a Welch test; please specify the unit of analysis).","section":"Figure 3 / ENA settings"},{"comment":"The phrase 'GPT-based text polishing' is underspecified. State which model version was used and whether polishing was done automatically or with human oversight; this matters for reproducibility and for the measurement concern above.","section":"§3.1, GPT polishing"},{"comment":"The paper cites 'Hila, A. (2025)' as the source for ENA, but this reference does not appear to be the methodological origin of Epistemic Network Analysis. Please cite the canonical ENA methodology papers (e.g., Shaffer et al.) in addition to any application-specific reference.","section":"References"},{"comment":"The 'core tutoring approach' is a key input to the simulation but is never defined. Provide an example or a formal description of what the distilled approach contains (e.g., sequence of pedagogical moves, solution steps) to support replication.","section":"§3.2 / Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the BERT-coder validity on out-of-domain AI text; the reviewer's stress-test concern lands squarely. The GPT-polishing issue is a second load-bearing point. Both are addressable with additional validation and sensitivity analyses, so I recommend major revision rather than rejection. The paper fits the ICCE venue and, if the measurement concerns are resolved, would be a useful contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper offers a genuinely useful matched comparison and a plausible qualitative conclusion—human tutors scaffold by questioning, AI tutors tend to explain and wait for a simple acknowledgment. But the specific ENA structure—the 'explanation-simplistic response' loop—depends on a BERT coder applied to AI-simulated text with no reported reliability. That is the soft spot that matters.\n\nWhat's new: the authors took 49 real one-on-one math tutoring dialogues, extracted the same teaching prompts, and used a three-agent LLM framework to generate matched AI dialogues. Both corpora were coded with the same fine-grained IRF scheme, then analyzed with ENA. The paired design and the focus on interactional loops rather than just frequency counts are a step up from earlier descriptive comparisons. The t-tests and centroid separation are reported with effect sizes. And the qualitative framing—human dialogue as cognitive guidance vs. AI as information transfer—matches what many in the field suspected.\n\nWhere I'd push back: Section 3.4 says human dialogues were coded by two researchers with Kappa .824, then a BERT model fine-tuned on those codes was used for the AI dialogues, with only 'manual verification.' No inter-coder reliability, confusion matrix, or error analysis for the AI text is given. Since AI-generated language is systematically different—more formulaic explanations, shorter student replies—the coder could be biased in exactly the direction of the headline finding. The paired t-tests in Table 2 inherit that risk because the proportions come from the same coder. This isn't about honesty; it's about whether the measurement instrument transfers.\n\nA second, smaller issue: human transcripts were polished with GPT before coding. That could smooth away the very disfluencies that distinguish human talk from AI talk, though it is unlikely to create the whole pattern. The sample is small (49 dialogues) and only one LLM and one framework, so generalizability is limited. The authors acknowledge some limitations, but not the coder transfer problem.\n\nOverall, this is worth engaging with. The framework is a solid contribution, and the central intuition is probably right. But the specific ENA loop should be treated as provisional until the authors show the BERT coder works on AI-simulated text—say, by human-coding a subsample and reporting reliability. I'd send it out for peer review with a clear request for that evidence.","headline":"Plausible and useful matched comparison, but the headline ENA contrast rests on a BERT coder applied to AI text with no reported reliability—needs that evidence before the specific loop claim is taken as established.","tokens_in":10997,"tokens_out":2865,"would_cite":false,"duration_ms":31030,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that AI-simulated one-on-one tutoring dialogues are structurally different from human ones: humans use a question-factual response-feedback loop, while AI uses an explanation-simplistic response loop.","keywords":["AI tutoring","dialogue structure","Initiation-Response-Feedback (IRF)","Epistemic Network Analysis","LLM simulation","one-on-one instruction","Socratic questioning","generative educational dialogue"],"falsifier":"A concrete falsifying check: have trained human coders independently code the same AI-simulated dialogues with the IRF scheme, then compare their codes with the automatic coder's on the F-E and R-SR categories. If agreement is poor on AI text and the ENA networks built from human codes no longer separate along the X-axis (or the F-E↔R-SR connection weakens), the claimed fundamental divergence is an artifact of automated coding.","tokens_in":10102,"feed_emoji":"🧑‍🏫","tokens_out":5634,"duration_ms":58369,"temperature":0.7,"pith_summary":"The paper asks whether large-language-model-generated teacher-student dialogues reproduce the structure of real one-on-one tutoring. Comparing authentic human dialogues with AI-simulated dialogues on the same math tutoring prompts, it finds they do not: human tutors initiate more, ask more questions, give more general feedback, and draw out longer factual student answers, while AI tutors explain more and receive short confirmatory replies. The signature result is a statistically separated pair of interaction loops in the network analysis: human dialogue centers on a question-factual response-feedback loop, AI dialogue on an explanation-simplistic response loop. If true, the finding gives a concrete target for improving generative educational dialogue systems—synthetic tutoring data is not just slightly different but structurally shallower—and offers a quantitative way to evaluate progress.","feed_headline":"AI tutoring runs on explain-and-confirm, humans on ask-and-answer","feed_subtitle":"Real and AI-simulated tutoring differ in structure: humans ask-and-guide, AI explains and receives short replies.","key_machinery":"The analysis rests on an IRF-based coding scheme (Initiation-Response-Feedback) with ten subcategories—Questioning, Hints, Modeling; Refusal, Simplistic Response, Factual Response, Open-ended Response; Feeding Back, Instructing, Explaining—combined with Epistemic Network Analysis, a method that projects how often coded behaviors co-occur into a low-dimensional space. The load-bearing edges are the I-Q↔R-FR connection in human dialogues and the F-E↔R-SR connection in AI dialogues; the separation of the two networks' centroids along the X-axis is the quantitative evidence for a fundamental divergence.","core_discovery":"The central claim is that the difference between AI-simulated and human tutoring is not merely fluency or utterance length but the underlying interaction structure. In the coded behavior space, human dialogues are organized around a tightly coupled Questioning–Factual Response pair (I-Q with R-FR), with feedback closing the loop; AI dialogues are organized around Explaining–Simplistic Response (F-E with R-SR). The network centroids separate significantly along one core dimension (t(84.35)=9.33, p<0.001, d=1.97), which the authors interpret as a single axis separating question-centered guided instruction from explanation-centered information transfer. The paper concludes that human dialogue i","pith_inferences":["Editorial inference: The same IRF+ENA contrast could serve as an automated outcome metric for attempts to improve AI tutoring—e.g., fine-tuning or preference-optimizing toward I-Q→R-FR transitions—rather than only measuring fluency or task success.","Editorial inference: Because the simulation was seeded with each human dialogue's distilled core tutoring approach, the observed divergence may understate how far an unconstrained LLM tutor would drift from human structure; a free-instructed condition would reveal whether the constraint helps or hides the gap.","Editorial inference: The coding pipeline is the most testable seam—a blind sample of AI dialogues double-coded by trained humans would separate real structural differences from classifier bias and either strengthen or qualify the central claim."],"forward_implications":["LLM-generated tutoring corpora should not be treated as drop-in substitutes for human teaching data in training or evaluation, because their interaction structure is demonstrably different.","Evaluation of generative educational dialogue systems should measure interaction-loop structure—whether the system supports question → factual response → feedback rather than explanation → simple acknowledgment—not just linguistic fluency.","The deficits are specific: AI dialogues are weakest in initiation, questioning, and general feedback, while strong in explanation, pointing to concrete behaviors to target with prompt design or fine-tuning.","The IRF+ENA signatures give a quantitative benchmark for tracking whether future AI tutoring systems become more human-like as their behavior changes.","The absence of the question-factual response-feedback loop is a measurable warning sign that an AI tutor is likely providing information transfer rather than Socratic cognitive guidance."],"supporting_citations":[{"why":"Supplies the IRF initiation-response-feedback discourse model that the coding scheme builds on.","marker":"(Sinclair & Coulthard, 2013)"},{"why":"Provides the six teacher-scaffolding types from which the Initiation and Feedback subcategories are adapted.","marker":"(Van de Pol et al., 2010)"},{"why":"Provides the hierarchical response classification that the Response dimension extends.","marker":"(Yang et al., 2023)"},{"why":"Supplies the tripartite SocraticLM agent framework used to generate the AI-simulated dialogues.","marker":"(Liu et al., 2024)"},{"why":"Establishes the IRF structure as the analytic frame for comparing teacher-student talk.","marker":"(Y. Liu, 2008)"},{"why":"Cited as the source of the Epistemic Network Analysis method used to compare interaction networks.","marker":"(Hila, 2025)"}],"fun_headline_variants":["Humans ask, AI tells — that's the tutoring gap","AI tutors explain, human tutors question","Tutoring split: humans guide with questions, AI with answers","Human tutoring is question-driven; AI is explanation-driven","The tutoring difference: asking vs telling"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The argument assumes the automatic coder, trained on human-labelled dialogues, labels AI-simulated dialogues without systematic bias; if it tags AI explanations and short replies differently for tool reasons, the central contrast could be an artifact of the coding rather than of the dialogues.","fun_headline_variants_meta":{"raw":{"variants":["Humans ask, AI tells — that's the tutoring gap","AI tutors explain, human tutors question","Tutoring split: humans guide with questions, AI with answers","Human tutoring is question-driven; AI is explanation-driven","The tutoring difference: asking vs telling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000466,"raw_usage":{"total_tokens":2164,"prompt_tokens":752,"completion_tokens":1412,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":1337}},"tokens_in":496,"tokens_out":1412,"duration_ms":14911,"temperature":1.0,"reasoning_tokens":1337,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:02:59.988849+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifying check: have trained human coders independently code the same AI-simulated dialogues with the IRF scheme, then compare their codes with the automatic coder's on the F-E and R-SR categories. If agreement is poor on AI text and the ENA networks built from human codes no longer separate along the X-axis (or the F-E↔R-SR connection weakens), the claimed fundamental divergence is an artifact of automated coding.","supporting_citations":[{"cited_title":"AI & SOCIETY","cited_arxiv_id":null,"evidence_quote":"Cited as the source of the Epistemic Network Analysis method used to compare interaction networks."}],"review_version":1}