{"id":"754fe6f0-aff4-4b42-8b6b-e78513044d77","arxiv_id":"2412.11388","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A student LLM that asks questions to a teacher LLM improves its quiz scores more than one that reads a static lesson, but the gain may be partly driven by the quiz-performance feedback it receives.","lead":"INTERACT lets a student LLM ask a teacher LLM questions about new concepts and then quizzes what the student learned. Across 1,347 post-2023 contexts, interactive questioning improved quiz accuracy by up to about 25 points, but a missing control for quiz feedback weakens the causal claim.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quiz-performance feedback (Listing 8) confounds the dynamic vs. static comparison; without a no-feedback ablation, the 'question-driven' causal claim is unsupported.","rationale":"Good-faith reading: the paper builds a sizable benchmark (1,347 post-cutoff contexts) and shows a consistent dynamic advantage across many architectures; the adversarial quiz filtering and the random-interaction ablation (Table 8) are real efforts to control for content relevance. However, the central causal claim — that student-initiated questioning, rather than feedback or extra information, drives improvement — is not identifiable from the current design. Listing 8 explicitly inserts quiz performance into the prompt that generates the next question, and the static baseline has no equivalent. The reader's weakest_assumption identifies this exact confound. I agree, with one refinement: unless quiz_performance includes per-question detail, the student may not be able to target 'weak areas' specifically, but the score can still modulate questioning effort, so the confound stands. The abstract's 'matching static baselines' claim is also contradicted by Table 2 (recovery 81–93%). These are internal-consistency and design concerns, not disagreements with consensus. A single ablation removing quiz feedback would settle the causal question; if the gain persists, the paper's main empirical claim survives in a weakened form. Hence I keep the reader's REJECT rather than moving it.","tokens_in":43996,"tokens_out":6286,"duration_ms":56037,"concrete_test":"Rerun the dynamic condition (both w/ and w/o lesson) with the 'Quiz Performance' field removed from the question-generation prompt (Listing 8), holding every other component fixed, across gpt-4o-mini and LLaMA-8B on all five domains. If the post-interaction quiz accuracy stays within a few points of the original (e.g., retains >90% of the original gain), the feedback confound is not load-bearing; if the gain drops substantially, the reported improvement is largely feedback-driven rather than question-driven. Additionally, a complementary control that provides the same quiz score but forces a fixed generic question each round would isolate whether question choice matters.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dynamic condition in §3.3/§4.2 is compared against a static baseline that differs on at least three axes: the ability to ask questions, access to additional teacher-generated answers, and receipt of quiz-performance feedback before each new question (Listing 8, Appendix D: 'Quiz Performance: {{ quiz_performance }}'). Static students receive no such feedback, so the reported gains (e.g., +25.77 for gpt-4o-mini in Table 2) cannot be attributed to student-initiated questioning per se. The paper's RQ4 borrowed-interaction control shows passive exposure to a transcript is not sufficient, but that control does not hold feedback fixed; a student who knows its quiz score can adapt its effort and focus even without per-question breakdowns. Because the central claim is specifically about question-driven learning, the absence of a no-feedback (or fixed-question) control makes the causal interpretation unsupported. The abstract further overstates the result by claiming cold-start students 'matching static learning baselines' while Table 2 reports recovery of only 81–93% (e.g., 91.23% for gpt-4o-mini).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces INTERACT, a framework in which a 'student' LLM learns a new concept by asking a 'teacher' LLM questions over multiple dialogue turns, optionally starting from a static generated lesson. The authors construct 1,347 contexts from after the models' knowledge cutoffs (song lyrics, news articles, movie plots, academic papers, and COCO images), generate nine-question quizzes per concept with gpt-4o plus adversarial filtering, and compare static-lesson, dynamic-without-lesson, and dynamic-with-lesson conditions across seven text LLMs and two multimodal models. They report that dynamic students without a lesson improve by about 12-26 points over five turns, that dynamic students with a lesson improve modestly over the static lesson, and that passive exposure to borrowed transcripts does not substitute for interaction. RQ5 attempts to predict learning gains from 53 transcript features.","tokens_in":44215,"tokens_out":8779,"duration_ms":78031,"significance":"The dataset and multi-model evaluation are potentially useful resources for studying interactive concept acquisition. The authors are careful about post-cutoff data, provide code and data, use adversarial quiz filtering, and include several ablations (random interactions, borrowed transcripts, teacher/lesson quality) that go beyond a simple static-versus-dynamic comparison. However, the central causal claim—that student-initiated questioning itself drives the gains—is not supported by the current design. The dynamic condition couples questioning with quiz-performance feedback and with additional teacher-generated content, so the reported improvements cannot be attributed specifically to question-driven learning. As a description of an interactive-learning pipeline, the resource has value; as evidence for the paper's headline claim, it needs substantial additional controls.","major_comments":[{"comment":"The dynamic-versus-static comparison is not a clean test of question-driven learning. The dynamic student's question-generation prompt (Listing 8) includes 'Quiz Performance: {{ quiz_performance }}' whenever a quiz score is available, while the static prompts (Listings 4 and 5) contain no such feedback. The dynamic condition therefore differs from the static baseline on at least three axes: the ability to pose questions, access to additional teacher answers grounded in the full context, and receipt of one's own quiz score before choosing the next question. Any of these could explain the gains in Figure 4 and Table 2, so the improvement cannot be attributed specifically to student-initiated questioning. The quiz-feedback channel is especially problematic because it lets the student target its remaining questions at the exact evaluation instrument, making part of the measured gain circular. Please add (a) a dynamic condition without quiz feedback, (b) a static condition with the same quiz feedback, and (c) a static full-context condition, or an equivalent factorial design, before claiming that interaction itself causes the improvement.","section":"§3.3/§4.2 and Appendix D Listing 8"},{"comment":"The abstract and introduction state that cold-start students 'match static learning baselines in as few as five dialogue turns,' but Table 2 reports recovery ratios of 83-94% of the static-lesson performance (e.g., 91.23% for gpt-4o-mini and 83.21% for Gemma-9B), not parity. Section 4.2 itself acknowledges that 'most dynamic students do not surpass the static-lesson baseline within five rounds.' The unqualified 'matching' claim should be replaced with a statement that interactive students approach but generally do not match the static-lesson baseline within five turns, with the qualification depending on domain and model.","section":"Abstract and Table 2"},{"comment":"The RQ4 borrowed-interaction control does not hold feedback fixed. The passive transcript condition provides the dialogue to a weaker student, but the original dynamic student used its own quiz-performance feedback to choose later questions, and the passive student never receives that feedback. The conclusion that 'passive exposure ... cannot substitute for pro-active engagement' therefore conflates activity with feedback access. A passive condition that receives the same feedback signal, or a factorial manipulation that separates feedback from questioning, is needed to support this conclusion.","section":"§4.4, Tables 6-7"},{"comment":"RQ5's predictive analysis uses features defined with respect to the quiz itself, including 'Student Quiz Coverage,' 'Teacher Quiz Coverage,' and 'Student Semantic Alignment' (overlap or embedding similarity with quiz question tokens). Because the quiz is the outcome measure, these features introduce label leakage, so the reported R2 values (up to 0.14 in Song Lyrics) cannot be interpreted as evidence for genuinely predictive interaction features. Please re-run the analysis excluding all quiz-overlap/quiz-similarity features and report both sets of results, or justify why the overlap is not leakage.","section":"§4.5 and Appendix Table 12"}],"minor_comments":[{"comment":"The caption contains a typo: 'pretaining' should be 'pretraining.'","section":"Figure 1"},{"comment":"The axis labels in the supplied PDF appear as escaped Unicode sequences such as '/uni0030/uni0025' instead of readable text; these need to be rendered correctly.","section":"Figures 3-8"},{"comment":"Appendix A cites Srivastava and Goodman (2021), Zhou et al. (2024), Wu et al. (2024), Kim and Rush (2016), Sanh et al. (2019), and Agarwal et al. (2023), but these references are missing from the reference list.","section":"Appendix A / References"},{"comment":"The text says 'Table 10 provides some example questions the gpt-4o student LLM asked,' but Section 3.4 states gpt-4o was not evaluated as a student due to cost, and the rows of Table 10 are for gpt-4o-mini. Please correct this discrepancy.","section":"Appendix B.3 / Table 10"},{"comment":"The column header 'Recovery of Student w/o Lesson (%) wrt S w/ L Start wrt Teacher' is ambiguous; clarify whether the denominator is the static-lesson start performance or the teacher performance, and state the convention in the caption.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has useful resources and the missing controls are well-defined, so I recommend major revision rather than rejection. The authors should be able to reuse their existing code and dataset to run the no-feedback and feedback-without-questioning conditions. The abstract's 'matching' claim should also be corrected, and RQ5 should be re-run without the quiz-overlap features. If the authors decline to add these controls, the central claim of the paper would remain unsupported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this one. First, the authors have put together a genuinely useful benchmark: 1,347 contexts released after the models' pretraining cutoff, across songs, news, movies, papers, and images, with code and data public. The adversarial filtering of quiz questions is careful—they throw out questions gpt-4o-mini can answer without context—and the manual validation on 234 questions is solid. Second, the paper's headline claim—that student-initiated questioning drives learning—is not supported by the experiments as designed. The dynamic student's question-generation prompt includes 'Quiz Performance: {{ quiz_performance }}' (Listing 8); the static baseline gets no such feedback. So the dynamic condition differs from static on at least three axes: the ability to ask questions, access to teacher follow-up answers, and knowledge of its own quiz score before each new question. The +25-point gains cannot be attributed to questioning per se. The RQ4 borrowed-interaction control shows passive transcript exposure isn't enough, but it doesn't hold feedback fixed; a student who knows its score can focus future questions on weak areas even without item-level breakdowns.\n\nWhat else is good: they run seven text models and three multimodal ones, include a random-interaction ablation that shows the transcript content matters, and the teacher-strength results (strong vs weak teacher barely changes final performance) are interesting. The abstract overstates the 'matching' result—Table 2 shows 81–93% recovery, not equality—but that's a minor sin compared to the confound.\n\nIs the paper worthless? No. The benchmark is a contribution, and the bug is fixable. A no-feedback dynamic control, or a fixed-question (non-adaptive) student that still receives quiz feedback, would separate the effects. As it stands, the central causal story is unsupported. I'd send this to referees rather than desk-reject, because the dataset and the question are worth airing and the fix is within reach. But I would not accept it in this form. For my own work, I'd be happy to cite the benchmark once the control runs are in; I wouldn't cite the causal claim.","headline":"Strong benchmark, load-bearing confound: the dynamic condition includes quiz-feedback that the static baseline lacks, so the 'question-driven' claim is unsupported as-is.","tokens_in":44719,"tokens_out":2724,"would_cite":true,"duration_ms":24800,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T50","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs learn new concepts more effectively when they actively ask a teacher questions, reaching up to 25% better quiz scores than students who passively receive a summary.","keywords":["interactive learning","question-driven learning","teacher-student dialogue","large language models","concept acquisition","static vs dynamic learning","benchmark dataset","cold-start learning"],"falsifier":"Run the dynamic condition with the same quiz-performance feedback but with the student forced to receive teacher answers to a fixed sequence of questions instead of asking its own; if quiz scores improve as much as in the free-question condition, the learning gains are due to feedback, not questioning.","tokens_in":43825,"feed_emoji":"💬","tokens_out":6563,"duration_ms":55795,"temperature":0.7,"pith_summary":"The paper sets out to show that large language models can learn new concepts through interactive, question-driven dialogue rather than passive absorption of static summaries. The INTERACT framework lets a student LLM ask a teacher LLM questions about a concept, take a quiz after each turn, and keep asking until it understands. Across 1,347 post-training-cutoff contexts — song lyrics, news articles, movie plots, academic papers, and images — interactive students consistently improved, by up to 25 percentage points, and cold-start students matched static-lesson baselines in as few as five turns. The authors argue that the student's own asking, not the quality of the teacher or the lesson, is what drives learning.","feed_headline":"LLMs learn up to 25% better by asking questions","feed_subtitle":"Students that query a teacher beat passive lessons, matching static baselines in five turns.","key_machinery":"The load-bearing mechanism is the INTERACT dialogue loop: at each turn the student LLM generates a single question, the teacher LLM answers using only the ground-truth context document, the student's conversation history (plus optionally a static lesson) is appended to its context, and the student completes a nine-question quiz before asking again. The framework isolates student-driven inquiry by testing on contexts published after the models' pretraining cutoff, so answers cannot be memorized, and by comparing against static-lesson and teacher-upper-bound baselines. The key ablation replaces real interactions with random ones, which collapses performance, showing the student actually uses the exchanged information.","core_discovery":"The central claim is that student-initiated questioning is a genuine mechanism for concept acquisition in LLMs. When a student asks its own questions of a teacher that answers from the ground-truth context, quiz performance rises across seven model families and five domains, with absolute gains of 12 to 26 percentage points over five turns. Swapping in a stronger teacher or a higher-quality static lesson changes final performance by only about one percent, and passively reading another student's high-quality interaction transcripts does not reproduce the benefit of asking. The authors conclude that the act of asking questions — not the quality of the content consumed — is what makes interactive learning work, although students still remain below teacher-level quiz scores.","pith_inferences":["The paper's prompt includes the student's own quiz score before question generation, so a control that gives identical feedback with fixed questions is needed to separate 'asking well' from 'knowing what to review'; this control is not reported.","If the gains are largely iterative exposure, then a five-hint or five-answer baseline that is not student-chosen might close much of the gap, making the comparison against a single static lesson a best-case framing.","A natural extension is to run more than five turns: the data suggest continued gains, so the 'matches static baselines in five turns' result may understate what interaction can achieve at ten or twenty turns.","The feature-analysis result (variance explained near zero except for lyrics) implies that the field still lacks a good predictive measure of what makes a question useful; measuring whether a question targets a quiz-covered fact the student got wrong might be more direct."],"forward_implications":["Interactive LLM students can bootstrap knowledge about a brand-new topic in about five questions, reaching most of the way to teacher-level performance without any curated lesson.","Since teacher strength matters little after interaction, even a smaller or weaker teacher model can support effective learning, lowering the compute barrier for tutors.","Because passively reading strong students' transcripts does not help, interactive learners should generate their own questions rather than receive explanations.","The 1,347-context benchmark provides a reusable testbed for comparing conversational learning methods across lyrical, journalistic, cinematic, scientific, and visual content."],"supporting_citations":[{"why":"Supplies the human-tutoring result that one-to-one interactive tutoring outperforms classroom instruction, motivating the interactive design.","marker":"(Bloom, 1984)"},{"why":"Provides evidence that student-led questioning and responsive explanation drive tutoring effectiveness, not just personalization.","marker":"(VanLehn, 2011)"},{"why":"Offers the theory of question asking as a core learning mechanism that INTERACT instantiates.","marker":"(Ram, 1991)"},{"why":"Grounds the teacher-student dialogue framing in social learning theory.","marker":"(Vygotsky and Cole, 1978)"},{"why":"Underlies the dataset design principle that pretraining knowledge contaminates evaluation, motivating post-cutoff contexts.","marker":"(Roberts et al., 2020)"},{"why":"Supplies the COCO image collection used for the visual domain of the benchmark.","marker":"(Lin et al., 2014)"}],"fun_headline_variants":["Asking questions, not just reading, boosts LLM learning by 25%","LLMs learn 25% better when they ask questions, not just read","Student LLMs that ask questions beat passive learning by 25%","Interactive questioning gives LLMs a 25% learning boost over reading","Question-driven learning makes LLMs 25% better, even with weak teachers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central causal claim would collapse if the same gains came from the quiz-performance feedback the student receives before each new question, rather than from the student's freedom to ask its own questions.","fun_headline_variants_meta":{"raw":{"variants":["Asking questions, not just reading, boosts LLM learning by 25%","LLMs learn 25% better when they ask questions, not just read","Student LLMs that ask questions beat passive learning by 25%","Interactive questioning gives LLMs a 25% learning boost over reading","Question-driven learning makes LLMs 25% better, even with weak teachers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000447,"raw_usage":{"total_tokens":2196,"prompt_tokens":823,"completion_tokens":1373,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":1275}},"tokens_in":439,"tokens_out":1373,"duration_ms":10099,"temperature":1.0,"reasoning_tokens":1275,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:58:31.237703+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the dynamic condition with the same quiz-performance feedback but with the student forced to receive teacher answers to a fixed sequence of questions instead of asking its own; if quiz scores improve as much as in the free-question condition, the learning gains are due to feedback, not questioning.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Offers the theory of question asking as a core learning mechanism that INTERACT instantiates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the teacher-student dialogue framing in social learning theory."}],"review_version":1}