REVIEW 12 cited by
Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Generative AI, particularly Language Models (LMs), has the potential to transform real-world domains with societal impact, particularly where access to experts is limited. For example, in education, training novice educators with expert guidance is important for effectiveness but expensive, creating significant barriers to improving education quality at scale. This challenge disproportionately harms students from under-served communities, who stand to gain the most from high-quality education. We introduce Tutor CoPilot, a novel Human-AI approach that leverages a model of expert thinking to provide expert-like guidance to tutors as they tutor. This study is the first randomized controlled trial of a Human-AI system in live tutoring, involving 900 tutors and 1,800 K-12 students from historically under-served communities. Following a preregistered analysis plan, we find that students working with tutors that have access to Tutor CoPilot are 4 percentage points (p.p.) more likely to master topics (p<0.01). Notably, students of lower-rated tutors experienced the greatest benefit, improving mastery by 9 p.p. We find that Tutor CoPilot costs only $20 per-tutor annually. We analyze 550,000+ messages using classifiers to identify pedagogical strategies, and find that tutors with access to Tutor CoPilot are more likely to use high-quality strategies to foster student understanding (e.g., asking guiding questions) and less likely to give away the answer to the student. Tutor interviews highlight how Tutor CoPilot's guidance helps tutors to respond to student needs, though they flag issues in Tutor CoPilot, such as generating suggestions that are not grade-level appropriate. Altogether, our study of Tutor CoPilot demonstrates how Human-AI systems can scale expertise in real-world domains, bridge gaps in skills and create a future where high-quality education is accessible to all students.
Forward citations
Cited by 12 Pith papers
-
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
A new benchmark finds low to moderate human agency support in 20 LLM assistants across six dimensions.
-
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.
-
Experimental Evidence on the Learning Impact of Generative AI
Off-the-shelf generative AI access raises student knowledge-test scores by 0.27 SD immediately and one week later, with delayed unaided essay gains concentrated among students who use AI to explain concepts rather tha...
-
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
A 30-day, knowledge-tracing-grounded simulated learner benchmark for tutoring agents finds that no base model or harness alone determines quality and that almost all tested combinations plateau within days.
-
Conversational Human Audio-visual Talking Dialogue Generation
CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.
-
A Causal Framework for Estimating Heterogeneous Effects of On-Demand Tutoring
Requesting on-demand human tutoring raises middle-school students' next-problem correctness by about 4 percentage points on average, with large session-level heterogeneity, when estimated from observational platform l...
-
Effective Red-Teaming of Policy-Adherent Agents
A policy-aware red-teaming system (CRAFT) induces policy violations in LLM customer service agents at much higher rates than generic jailbreak prompts, using a new security-focused benchmark (tau-break) built from tau-bench.
-
Evaluating Gemini in an arena for learning
A blind expert-judged arena found Gemini 2.5 Pro preferred for tutoring in 73.2% of non-tied matchups against four competing models.
-
A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics
Students rated locally-run Llama-3.1 feedback roughly on par with GPT-4 and often above human TAs in technical courses, but preferred human feedback for specialized writing.
-
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
LLMs reach only 49% F1 (MathDial) and 27% F1 (AlgebraNation) when predicting the next tutor move, while tutor moves help predict dialogue success in AlgebraNation but not consistently on MathDial.
-
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
A large LLM-to-LLM math tutoring simulation across 11 languages shows English-language hints often yield the largest accuracy gains for student models, but the low-resource-language results lack statistical support.
-
Human Authenticity and Flourishing in an AI-Driven World: Edmund's Journey and the Call for Mindfulness
An HCI position paper argues for a Human Flourishing Benchmark that scores AI systems on cognitive preservation, autonomy, skill development, and relational authenticity.
Discussion (0). Sign in to comment.