Pith. sign in

REVIEW 12 cited by

Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03017 v2 pith:X5PSIEQ4 submitted 2024-10-03 cs.CL

classification cs.CL
keywords tutorcopilottutorsstudentseducationhuman-aiaccessfind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Generative AI, particularly Language Models (LMs), has the potential to transform real-world domains with societal impact, particularly where access to experts is limited. For example, in education, training novice educators with expert guidance is important for effectiveness but expensive, creating significant barriers to improving education quality at scale. This challenge disproportionately harms students from under-served communities, who stand to gain the most from high-quality education. We introduce Tutor CoPilot, a novel Human-AI approach that leverages a model of expert thinking to provide expert-like guidance to tutors as they tutor. This study is the first randomized controlled trial of a Human-AI system in live tutoring, involving 900 tutors and 1,800 K-12 students from historically under-served communities. Following a preregistered analysis plan, we find that students working with tutors that have access to Tutor CoPilot are 4 percentage points (p.p.) more likely to master topics (p<0.01). Notably, students of lower-rated tutors experienced the greatest benefit, improving mastery by 9 p.p. We find that Tutor CoPilot costs only $20 per-tutor annually. We analyze 550,000+ messages using classifiers to identify pedagogical strategies, and find that tutors with access to Tutor CoPilot are more likely to use high-quality strategies to foster student understanding (e.g., asking guiding questions) and less likely to give away the answer to the student. Tutor interviews highlight how Tutor CoPilot's guidance helps tutors to respond to student needs, though they flag issues in Tutor CoPilot, such as generating suggestions that are not grade-level appropriate. Altogether, our study of Tutor CoPilot demonstrates how Human-AI systems can scale expertise in real-world domains, bridge gaps in skills and create a future where high-quality education is accessible to all students.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

    cs.CY 2025-09 conditional novelty 7.0 of 10

    A new benchmark finds low to moderate human agency support in 20 LLM assistants across six dimensions.

  2. When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration

    cs.AI 2025-06 conditional novelty 7.0 of 10

    Model benchmark performance only weakly predicts how well people learn from AI explanations, with notable outliers across code and math.

  3. Experimental Evidence on the Learning Impact of Generative AI

    econ.GN 2026-07 conditional novelty 6.5 of 10

    Off-the-shelf generative AI access raises student knowledge-test scores by 0.27 SD immediately and one week later, with delayed unaided essay gains concentrated among students who use AI to explain concepts rather tha...

  4. EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

    cs.CY 2026-08 conditional novelty 6.0 of 10

    A 30-day, knowledge-tracing-grounded simulated learner benchmark for tutoring agents finds that no base model or harness alone determines quality and that almost all tested combinations plateau within days.

  5. Conversational Human Audio-visual Talking Dialogue Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.

  6. A Causal Framework for Estimating Heterogeneous Effects of On-Demand Tutoring

    cs.HC 2026-02 conditional novelty 6.0 of 10

    Requesting on-demand human tutoring raises middle-school students' next-problem correctness by about 4 percentage points on average, with large session-level heterogeneity, when estimated from observational platform l...

  7. Effective Red-Teaming of Policy-Adherent Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    A policy-aware red-teaming system (CRAFT) induces policy violations in LLM customer service agents at much higher rates than generic jailbreak prompts, using a new security-focused benchmark (tau-break) built from tau-bench.

  8. Evaluating Gemini in an arena for learning

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A blind expert-judged arena found Gemini 2.5 Pro preferred for tutoring in 73.2% of non-tied matchups against four competing models.

  9. A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

    cs.HC 2025-12 conditional novelty 5.0 of 10

    Students rated locally-run Llama-3.1 feedback roughly on par with GPT-4 and often above human TAs in technical courses, but preferred human feedback for specialized writing.

  10. Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LLMs reach only 49% F1 (MathDial) and 27% F1 (AlgebraNation) when predicting the next tutor move, while tutor moves help predict dialogue success in AlgebraNation but not consistently on MathDial.

  11. Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A large LLM-to-LLM math tutoring simulation across 11 languages shows English-language hints often yield the largest accuracy gains for student models, but the low-resource-language results lack statistical support.

  12. Human Authenticity and Flourishing in an AI-Driven World: Edmund's Journey and the Call for Mindfulness

    cs.HC 2025-05 conditional novelty 5.0 of 10

    An HCI position paper argues for a Human Flourishing Benchmark that scores AI systems on cognitive preservation, autonomy, skill development, and relational authenticity.

Pith tools