Pith. sign in

REVIEW 19 cited by

LearnLM: Improving Gemini for Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.16429 v3 pith:IXCFBIQK submitted 2024-12-21 cs.CY cs.AIcs.LG

LearnLM: Improving Gemini for Learning

classification cs.CY cs.AIcs.LG
keywords learningmodelpedagogicalgeminilearnlmbehaviordesiredfollowing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Today's generative AI systems are tuned to present information by default, rather than engage users in service of learning as a human tutor would. To address the wide range of potential education use cases for these systems, we reframe the challenge of injecting pedagogical behavior as one of \textit{pedagogical instruction following}, where training and evaluation examples include system-level instructions describing the specific pedagogy attributes present or desired in subsequent model turns. This framing avoids committing our models to any particular definition of pedagogy, and instead allows teachers or developers to specify desired model behavior. It also clears a path to improving Gemini models for learning -- by enabling the addition of our pedagogical data to post-training mixtures -- alongside their rapidly expanding set of capabilities. Both represent important changes from our initial tech report. We show how training with pedagogical instruction following produces a LearnLM model (available on Google AI Studio) that experts substantially prefer across a diverse set of learning scenarios, with average preference strengths of +31\% over GPT-4o, +11\% over Claude 3.5 Sonnet, and +13\% over the Gemini 1.5 Pro model on which LearnLM was based.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Curiosity as Linguistic Intervention: Using LLM Tutoring Dialogues to Influence Exploratory Learning Behavior

    cs.CL 2026-06 unverdicted novelty 7.0

    Curiosity-oriented linguistic interventions in LLM tutoring dialogues increased exploratory learner behaviors up to 2.4x across 270 conversations spanning multiple models and domains.

  2. Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild

    cs.CL 2026-06 unverdicted novelty 7.0

    A PPO policy for deciding topic order and duration on a prerequisite knowledge graph, paired with an LLM for Socratic dialogue, improves student mastery rates and reduces turns compared to baselines and scaled models ...

  3. Experimental Evidence on the Learning Impact of Generative AI

    econ.GN 2026-07 conditional novelty 6.5

    Off-the-shelf generative AI access raises student knowledge-test scores by 0.27 SD immediately and one week later, with delayed unaided essay gains concentrated among students who use AI to explain concepts rather tha...

  4. Auditable Release Control for Pedagogical Leakage in LLM Tutors

    cs.CR 2026-08 conditional novelty 6.0

    A release gate with deterministic fallback reduces unauthorized answer leakage in LLM tutors and makes failure attribution replayable, but it costs helpfulness and does not improve learning.

  5. The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals

    cs.CY 2026-05 unverdicted novelty 6.0

    The Tutoring Effectiveness Index (TEI) uses four signals from LLM conversations to select math tutoring responses, raising student improvement rates from 59.0% to 81.9% at N=8 on a frozen DeepSeek-R1-8B model without ...

  6. LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring

    cs.CL 2026-05 unverdicted novelty 6.0

    Training-free prompt optimization methods, including five new education-focused ones, surpass the strongest RL-trained baseline across five conditions on two OOD suites while showing distinct teaching behavior patterns.

  7. TeachArena: Are Language Agents Ready for Realistic Teaching Work?

    cs.AI 2026-05 unverdicted novelty 6.0

    EduAgentBench is a new source-grounded benchmark that evaluates tutor agents across pedagogical judgment, situated multi-turn tutoring, and Canvas-style workflow completion, finding frontier models capable of basic ju...

  8. TeachArena: Are Language Agents Ready for Realistic Teaching Work?

    cs.AI 2026-05 conditional novelty 6.0

    A single three-surface benchmark shows frontier LLM agents pass ≈90% of bounded pedagogical-judgment tasks but only ≈36% of situated tutoring and ≈33% of LMS workflow tasks.

  9. Behavior Latticing: Inferring User Motivations from Unstructured Interactions

    cs.HC 2026-04 unverdicted novelty 6.0

    Behavior latticing synthesizes connections across unstructured user interactions to generate insights into underlying motivations, yielding deeper and more accurate user understanding than task-only models.

  10. Mitigating LLM biases toward spurious social contexts using direct preference optimization

    cs.AI 2026-04 unverdicted novelty 6.0

    Debiasing-DPO reduces bias to spurious social contexts by 84% and improves predictive accuracy by 52% on average for LLMs evaluating U.S. classroom transcripts.

  11. Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

    cs.CL 2026-02 unverdicted novelty 6.0

    LLMs simulating student think-alouds in multi-step chemistry tutoring produce overly coherent, verbose, and confident reasoning that overestimates learner success compared to 630 human utterances.

  12. LLMs cannot spot math errors, even when allowed to peek into the solution

    cs.CL 2025-09 conditional novelty 6.0

    Even with the gold solution in hand, large language models locate the first error step in student math solutions poorly; providing a generated corrected student solution improves accuracy.

  13. ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

    cs.HC 2026-06 conditional novelty 5.0

    A Vietnamese education LLM (Qwen3-8B + SFT/DPO) with an ADDIE-style lesson-planning workflow reports 87% exam accuracy and 30–45 min prep, with small satisfaction surveys.

  14. Reinforcement Learning for Special Education: Aligning LLM Tutors to Diverse Learners through Disability-Adaptive Training

    cs.CY 2026-05 unverdicted novelty 5.0

    Special-R1 combines two-dimensional adaptive prompts and a disability-conditioned Thinking Reward in RL training, lifting persona-aware Fit by 1.65 and SPED Helpfulness by 0.048 on a 690-dialogue test set while stayin...

  15. IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations

    cs.CL 2025-09 conditional novelty 5.0

    IDEAlign uses a pick-the-odd-one-out triplet task to measure idea-level similarity between LLMs and expert human annotations, and shows LLM judges using this protocol outperform lexical and vector-based baselines.

  16. AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations

    cs.AI 2026-07 conditional novelty 4.0

    A demo audits K-12 explanations for five pedagogical risks with localized evidence and rationales, reporting that a locally fine-tuned Llama-3.1-8B beats GPT-5.5 on most metrics—on a benchmark the authors built themselves.

  17. Uncertainty-Aware Generation and Decision-Making Under Ambiguity

    cs.CL 2026-06 unverdicted novelty 4.0

    Uncertainty-aware algorithms based on Bayesian decision theory improve generation utility on tutoring and reviewing tasks while risk-averse methods can degrade performance under high ambiguity, with conformal predicti...

  18. Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems

    cs.HC 2026-04 unverdicted novelty 4.0

    AI explanations in language learning often fail across six dimensions like diagnostic accuracy and self-regulation support, creating hidden risks that demand better evaluation frameworks such as L2-Bench.

  19. Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems

    cs.HC 2026-04 unverdicted novelty 4.0

    Introduces L2-Bench benchmark for AI feedback in language education across six dimensions and identifies explainability pitfalls in AI-generated explanations that appear helpful but are flawed.