Sherpa: Teaching LLMs to Teach Adaptively
Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang
Abstract
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
Create a lesson
Related papers
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Zewei Zhou, Rachel Luo, Yulong Cao et al.
Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Egor Pakhomov, Erik Nijkamp
WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
Siru Jiang, Yongzhe Lyu, Shuo Lu et al.
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Guangyi Liu, Yong Liu, Jiangning Zhang
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
Mingda Zhang, Wenjin Liu, Tiesunlong Shen et al.