Distilling Black-Box Machine Learning into a Small, Self-Explaining Language Model for Learning Analytics
Chenguang Pan, Airui Meng, Youmi Suk
Abstract
Learning analytics increasingly relies on flexible machine learning (ML), but the model opacity and the burden of deployment prevent these tools from reaching educational practice. We propose a two-stage fine-tuning pipeline that distills a fitted black-box estimator and its post hoc interpretation (the mentor) into a small, open-weight large language model (LLM; the mentee) that returns an individual-level estimate and explains in natural language. The design is estimator-agnostic and paired with a faithfulness-first evaluation framework that audits every narration against the attribution it claims to describe. We design a simulation study that separates distillation loss from estimator loss by comparing an oracle mentor with a realistic ML mentor. Given an oracle signal, distillation with a two-billion-parameter LLM model is nearly lossless in recovering the effect surface (r > .90), perfectly ranking the important variables, and citing no spurious covariate. Under a realistic estimator, almost all remaining error originates upstream. We find that fluency is no evidence of correctness since narration quality is independent of signal quality, and decision quality collapses toward the majority action in severely imbalanced settings. Applied to a nationally representative dataset, the pipeline recovers the finding that advanced mathematics coursework benefits students least likely to enroll in four-year college the most, with 98.8% of narrations passing the audit and no fabricated quantities. The result is a single fine-tuned LLM that predicts and explains offline on a commodity laptop, so student records never leave the machine.
Create a lesson
Related papers
Calmables: Demonstrating Closed-Loop Infrared Earables for Thermal Biofeedback and Relaxation Support
Valeria Zitz, Michael Küttner, Jonas Hummel et al.
"Okay, I've Actually Softened My Take on This": How People in Decentralized Social Media Reason about the Appropriateness of Generative AI
Romina Mahinpei, Manoel Horta Ribeiro, Andrés Monroy-Hernández et al.
Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course
Hong Qu, Zichao Ling, Yadie Yang
EasyFashion: A Human-AI Co-Creation System for Personalized Fashion Design and Sewing Pattern Generation
Hong Qu, Zhaoxiang Xu, Jinbo Luo et al.
Verify, Offload, Extend & Recommend: Selective Complementarity in AI Support for Physical Activity Planning with Longitudinal Patient Data
Pavithren V S Pakianathan, Rania Islambouli, Diogo Branco et al.
Building a Cultural Perspective on Doctor-Patient Conversations
Krithi Shailya, Siddharth D Jaiswal, Ashish Makani et al.