Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
Adam Coscia, Sujata Duwal, Langdon Holmes, Scott Crossley, Alex Endert
Abstract
LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we: (1) co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; (2) designed visualizations that map trustworthiness violations onto LLM responses; and (3) evaluated how these tools help learning engineers make A/B comparisons of LLM responses. Making trustworthiness explicit increased inter-rater reliability while helping learning engineers resolve conflicting objectives and produce more consistent judgments. We discuss the emergent benefits of trustworthiness as a lens for evaluating LLMs in education and propose new design guidelines for future evaluation tools that enable pedagogically-aligned, LLM-powered learning tools.
Create a lesson
Related papers
Calmables: Demonstrating Closed-Loop Infrared Earables for Thermal Biofeedback and Relaxation Support
Valeria Zitz, Michael Küttner, Jonas Hummel et al.
"Okay, I've Actually Softened My Take on This": How People in Decentralized Social Media Reason about the Appropriateness of Generative AI
Romina Mahinpei, Manoel Horta Ribeiro, Andrés Monroy-Hernández et al.
Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course
Hong Qu, Zichao Ling, Yadie Yang
EasyFashion: A Human-AI Co-Creation System for Personalized Fashion Design and Sewing Pattern Generation
Hong Qu, Zhaoxiang Xu, Jinbo Luo et al.
Verify, Offload, Extend & Recommend: Selective Complementarity in AI Support for Physical Activity Planning with Longitudinal Patient Data
Pavithren V S Pakianathan, Rania Islambouli, Diogo Branco et al.
Building a Cultural Perspective on Doctor-Patient Conversations
Krithi Shailya, Siddharth D Jaiswal, Ashish Makani et al.