A Ranking Approach for Measuring Calibration
Anirban Chatterjee, Rina Foygel Barber
Abstract
When providing forecasted probabilities with a predictive model, the ideal model offers perfect calibration: the true probability of the outcome (i.e., the probability that Y=1) exactly matches the forecasted probability f(X). In practice, models inevitably exhibit calibration error, and it is therefore important to be able to measure this miscalibration to assess a model's reliability. The Expected Calibration Error (ECE) is the most widely used measure of miscalibration, but is known to be impossible to estimate the ECE with guaranteed accuracy in an assumption-free setting. In this work, we propose an alternative measure, the rankECE, that is based on comparing points with neighboring values of the predicted probability f(X). Our theoretical guarantees and empirical results establish that rankECE provides a better proxy for ECE as compared to binned approximations to ECE, which are the most commonly-used approximations in practice.
Create a lesson
Related papers
Design-Assisted Regression
Shangyuan Ye, Guanbo Wang, Cong Zhang et al.
Feedback-Aware Tuning of Recursive Q-Learning
Masahiro Kojima
Recoverability Is a Subspace Property: A Benchmark for Certified State Estimation from Partial PDE Observations
Qingwei Dong, Peng Zeng, Guangxi Wan et al.
Gibbs Sampling for Bayesian Generalized Poisson Matrix Factorization
Fumitake Sakaori, Hiroyasu Abe
Dynamic Amplification of Risk-Estimate Bias Through Differential Detection: A Markov Model for History-Based Covariates
Hadar Sharvit, Micha Mandel
The Anatomy and Boundary of Adaptation under Temporal Tabular Shift
Tianyu Wang, Xi Vincent Wang, Lihui Wang et al.