OSCAR: Order-aware Scoring and Calibration for AI Rankings
You Liu, Yue Liu, Quanchao Lu, Nick Shipilov
Abstract
Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating AI rankings, and study position as one such effect. In released judgments from 18 evaluators, the all-response A-minus-B score difference ranges from -63.11 to 98.31 percentage points. Matching question text, response texts, candidate identities, and judge within the released table gives an overall difference of 24.22 points (95% interval [22.90,25.54]), conditional on the released text mapping. A controlled calculation isolates the potential consequence: with true sensitivity fixed at one, omitting a position intercept of four reduces the population-optimal slope to 0.0771. We extend sensitivity-based ranking with judge-specific position, length, and family terms, characterize local omission-induced displacement and an identification failure, and propagate prompt-cluster uncertainty to adjusted comparisons. Across four released datasets, position provides the largest stand-alone predictive improvement. Refitting bootstrap comparisons show more selective gains from the full model over position-only adjustment. In dependent binary simulations, adjusting both the mean and covariance yields 94.4--95.2% coverage; correcting either alone is insufficient. At N=10,000, OSCAR reduces mean neutral-target RMSE from 0.1158 under the sensitivity-only model to 0.0237.
Create a lesson
Related papers
Empirical Auditing of Edge-Private Graph Generators
Anum Fatima, Stratis Limnios, James Adams et al.
Variational objectives for amortized Bayesian inference in inverse problems: The role of posterior conditioning
Abhishek Srivastava, Arijit Hazra, Rajesh Dubbaku
JAREX: An Acquisition Function for Multi-Objective Algorithmic Process Characterization
Xinyang Li, Kevin Stone, Ajit Vikram
Identifying Representational Biases in Datasets Using PCA: A Max-Disparity Partition Framework
Arjun KM, Shashi Jain
Beyond Point Prediction: Artificial Representative Trees with Uncertainty
Lea L. Mairhöfer, Silke Szymczak, Björn-Hergen Laabs et al.
Adversarially Robust PAC Learning with Optimal VC Rates
Steve Hanneke, Amirreza Shaeiri