Skip to content

OSCAR: Order-aware Scoring and Calibration for AI Rankings

You Liu, Yue Liu, Quanchao Lu, Nick Shipilov

stat.MLarXiv:2609.24128

Abstract

Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating AI rankings, and study position as one such effect. In released judgments from 18 evaluators, the all-response A-minus-B score difference ranges from -63.11 to 98.31 percentage points. Matching question text, response texts, candidate identities, and judge within the released table gives an overall difference of 24.22 points (95% interval [22.90,25.54]), conditional on the released text mapping. A controlled calculation isolates the potential consequence: with true sensitivity fixed at one, omitting a position intercept of four reduces the population-optimal slope to 0.0771. We extend sensitivity-based ranking with judge-specific position, length, and family terms, characterize local omission-induced displacement and an identification failure, and propagate prompt-cluster uncertainty to adjusted comparisons. Across four released datasets, position provides the largest stand-alone predictive improvement. Refitting bootstrap comparisons show more selective gains from the full model over position-only adjustment. In dependent binary simulations, adjusting both the mean and covariance yields 94.4--95.2% coverage; correcting either alone is insufficient. At N=10,000, OSCAR reduces mean neutral-target RMSE from 0.1158 under the sensitivity-only model to 0.0237.

Create a lesson