Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
Binyan Xu, Xilin Dai, Fan Yang, Kehuan Zhang
Abstract
Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021--2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper's acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.
Create a lesson
Related papers
Shifting Research Funding Priorities under Geopolitical Pressure: Evidence from Estonia
Yunfeng Gao, Yang Ding
Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
Daniel Ebanks, Devika Jain
Quantifying the impact of clinical-academic collaborations
Mohamad Zeina, Nick McNally, Karl S. Peggs et al.
Testing Our Foundations: Citation Trends, Errors, and Emerging Hallucinations in the Computing Education Literature
Paul Denny, Gweneth Barbre, Musa Blake et al.
Toward non-textual representation of social anthropology: Modeling cultures as knowledge graphs
Manolis Peponakis, Sarantos Kapidakis, Martin Doerr et al.
Geometric Signatures of Conceptual Reorganization: A Counterfactual Embedding Framework for Detecting Scientific Revolutions
Dimitris Ntounis, Ariel Schwartzman, Chris Chafe et al.