Skip to content

Exact Risk Ratios for Weighted Data Selection in Linear Regression

Guangjian Zhang

cs.LGarXiv:2608.28007

Abstract

Hanneke, Moran, Shlimovich and Yehudayoff (COLT 2025) posed the following open problem. A selector sees a finite dataset D ⊂eq Rd × R, picks at most n examples together with nonnegative weights, and hands the weighted least squares objective to the minimum-norm ERM. Writing Fw(d,n) for the worst-case ratio between the loss of the returned predictor on all of D and the optimal loss, they proved Fw(d,n)=∞ for n<d, Fw(d,d)=d+1 and Fw(d,n)=1 for n 2d, and asked for the value in the open regime d<n<2d. We determine this value in several cases. For every d we prove Fw(d,2d-1)=1+1/d, which confirms a claim stated without proof in the original note. We further prove Fw(3,4)=5/3 and Fw(4,5)=2, the two smallest cells not covered by the endpoint formula. For every intermediate budget n=d+k we prove the lower bound Fw(d,d+k) 1+Γd,k, where Γd,k is an explicit harmonic quantity over balanced partitions, and we show that this bound is the exact minimax value over the class of datasets whose whitened gradient systems carry an orthogonal circuit-block structure. All three exact values match 1+Γd,k, and we conjecture that equality holds throughout the open regime. The upper bound proofs run on a common geometric spine: a rigidity theorem for positive spanning configurations of loss gradients, classifications and structural reductions of small positive bases in R3 and R4, and a dimension-free extremal-basis argument that converts sign-cone geometry into five-point selections. We also give explicit counterexamples showing that several shorter routes fail, and constructive polynomial-time selection algorithms for all proved cases.

Create a lesson