Skip to content

Exact Recovery Thresholds for Weighted Data Selection in Vector-Valued Linear Regression

Guangjian Zhang

cs.LGarXiv:2608.30254

Abstract

We resolve the threshold part of Question 4 of the COLT 2025 open problem &#34;Data Selection for Regression Tasks&#34; of Hanneke, Moran, Shlimovich and Yehudayoff. In vector-valued linear regression with square loss (x,y)(W)=|Wx-y|22, where x∈Rd, y∈Rm and the learner is the empirical risk minimizer of minimal Frobenius norm, we prove that the minimal budget of weighted examples that recovers the full-data loss on every finite dataset is exactly n*(d,m)=(m+1)d. We further determine two more values of the weighted selection profile Fw(d,m,n): at the near-threshold budget, Fw(d,m,(m+1)d-1)=1+1dm2, and at the spanning budget, Fw(d,m,d)=d+1 for every m, while Fw(d,m,n)=∞ for n<d. For the smallest open intermediate cell (d,m)=(2,2) we prove Fw(2,2,3)∈[13/8,15/8] and Fw(2,2,4)∈[5/4,3/2], reduce the conjectured exact values 13/8 and 5/4 to a finite moment problem on the circle with at most seven atoms, and establish strong structural evidence for the conjecture. The upper-bound techniques (a fixed-basis conic compression lemma, a determinant-facet rigidity theorem for maximal certificates, and sharp sparsification lemmas for zero-mean weighted point systems) are of independent interest. As a byproduct we correct an erroneous claim circulating in a recent unrefereed preprint, exhibiting an explicit dataset with m=2 on which no weighted selection of 2d points recovers the optimal loss. All results are new only for m 2; the scalar case m=1 is due to Hanneke et al.

Create a lesson