Research

Exact Rank and Convex Calibration Dimension Lower Bounds for the Multi-Label F1 Loss

arXiv:2608.08399v1 Announce Type: new Abstract: The instance-wise F_1 measure is a central performance measure for multi-label classification. For a problem with s labels, it defines a 2^simes 2^s los

DGX agentpaper
researcharxiv-cs-lg

arXiv:2608.08399v1 Announce Type: new Abstract: The instance-wise F_1 measure is a central performance measure for multi-label classification. For a problem with s labels, it defines a 2^simes 2^s loss matrix. Previous work exhibited s^2+1-coordinate affine and shifted low-rank representations and used them to construct quadratic-dimensional convex calibrated surrogates. We determine the exact rank. Under the convention F_1(arnothing,arnothing)=1, the F_1 score matrix, the shifted loss matrix, and the unshifted loss matrix all have rank s^2-s+2, while the column-affine dimension of the loss is s^2-s+1. The proof factors the nonempty score matrix through subset-incidence matrices and a positive-definite Cauchy matrix. Exact rank does not, by itself, lower-bound the dimension of an arbitrary convex calibrated surrogate. We therefore analyze the Bayes geometry of F_1 directly. We construct a distribution for which precisely all supersets of a fixed core label set are Bayes optimal, and show that the corresponding active loss columns, restricted to the witness support, have affine dimension hn, where n=s-lfloor s/3rfloor and h=lceil(slfloor s/3rfloor)^{1/2}rceil-1. Applying the feasible-subspace lower bound for convex calibration dimension gives [ operatorname{CCdim}(L^{F_1}) ge left(frac{2}{3sqrt{3}}-o(1)right)s^2. ] Together with the quadratic upper bound, this establishes operatorname{CCdim}(L^{F_1})=Theta(s^2).

Source: arXiv cs.LG | 2026-08-11

Loading related sources…