Applications
Contrasting Cost-Agnostic and Cost-Sensitive Losses under Limited Model Capacity via mathcal H-consistency
arXiv:2502.19522v2 Announce Type: replace Abstract: There is a prevalent debate in machine learning about whether practitioners should train models to optimize a task-agnostic objective (e.g., cross e
arXiv:2502.19522v2 Announce Type: replace Abstract: There is a prevalent debate in machine learning about whether practitioners should train models to optimize a task-agnostic objective (e.g., cross entropy) or incorporate the downstream decision task into the optimization objective (e.g., weighted cross entropy). In ideal settings, like those with infinite data and infinite model capacity, the two approaches are statistically equivalent for the downstream decision task. In practice, however, incorporating the decision task into model training has been shown to empirically improve task-specific performance in certain real-world scenarios. The cause of these benefits has not been theoretically studied to date. Focusing on the setting with limited model capacity through the model class mathcal H, we establish a strict performance gap between post-processing a model learned with a cost-agnostic objective (e.g., thresholding a risk score prediction) and models learned by optimizing a cost-sensitive model without a threshold search. In particular, we establish this gap when there is a hypothesis recovering the optimal decision boundary for the discrete task, but it does not align with the optimal cost-agnostic hypothesis, and give a simple example demonstrating the plausibility of this setting. With assumptions that are hard to verify in practice, we demonstrate that this gap generally exists on classification datasets from the UCI repository, particularly with very simple models.
Related
- A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning
- The Cost of Learning Under Multiple Change Points
- Metric-agnostic Learning-to-Rank via Boosting and Rank Approximation
Source: arXiv cs.LG | 2026-08-20