Pith. sign in

REVIEW

Contrasting Cost-Agnostic and Cost-Sensitive Losses under Limited Model Capacity via $\mathcal H$-consistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.19522 v2 pith:MEVMT7YW submitted 2025-02-26 cs.LG

classification cs.LG
keywords modeldecisiontaskcapacitycost-agnosticmodelsobjectivebeen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

There is a prevalent debate in machine learning about whether practitioners should train models to optimize a task-agnostic objective (e.g., cross entropy) or incorporate the downstream decision task into the optimization objective (e.g., weighted cross entropy). In ideal settings, like those with infinite data and infinite model capacity, the two approaches are statistically equivalent for the downstream decision task. In practice, however, incorporating the decision task into model training has been shown to empirically improve task-specific performance in certain real-world scenarios. The cause of these benefits has not been theoretically studied to date. Focusing on the setting with limited model capacity through the model class $\mathcal H$, we establish a strict performance gap between post-processing a model learned with a cost-agnostic objective (e.g., thresholding a risk score prediction) and models learned by optimizing a cost-sensitive model without a threshold search. In particular, we establish this gap when there is a hypothesis recovering the optimal decision boundary for the discrete task, but it does not align with the optimal cost-agnostic hypothesis, and give a simple example demonstrating the plausibility of this setting. With assumptions that are hard to verify in practice, we demonstrate that this gap generally exists on classification datasets from the UCI repository, particularly with very simple models.

Discussion (0). Continue with ORCID to comment.

Pith tools