REVIEW 5 cited by
Omnipredictors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Loss minimization is a dominant paradigm in machine learning, where a predictor is trained to minimize some loss function that depends on an uncertain event (e.g., "will it rain tomorrow?''). Different loss functions imply different learning algorithms and, at times, very different predictors. While widespread and appealing, a clear drawback of this approach is that the loss function may not be known at the time of learning, requiring the algorithm to use a best-guess loss function. We suggest a rigorous new paradigm for loss minimization in machine learning where the loss function can be ignored at the time of learning and only be taken into account when deciding an action. We introduce the notion of an (${\mathcal{L}},\mathcal{C}$)-omnipredictor, which could be used to optimize any loss in a family ${\mathcal{L}}$. Once the loss function is set, the outputs of the predictor can be post-processed (a simple univariate data-independent transformation of individual predictions) to do well compared with any hypothesis from the class $\mathcal{C}$. The post processing is essentially what one would perform if the outputs of the predictor were true probabilities of the uncertain events. In a sense, omnipredictors extract all the predictive power from the class $\mathcal{C}$, irrespective of the loss function in $\mathcal{L}$. We show that such "loss-oblivious'' learning is feasible through a connection to multicalibration, a notion introduced in the context of algorithmic fairness. In addition, we show how multicalibration can be viewed as a solution concept for agnostic boosting, shedding new light on past results. Finally, we transfer our insights back to the context of algorithmic fairness by providing omnipredictors for multi-group loss minimization.
Forward citations
Cited by 5 Pith papers
-
Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening
If an optimal (even intractable) protocol achieves utility α in k bits, a polynomial-time algorithm can find a protocol achieving α−ε using 2^{O(k)}/ε^2 bits, and this is tight up to a constant in the exponent.
-
Representative Language Generation
A new 'representative generation' requirement is formalized, characterized by a group closure dimension, with a feasibility result under finite support and a membership-query impossibility.
-
Discretization-free Multicalibration through Loss Minimization over Tree Ensembles
A one-shot ERM over depth-two tree ensembles on the base predictor and group indicators yields multicalibration whenever squared loss is saturated, a condition verified empirically on six datasets.
-
Persuasive Prediction via Decision Calibration
A data-driven sender can learn a near-optimal decision-calibrated predictor without knowing the prior, but the proof as written has a critical Lagrangian error and the Bayesian benchmark is restricted by construction.
-
Calibration through the Lens of Indistinguishability
A survey arguing that calibration error is best understood as the degree to which two worlds, the predictor's and nature's, can be distinguished, and that this view unifies ECE, smooth calibration, CDL, and distance t...
Discussion (0). Continue with ORCID to comment.