REVIEW 8 cited by
Distributionally Robust Optimization and Robust Statistics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We review distributionally robust optimization (DRO), a principled approach for constructing statistical estimators that hedge against the impact of deviations in the expected loss between the training and deployment environments. Many well-known estimators in statistics and machine learning (e.g. AdaBoost, LASSO, ridge regression, dropout training, etc.) are distributionally robust in a precise sense. We hope that by discussing the DRO interpretation of well-known estimators, statisticians who may not be too familiar with DRO may find a way to access the DRO literature through the bridge between classical results and their DRO equivalent formulation. On the other hand, the topic of robustness in statistics has a rich tradition associated with removing the impact of contamination. Thus, another objective of this paper is to clarify the difference between DRO and classical statistical robustness. As we will see, these are two fundamentally different philosophies leading to completely different types of estimators. In DRO, the statistician hedges against an environment shift that occurs after the decision is made; thus DRO estimators tend to be pessimistic in an adversarial setting, leading to a min-max type formulation. In classical robust statistics, the statistician seeks to correct contamination that occurred before a decision is made; thus robust statistical estimators tend to be optimistic leading to a min-min type formulation.
Forward citations
Cited by 8 Pith papers
-
Approximating Rockafellians Mitigate Distributional Perturbations: Discontinuous Integrands and Chance-Constrained Applications
Approximating Rockafellians restore convergence of stochastic programs under distributional perturbations for discontinuous integrands and general Borel measures, with quantitative rates for chance-constrained programs.
-
Distributionally Robust Shape and Topology Optimization
The paper derives tractable single-level reformulations of distributionally robust shape and topology optimization for Wasserstein, moment, and CVaR ambiguity sets, and demonstrates them numerically.
-
Doubly Smoothed Optimistic Gradients: A Universal Approach for Smooth Minimax Problems
DS-OGDA is a single-loop first-order method claimed to handle a broad class of smooth minimax problems with a universal step-size schedule and best-known iteration complexity in each subclass.
-
Unregularized limit of stochastic gradient method for Wasserstein distributionally robust optimization
Gradients of the entropically smoothed and sampled WDRO objective converge to Clarke subgradients of the unregularized objective as regularization vanishes, yielding O(log N/√N) SGD convergence rates up to sampling error.
-
DRO: A Python Library for Distributionally Robust Optimization in Machine Learning
The dro library provides a unified implementation of 14 DRO formulations across 9 model backbones, with claims of large speedups from vectorization and approximation.
-
Is Noisy Data a Blessing in Disguise? A Distributionally Robust Optimization Perspective
A proposed inverse-image Wasserstein DRO for noisy data is shown to contain a false equivalence in its reformulation, invalidating the paper's main 'blessing in disguise' result.
-
Gromov-Wasserstein and optimal transport: from assignment problems to probabilistic numeric
A largely expository paper connecting assignment problems to optimal transport and Gromov-Wasserstein distances, with a benchmark claiming a multi-start GW heuristic finds near-optimal capacitated QAP solutions; the b...
-
Data Heterogeneity Modeling for Trustworthy Machine Learning
A survey that frames heterogeneity-aware machine learning as a paradigm spanning data collection, training, evaluation, and deployment, drawing mostly on the authors' prior results.
Discussion (0). Continue with ORCID to comment.