Pith. sign in

REVIEW 12 cited by

Certifying Some Distributional Robustness with Principled Adversarial Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.10571 v5 pith:P5MCRTQ3 submitted 2017-10-29 stat.ML cs.LG

classification stat.MLcs.LG
keywords adversarialperturbationsrobustnesstrainingdataguaranteesheuristicprincipled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural networks are vulnerable to adversarial examples and researchers have proposed many heuristic attack and defense mechanisms. We address this problem through the principled lens of distributionally robust optimization, which guarantees performance under adversarial input perturbations. By considering a Lagrangian penalty formulation of perturbing the underlying data distribution in a Wasserstein ball, we provide a training procedure that augments model parameter updates with worst-case perturbations of training data. For smooth losses, our procedure provably achieves moderate levels of robustness with little computational or statistical cost relative to empirical risk minimization. Furthermore, our statistical guarantees allow us to efficiently certify robustness for the population loss. For imperceptible perturbations, our method matches or outperforms heuristic approaches.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 346 citations worldwide. Full citation record

  1. A first-order method for constrained nonconvex-nonconcave minimax optimization

    math.OC 2025-10 conditional novelty 6.0 of 10

    Under a local Kurdyka-Łojasiewicz condition, the constrained nonconvex-nonconcave minimax value function is locally generalized Hölder smooth, and an interleaved SCP/proximal-gradient method achieves Õ(ε^{−max{1/(1−θ)...

  2. An Optimistic Gradient Tracking Method for Distributed Minimax Optimization

    math.OC 2025-08 conditional novelty 6.0 of 10

    DOGT and its accelerated variant ADOGT achieve optimal communication complexity O(κ log(1/ε)/√(1-√ρ_W)) for distributed strongly convex-strongly concave minimax optimization over networks.

  3. A New Perspective On AI Safety Through Control Theory Methodologies

    cs.AI 2025-06 conditional novelty 6.0 of 10

    This paper outlines a new conceptual paradigm, data control, which transfers control-theoretic system analysis and properties to AI systems to support generic AI safety assurance.

  4. When Shift Happens - Confounding Is to Blame

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Under hidden confounding shifts, predictive information reduces to conditional informativeness minus a residual, a result the authors use to explain ERM's surprising OOD competitiveness and the value of all-covariate models.

  5. Bigger Is Safer: Provable Robustness in In-Context Learning Scales with Capacity

    cs.LG 2026-02 reject novelty 5.0 of 10

    For linear self-attention Transformers, the paper claims worst-case risk under Wasserstein adversarial shifts is bounded by L0 + C1 ρ√(d/m) + C2 ρ²/√N, giving ρmax∝√m and Nρ−N0∝ρ², but the m-dependence is asserted rat...

  6. Gradient Flow Sampler-based Distributionally Robust Optimization

    math.OC 2025-10 conditional novelty 5.0 of 10

    Entropy-regularized Wasserstein DRO can be solved by sampling from a Gibbs worst-case distribution with gradient-flow samplers, giving new WFR/SVGD algorithms and a principled recovery of WRM.

  7. Improved Stochastic Optimization of LogSumExp

    math.OC 2025-09 conditional novelty 5.0 of 10

    A rescaled SoftPlus family approximates LogSumExp with O(ρ) error, enabling stable stochastic optimization in entropic OT and KL-DRO.

  8. Robustifying Diffusion-Denoised Smoothing Against Covariate Shift

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Adversarially perturbing the noise term of a diffusion denoiser during training improves the certified l2 robustness of denoised randomized smoothing on MNIST, CIFAR-10, and ImageNet, with the largest gains at large p...

  9. Group Distributionally Robust Machine Learning under Group Level Distributional Uncertainty

    cs.LG 2025-09 reject novelty 5.0 of 10

    A min-max-sup extension of Group DRO that adds a Wasserstein ball around each group's empirical distribution, with a descent-mirror-ascent algorithm and Adult income experiments.

  10. FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FOCoOp uses global, local, and OOD prompts with bi-level distributionally robust optimization and semi-unbalanced optimal transport to improve OOD robustness in federated prompt learning.

  11. Distributionally Robust and Safe Imitation Learning

    cs.LG 2026-07 reject novelty 4.0 of 10

    A distributionally robust and safe imitation-learning loss is proposed by adding a Wasserstein-ambiguity-set objective and a CVaR safety penalty to TaSIL, validated only by a qualitative UAV simulation.

  12. Understanding Knowledge Transferability for Transfer Learning: A Survey

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A survey that classifies transferability metrics by knowledge modality (dataset vs. model) and granularity (task vs. instance), with a theoretical primer and applications to eight learning paradigms.

Pith tools