The work gives the first algorithms for general robust Markov games with linear function approximation whose sample complexity breaks the curse of multiagency for large state spaces in both generative and online settings.
hub
Distributionally robust optimization: A review
23 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
other 1polarities
unclear 1representative citing papers
Introduces PowerPhase benchmark for massive-variate power-system forecasting and PowerForge model that achieves best average rank on safety-fidelity metrics across all tested grids.
Conformal Risk Sharing combines an interpretable sharing policy tuned on training data with split conformal calibration on held-out data to produce certified obligation caps and bounded aggregate harm without distributional assumptions.
Action-conditional conformal prediction sets provide per-action safety guarantees for risk-averse policies that optimize conditional value-at-risk through pinball-loss minimization.
CorrDP relaxes standard differential privacy by incorporating feature correlations, enabling distance-dependent noise in DP-ERM for better privacy-utility tradeoffs.
Presents the first algorithm to identify an ε-optimal policy in robust constrained MDPs via epigraph form and bisection search with Õ(ε^{-4}) robust policy evaluations.
Proposes APUB optimization framework for stochastic programming, proves asymptotic correctness and consistency of the new bound, and develops bootstrap and L-shaped solvers for two-stage linear problems with empirical tests on a product mix example.
A systematic review of over 200 studies concludes that LLMs in recommender systems act as a double-edged sword, creating both opportunities and new risks for trustworthiness.
SW-DRSO optimizes a tractable surrogate of worst-case expected loss over plausible inference-time corruptions using a barycentric adversary approximated via simplex weights.
CVaR-constrained TD3 policies for robot navigation show larger safety margins and higher post-training reachability verification rates than average-cost baselines across simulated scenarios and real-robot tests.
Expected regret equals covariance between costs and optimal decisions for linear and quadratic stochastic programs, with explicit bounds on the residual.
Learned primal and dual maps conditioned on population summaries enable reliable coordination across composition shifts in large multi-agent systems, cutting forecast error 16-19% and violations 20-51% in a supply-chain case study.
RACER routes between reasoning and non-reasoning LLM judges via constrained distributionally robust optimization to achieve better accuracy-cost trade-offs under distribution shift.
Q-MMR introduces recursive reweighting and moment matching for off-policy evaluation, delivering dimension-free error bounds under Q^π realizability alone.
The authors create a distributionally robust formulation for the cyclic inventory routing problem that admits a deterministic reformulation via multi-point worst-case distributions and chance-constraint equivalents, solved by nested branch-and-price and tested on real automotive data.
The authors introduce (ηx,ηy,δ,ε)-GSSP as a convergence criterion and develop projected gradient-free descent-ascent methods achieving non-asymptotic rates for nonsmooth nonconvex-concave minimax optimization without weak convexity assumptions.
A new disturbance-affine distributionally robust MPC framework for uncertain linear systems that is less conservative than tube-based approaches while guaranteeing recursive feasibility and stability.
PAC learning-based DR-MPC framework interpolates between robust MPC and stochastic MPC for interactive trajectory planning under agent decision uncertainty.
The authors develop a conceptual framework for assured autonomy in generative AI by using flow-based models for auditable generation and adversarial robustness for operational safety, repositioning operations research as a system architect.
PECO strengthens chance constraints by mandating feasibility for all high-probability events and is solved via a data-embedded deterministic program that works for nonlinear nonconvex instances when the size of the solution-determining data family can be estimated by machine learning.
A target-based DRO model for MST under distributional uncertainty is solved exactly via Benders decomposition and a modified Prim algorithm.
citing papers explorer
-
Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation
The work gives the first algorithms for general robust Markov games with linear function approximation whose sample complexity breaks the curse of multiagency for large state spaces in both generative and online settings.
-
Navigating the Safety-Fidelity Trade-off: Massive-Variate Time Series Forecasting for Power Systems via Probabilistic Scenarios
Introduces PowerPhase benchmark for massive-variate power-system forecasting and PowerForge model that achieves best average rank on safety-fidelity metrics across all tested grids.
-
Conformal Risk Sharing: Certified Cost Allocation with Participation Guarantees
Conformal Risk Sharing combines an interpretable sharing policy tuned on training data with split conformal calibration on held-out data to produce certified obligation caps and bounded aggregate harm without distributional assumptions.
-
Conformal Risk-Averse Decision Making with Action Conditional Guarantee
Action-conditional conformal prediction sets provide per-action safety guarantees for risk-averse policies that optimize conditional value-at-risk through pinball-loss minimization.
-
Integrating Feature Correlation in Differential Privacy with Applications in DP-ERM
CorrDP relaxes standard differential privacy by incorporating feature correlations, enabling distance-dependent noise in DP-ERM for better privacy-utility tradeoffs.
-
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
Presents the first algorithm to identify an ε-optimal policy in robust constrained MDPs via epigraph form and bisection search with Õ(ε^{-4}) robust policy evaluations.
-
Minimizing Upper Confidence Bounds: A Data-Driven Framework for Stochastic Programming
Proposes APUB optimization framework for stochastic programming, proves asymptotic correctness and consistency of the new bound, and develops bootstrap and L-shaped solvers for two-stage linear problems with empirical tests on a product mix example.
-
Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges
A systematic review of over 200 studies concludes that LLMs in recommender systems act as a double-edged sword, creating both opportunities and new risks for trustworthiness.
-
Distributionally Robust Set Representation Learning Under Inference-Time Element Corruption
SW-DRSO optimizes a tractable surrogate of worst-case expected loss over plausible inference-time corruptions using a barycentric adversary approximated via simplex weights.
-
Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation
CVaR-constrained TD3 policies for robot navigation show larger safety margins and higher post-training reachability verification rates than average-cost baselines across simulated scenarios and real-robot tests.
-
Regret Equals Covariance: A Closed-Form Characterization for Stochastic Optimization
Expected regret equals covariance between costs and optimal decisions for linear and quadratic stochastic programs, with explicit bounds on the residual.
-
Ready from Day 1: Population-Aware Coordination for Large-Scale Constrained Multi-Agent Systems
Learned primal and dual maps conditioned on population summaries enable reliable coordination across composition shifts in large multi-agent systems, cutting forecast error 16-19% and violations 20-51% in a supply-chain case study.
-
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
RACER routes between reasoning and non-reasoning LLM judges via constrained distributionally robust optimization to achieve better accuracy-cost trade-offs under distribution shift.
-
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
Q-MMR introduces recursive reweighting and moment matching for off-policy evaluation, delivering dimension-free error bounds under Q^π realizability alone.
-
The Distributionally Robust Cyclic Inventory Routing Problem
The authors create a distributionally robust formulation for the cyclic inventory routing problem that admits a deterministic reformulation via multi-point worst-case distributions and chance-constraint equivalents, solved by nested branch-and-price and tested on real automotive data.
-
Nonsmooth Nonconvex-Concave Minimax Optimization: Convergence Criteria and Algorithms
The authors introduce (ηx,ηy,δ,ε)-GSSP as a convergence criterion and develop projected gradient-free descent-ascent methods achieving non-asymptotic rates for nonsmooth nonconvex-concave minimax optimization without weak convexity assumptions.
-
Distributionally Robust Stochastic MPC under Disturbance-Affine Feedback Policies
A new disturbance-affine distributionally robust MPC framework for uncertain linear systems that is less conservative than tube-based approaches while guaranteeing recursive feasibility and stability.
-
Interactive Trajectory Planning with Learning-based Distributionally Robust Model Predictive Control and Markov Systems
PAC learning-based DR-MPC framework interpolates between robust MPC and stochastic MPC for interactive trajectory planning under agent decision uncertainty.
-
Assured autonomy: How operations research powers and orchestrates generative AI systems
The authors develop a conceptual framework for assured autonomy in generative AI by using flow-based models for auditable generation and adversarial robustness for operational safety, repositioning operations research as a system architect.
-
A Data-embedded Solution Paradigm for Nonconvex Probable Event Constrained Optimization
PECO strengthens chance constraints by mandating feasibility for all high-probability events and is solved via a data-embedded deterministic program that works for nonlinear nonconvex instances when the size of the solution-determining data family can be estimated by machine learning.
-
Target-based Distributionally Robust Minimum Spanning Tree Problem
A target-based DRO model for MST under distributional uncertainty is solved exactly via Benders decomposition and a modified Prim algorithm.
- Distributionally Robust Safety Under Arbitrary Uncertainties: A Safety Filtering Approach
- Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback