Pith. sign in

REVIEW 3 cited by

Risk-Sensitive Soft Actor-Critic for Robust Deep Reinforcement Learning under Distribution Shifts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09992 v1 pith:OFRXNAW3 submitted 2024-02-15 cs.LG cs.SYeess.SY

classification cs.LGcs.SYeess.SY
keywords learningreinforcementdistributionalgorithmdeepshiftsactor-criticcombinatorial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the robustness of deep reinforcement learning algorithms against distribution shifts within contextual multi-stage stochastic combinatorial optimization problems from the operations research domain. In this context, risk-sensitive algorithms promise to learn robust policies. While this field is of general interest to the reinforcement learning community, most studies up-to-date focus on theoretical results rather than real-world performance. With this work, we aim to bridge this gap by formally deriving a novel risk-sensitive deep reinforcement learning algorithm while providing numerical evidence for its efficacy. Specifically, we introduce discrete Soft Actor-Critic for the entropic risk measure by deriving a version of the Bellman equation for the respective Q-values. We establish a corresponding policy improvement result and infer a practical algorithm. We introduce an environment that represents typical contextual multi-stage stochastic combinatorial optimization problems and perform numerical experiments to empirically validate our algorithm's robustness against realistic distribution shifts, without compromising performance on the training distribution. We show that our algorithm is superior to risk-neutral Soft Actor-Critic as well as to two benchmark approaches for robust deep reinforcement learning. Thereby, we provide the first structured analysis on the robustness of reinforcement learning under distribution shifts in the realm of contextual multi-stage stochastic combinatorial optimization problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

    cs.LG 2026-08 accept novelty 7.0 of 10

    The one- and two-timescale algorithms for discounted exponential-utility RL achieve O~(1/sqrt(n)) finite-time rates under Markovian sampling with parameter-free stepsizes.

  2. Diffusion-Modeled Reinforcement Learning for Carbon and Risk-Aware Microgrid Optimization

    cs.LG 2025-07 reject novelty 4.0 of 10

    DiffCarl, a diffusion-actor variant of SAC with carbon pricing and CVaR risk terms, is reported to lower microgrid operating cost by 2.3-30.1% versus baselines, though the paper's own numbers contradict its 28.7% carb...

  3. Risk-Averse Reinforcement Learning with Itakura-Saito Loss

    cs.LG 2025-05 conditional novelty 4.0 of 10

    The Itakura-Saito loss, derived from Bregman divergence, learns risk-averse value functions that match the exponential-utility Bellman equations and trains more stably than exponential MSE in the tested benchmarks.

Pith tools