Critic-Driven Voronoi State Partitioning distills deep RL policies into piecewise-linear models by iteratively adding linear subpolicies in high-value-error regions identified by the critic.
A survey of reinforcement learning algorithms for dynamically varying environments
2 Pith papers cite this work, alongside 214 external citations. Polarity classification is still indexing.
2
Pith papers citing it
214
external citations · OpenAlex
verdicts
UNVERDICTED 2representative citing papers
CL-MARL uses an adaptive curriculum scheduler called FlexDiff and Counterfactual Group Relative Policy Advantage to break static-difficulty training in MARL and achieve higher win rates on hard StarCraft maps.
citing papers explorer
-
Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models
Critic-Driven Voronoi State Partitioning distills deep RL policies into piecewise-linear models by iteratively adding linear subpolicies in high-value-error regions identified by the critic.
-
Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage
CL-MARL uses an adaptive curriculum scheduler called FlexDiff and Counterfactual Group Relative Policy Advantage to break static-difficulty training in MARL and achieve higher win rates on hard StarCraft maps.