REVIEW 13 cited by
Soft Actor-Critic for Discrete Action Settings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an alternative version of the Soft Actor-Critic algorithm that is applicable to discrete action settings. We then show that, even without any hyperparameter tuning, it is competitive with the tuned model-free state-of-the-art on a selection of games from the Atari suite.
Forward citations
Cited by 13 Pith papers
-
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
ACPO decomposes the joint policy gradient into per-agent terms allowing independent actor training that collectively forms a joint gradient step in CTDE-based MARL.
-
Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance
A per-task guide policy selects other tasks' control policies to generate training trajectories, boosting multi-task RL performance across five baselines.
-
Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening
New game-theoretic and optimization models for honeypot placement, temporal decoy placement, and human-in-the-loop edge removal on Active Directory attack graphs, with hardness proofs and scalable heuristics.
-
Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding
Hybrid RL-MPC trained on the full hybrid action space parametrizes continuous MPC via discrete rollouts and a critic terminal cost, yielding near-MINLP F1 strategies with recursive feasibility under a structural assumption.
-
Generalizable Pareto-Optimal Offloading with Reinforcement Learning in Mobile Edge Computing
A single discrete soft actor-critic policy conditioned on preference and system context produces near-Pareto-optimal offloading decisions that generalize to unseen server counts and CPU frequencies.
-
Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing
A framework that uses LLM-annotated knowledge concepts and a knowledge tracing simulator to train RL policies for exercise recommendation, with a model-based value estimator that improves simulated knowledge gains.
-
Learning To Communicate Over An Unknown Shared Network
A DRL-based querying policy trained only on a single-parameter queue simulation transfers zero-shot to real WiFi (5-50 agents) and cellular networks and adapts its query rate to congestion.
-
Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits
A multi-agent bandit RL formulation for binary photonic topology optimization produces designs that outperform gradient-based optimization on eight simulated 2D and 3D tasks.
-
Hierarchical Learning-Enhanced MPC for Safe Crowd Navigation with Heterogeneous Constraints
A hierarchical planner using a GNN-based local-goal recommender, spatio-temporal search, and MPC achieves high success rates in simulated and real crowd navigation, at the cost of slower navigation.
-
SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training
SLAC learns a latent action space in a low-fidelity simulator and uses it for real-world reinforcement learning, solving whole-body mobile manipulation tasks in under an hour without demonstrations.
-
DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning
A PPO-trained agent outputs a degeneracy factor that blends the laser-based particle cloud toward the motion model, reducing drift and mapping errors in feature-less corridors.
-
Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections
A hierarchical RL agent with a goal-conditioned collision prediction module achieves 94.7% success and 3.3% collisions in SMARTS intersection tasks, outperforming flat RL baselines.
-
Deep Reinforcement Learning: From First Principles to Reasoning Models
A textbook survey of deep reinforcement learning, from Bellman foundations to DQN, PPO, MuZero, offline RL, and reasoning models, with UAV/SD-WAN examples throughout.
Discussion (0). Sign in to comment.