Pith. sign in

REVIEW 13 cited by

Soft Actor-Critic for Discrete Action Settings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.07207 v2 pith:Q6O7JZUZ submitted 2019-10-16 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords settingsactiondiscreteactor-criticsoftalgorithmapplicablestate-of-the-art
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Soft Actor-Critic is a state-of-the-art reinforcement learning algorithm for continuous action settings that is not applicable to discrete action settings. Many important settings involve discrete actions, however, and so here we derive an alternative version of the Soft Actor-Critic algorithm that is applicable to discrete action settings. We then show that, even without any hyperparameter tuning, it is competitive with the tuned model-free state-of-the-art on a selection of games from the Atari suite.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    ACPO decomposes the joint policy gradient into per-agent terms allowing independent actor training that collectively forms a joint gradient step in CTDE-based MARL.

  2. Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A per-task guide policy selects other tasks' control policies to generate training trajectories, boosting multi-task RL performance across five baselines.

  3. Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening

    cs.CR 2026-07 conditional novelty 6.0 of 10

    New game-theoretic and optimization models for honeypot placement, temporal decoy placement, and human-in-the-loop edge removal on Active Directory attack graphs, with hardness proofs and scalable heuristics.

  4. Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Hybrid RL-MPC trained on the full hybrid action space parametrizes continuous MPC via discrete rollouts and a critic terminal cost, yielding near-MINLP F1 strategies with recursive feasibility under a structural assumption.

  5. Generalizable Pareto-Optimal Offloading with Reinforcement Learning in Mobile Edge Computing

    eess.SY 2025-08 conditional novelty 6.0 of 10

    A single discrete soft actor-critic policy conditioned on preference and system context produces near-Pareto-optimal offloading decisions that generalize to unseen server counts and CPU frequencies.

  6. Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing

    cs.AI 2025-07 reject novelty 6.0 of 10

    A framework that uses LLM-annotated knowledge concepts and a knowledge tracing simulator to train RL policies for exercise recommendation, with a model-based value estimator that improves simulated knowledge gains.

  7. Learning To Communicate Over An Unknown Shared Network

    cs.MA 2025-07 conditional novelty 6.0 of 10

    A DRL-based querying policy trained only on a single-parameter queue simulation transfers zero-shot to real WiFi (5-50 agents) and cellular networks and adapts its query rate to congestion.

  8. Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A multi-agent bandit RL formulation for binary photonic topology optimization produces designs that outperform gradient-based optimization on eight simulated 2D and 3D tasks.

  9. Hierarchical Learning-Enhanced MPC for Safe Crowd Navigation with Heterogeneous Constraints

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A hierarchical planner using a GNN-based local-goal recommender, spatio-temporal search, and MPC achieves high success rates in simulated and real crowd navigation, at the cost of slower navigation.

  10. SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training

    cs.RO 2025-06 conditional novelty 5.0 of 10

    SLAC learns a latent action space in a low-fidelity simulator and uses it for real-world reinforcement learning, solving whole-body mobile manipulation tasks in under an hour without demonstrations.

  11. DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A PPO-trained agent outputs a degeneracy factor that blends the laser-based particle cloud toward the motion model, reducing drift and mapping errors in feature-less corridors.

  12. Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A hierarchical RL agent with a goal-conditioned collision prediction module achieves 94.7% success and 3.3% collisions in SMARTS intersection tasks, outperforming flat RL baselines.

  13. Deep Reinforcement Learning: From First Principles to Reasoning Models

    eess.SY 2026-07 unverdicted novelty 1.0 of 10

    A textbook survey of deep reinforcement learning, from Bellman foundations to DQN, PPO, MuZero, offline RL, and reasoning models, with UAV/SD-WAN examples throughout.

Pith tools