REVIEW 6 cited by
A Survey of Deep Reinforcement Learning in Video Games
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep reinforcement learning (DRL) has made great achievements since proposed. Generally, DRL agents receive high-dimensional inputs at each step, and make actions according to deep-neural-network-based policies. This learning mechanism updates the policy to maximize the return with an end-to-end method. In this paper, we survey the progress of DRL methods, including value-based, policy gradient, and model-based algorithms, and compare their main techniques and properties. Besides, DRL plays an important role in game artificial intelligence (AI). We also take a review of the achievements of DRL in various video games, including classical Arcade games, first-person perspective games and multi-agent real-time strategy games, from 2D to 3D, and from single-agent to multi-agent. A large number of video game AIs with DRL have achieved super-human performance, while there are still some challenges in this domain. Therefore, we also discuss some key points when applying DRL methods to this field, including exploration-exploitation, sample efficiency, generalization and transfer, multi-agent learning, imperfect information, and delayed spare rewards, as well as some research directions.
Forward citations
Cited by 6 Pith papers
-
Counterfactual Shapley Credit Assignment
Counterfactual Shapley values, computed by simulated 'what-if' action replacements, redistribute RL rewards without changing the optimal policy and improve credit assignment in stochastic, sparse, delayed-reward tasks.
-
Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
PEARL trains a policy by differentiating a short horizon of the simulator and using a learned adjoint network to account for long-term return gradients, beating PPO, TD3, BPTT, and SHAC on two double-gyre navigation tasks.
-
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
GraphAllocBench provides a graph-based resource-allocation benchmark and two supplemental metrics that expose preference-consistency failures in multi-objective RL policies.
-
Enhancing Reinforcement Learning in 3D Environments through Semantic Segmentation: A Case Study in ViZDoom
Semantic-segmentation masks can replace RGB input to ViZDoom RL agents with comparable performance and far lower buffer memory, and adding them as an extra channel improves frag scores.
-
Causal-aware Large Language Models: Enhancing Decision-Making Through Learning, Adapting and Acting
A learning-adapting-acting framework that uses an LLM to build and update a causal graph of the game world, then uses that graph to generate goals and shape rewards for an RL agent in Crafter.
-
A Comprehensive Review of Multi-Agent Reinforcement Learning in Video Games
A survey of multi-agent reinforcement learning in video games, plus a proposed five-dimension, MDP-based classification for comparing game complexity.
Discussion (0). Continue with ORCID to comment.