REVIEW 3 cited by
Stabilizing Transformers for Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown breakthrough success in natural language processing (NLP), achieving state-of-the-art results in domains such as language modeling and machine translation. Harnessing the transformer's ability to process long time horizons of information could provide a similar performance boost in partially observable reinforcement learning (RL) domains, but the large-scale transformers used in NLP have yet to be successfully applied to the RL setting. In this work we demonstrate that the standard transformer architecture is difficult to optimize, which was previously observed in the supervised learning setting but becomes especially pronounced with RL objectives. We propose architectural modifications that substantially improve the stability and learning speed of the original Transformer and XL variant. The proposed architecture, the Gated Transformer-XL (GTrXL), surpasses LSTMs on challenging memory environments and achieves state-of-the-art results on the multi-task DMLab-30 benchmark suite, exceeding the performance of an external memory architecture. We show that the GTrXL, trained using the same losses, has stability and performance that consistently matches or exceeds a competitive LSTM baseline, including on more reactive tasks where memory is less critical. GTrXL offers an easy-to-train, simple-to-implement but substantially more expressive architectural alternative to the standard multi-layer LSTM ubiquitously used for RL agents in partially observable environments.
Forward citations
Cited by 3 Pith papers
-
Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games
A new game model for competitive prize collecting on graphs, with an ordinal-rank conditioning trick that improves multi-agent policy scaling and generalization in road-network simulations.
-
Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes
DBGFQN swaps transformer feed-forward networks for a single BiGRU layer and reports improved average success rate on 23 POMDP environments, but the headline gains and parameter reductions are not reproducible from the...
-
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation
A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.
Discussion (0). Continue with ORCID to comment.