REVIEW 1 cited by
Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We explore the capability of evolution strategies to train an agent with a policy based on a transformer architecture in a reinforcement learning setting. We performed experiments using OpenAI's highly parallelizable evolution strategy to train Decision Transformer in the MuJoCo Humanoid locomotion environment and in the environment of Atari games, testing the ability of this black-box optimization technique to train even such relatively large and complicated models (compared to those previously tested in the literature). The examined evolution strategy proved to be, in general, capable of achieving strong results and managed to produce high-performing agents, showcasing evolution's ability to tackle the training of even such complex models.
Forward citations
Cited by 1 Pith paper
-
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Evolution strategies can full-parameter fine-tune billion-parameter LLMs, outperforming PPO and GRPO on the Countdown task and reward robustness in a conciseness task.
Discussion (0). Continue with ORCID to comment.