REVIEW 1 cited by
Proximal Policy Optimization and its Dynamic Version for Sequence Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In sequence generation task, many works use policy gradient for model optimization to tackle the intractable backpropagation issue when maximizing the non-differentiable evaluation metrics or fooling the discriminator in adversarial learning. In this paper, we replace policy gradient with proximal policy optimization (PPO), which is a proved more efficient reinforcement learning algorithm, and propose a dynamic approach for PPO (PPO-dynamic). We demonstrate the efficacy of PPO and PPO-dynamic on conditional sequence generation tasks including synthetic experiment and chit-chat chatbot. The results show that PPO and PPO-dynamic can beat policy gradient by stability and performance.
Forward citations
Cited by 1 Pith paper
-
Generative Question Refinement with Deep Reinforcement Learning in Retrieval-based QA System
QREFINE, a BERT- and character-aware Seq2Seq model trained with PPO and answer-aware rewards, generates cleaned questions that improve answer retrieval over previous refinement methods.
Discussion (0). Continue with ORCID to comment.