Automatic Jump Start uses Fitted Q Evaluation to adapt the Jump-Start exploration schedule, reducing fine-tuning performance degradation without tuning a tolerance threshold.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Fine-Tuning without Performance Degradation
Automatic Jump Start uses Fitted Q Evaluation to adapt the Jump-Start exploration schedule, reducing fine-tuning performance degradation without tuning a tolerance threshold.