The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL

Baoxiang Wang; Ge Zhang; Jiacai Liu; Jiawei Xu; Longtao Zheng; Qian Liu; Tianle Cai; Wei Liu; Yaxiang Zhang; Yingru Li

arxiv: 2602.07078 · v2 · pith:4WDAUHA4new · submitted 2026-02-06 · 💻 cs.LG · cs.AI

The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL

Yingru Li , Jiawei Xu , Ziniu Li , Jiacai Liu , Wei Liu , Yuxuan Tong , Longtao Zheng , Zhenghai Xue

show 5 more authors

Yaxiang Zhang Tianle Cai Ge Zhang Qian Liu Baoxiang Wang

This is my paper

classification 💻 cs.LG cs.AI

keywords baselinegradienttokenoptimalvariancecomputationheterogeneitylarge

0 comments

read the original abstract

Reinforcement Learning (RL) for Large Language Models (LLMs) often suffers from training collapse in long-horizon tasks due to exploding gradient variance. To mitigate this, a baseline is commonly introduced for advantage computation; however, traditional value models remain difficult to optimize, and standard group-based baselines overlook sequence heterogeneity. Although classic optimal baseline theory can achieve global variance reduction, it neglects token heterogeneity and requires prohibitive gradient-based computation. In this work, we derive the Optimal Token Baseline (OTB) from first principles, proving that gradient updates should be weighted inversely to their cumulative gradient norm. To ensure efficiency, we propose the Logit-Gradient Proxy that approximates the gradient norm using only forward-pass probabilities. Our method achieves training stability and matches the performance of large group sizes ($N=32$) with only $N=4$, reducing token consumption by over 65\% across single-turn and tool-integrated reasoning tasks.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate
cs.LG 2026-05 unverdicted novelty 7.0

ARCA assigns token credit in LoRA-based LLM RL from the norm of adapter-induced hidden state changes, yielding non-degenerate distributions and competitive performance on MATH tasks with Qwen3-1.7B under GRPO.
Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
cs.LG 2026-05 unverdicted novelty 6.0

Optimal Coefficient Calibration (OCC) enables joint MTP-RL training to match or exceed the detach baseline on six math reasoning benchmarks by tracking the optimal coefficient online.
Policy Improvement Reinforcement Learning
cs.LG 2026-04 unverdicted novelty 6.0

PIRL maximizes cumulative policy improvement across iterations instead of surrogate rewards and is proven aligned with final performance; PIPO implements it via retrospective verification for stable closed-loop optimization.