pith. sign in

Optimizing chain-of-thought reasoners via gradient variance minimization in rejection sampling and rl

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.LG 3 cs.AI 1

years

2026 4

representative citing papers

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

cs.LG · 2026-06-03 · conditional · novelty 6.0

Rollout-level advantage-prioritized experience replay for GRPO recycles high-advantage individual rollouts with age eviction and fresh-anchored batches to outperform standard GRPO on math benchmarks, with gains increasing with model size.

citing papers explorer

Showing 4 of 4 citing papers.