Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

· 2026 · cs.LG · arXiv 2605.06387

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy distillation and standard reinforcement learning. However, we find that its standard advantage weighted policy gradient suffers from three structural weaknesses, including high variance updates, vanishing gradients in zero-advantage regions, and exploration bottlenecks when corrective signals are insufficient. We therefore propose Asymmetric On-Policy Distillation (AOPD), which replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning. Experiments on mathematical reasoning benchmarks show that AOPD consistently outperforms standard OPD, with average gains of 4.09 / 8.34 under strong / weak initialization, respectively. AOPD also maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.

representative citing papers

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling

cs.CL · 2026-06-08 · unverdicted · novelty 4.0

SG-OPD adds sign-consistency gating and phased teacher sampling to on-policy distillation, reporting average gains of 1.98 per sample and 7.50 per question over standard OPD on math benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling cs.CL · 2026-06-08 · unverdicted · none · ref 3 · internal anchor
SG-OPD adds sign-consistency gating and phased teacher sampling to on-policy distillation, reporting average gains of 1.98 per sample and 7.50 per question over standard OPD on math benchmarks.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

fields

years

verdicts

representative citing papers

citing papers explorer