← back to paper
arxiv: 2607.28076 · 2 revisions
Group-Reflective Self-Distillation for Agentic Reinforcement Learning