REVIEW 2 minor 3 cited by
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation
T0 review · 0 major / 2 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read OrderGrad supplies unbiased gradient estimates for any fixed-sample order-statistic objective by a simple reward transformation before a standard policy-gradient step.
desk verdict OrderGrad turns order-statistic objectives into a fixed reward transform that plugs into standard policy gradients, with unbiasedness holding for fixed N and weights. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The finite-sample L-statistic defined by a fixed rank-weight vector applied to a fixed number of sorted samples; the gradient estimator is obtained by weighting each sample's contribution by its rank-dependent transformed value.
What would settle it
For a simple differentiable policy and a known order-statistic objective, compute the true gradient analytically and compare it to the Monte-Carlo average of OrderGrad estimates over many independent batches of fixed size; any nonzero bias would falsify the claim.
Extended reading notes
Core claim
For any fixed sample size and any fixed rank-weight vector, OrderGrad yields an unbiased gradient estimator of the corresponding finite-sample L-statistic objective; the estimator is realized simply by transforming each reward according to its rank within the batch before applying a standard policy-gradient or reparameterized update.
Load-bearing premise
The number of samples used to form the order statistics must stay fixed and the rank weights must be chosen independently of the realized sample values.
Editorial extensions
If this is right
- Any existing policy-gradient or reparameterized algorithm can optimize VaR, CVaR, trimmed means, or best-of-K criteria after only a reward transformation.
- The same estimator applies unchanged to both on-policy likelihood-ratio and off-policy or reparameterized settings.
- Variance of the estimator can be controlled by choice of rank weights without altering the underlying optimizer.
- Tasks whose deployment objective differs from mean return, such as LLM math post-training, become directly addressable.
Reading between the lines
- If a bias-correction term could be derived, the rank weights might be allowed to adapt to the data without losing unbiasedness.
- The batch-sorting construction may extend to continuous or infinite-horizon settings by replacing exact order statistics with suitable approximations.
- The same reward-transformation idea could be applied outside reinforcement learning to any gradient-based optimizer whose loss is an order statistic.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces OrderGrad, a family of likelihood-ratio and reparameterization gradient estimators for policy optimization of finite-sample order-statistic (L-statistic) objectives. For any fixed sample size N and fixed rank-weight vector w independent of the data, it claims the resulting estimators are unbiased for the corresponding weighted sum of order statistics, recovering objectives such as VaR, CVaR, trimmed means, medians, and top-m criteria via a simple reward transformation. The work includes variance analysis and empirical evaluation on tasks where mean optimization is mismatched to the deployment goal, including LLM math post-training.
Significance. If the unbiasedness result holds under the stated conditions, OrderGrad supplies a unified, plug-and-play route to optimizing non-mean objectives in reinforcement learning and policy gradients. This is significant for risk-averse, robust, and exploratory learning settings. The open-source code link is a positive contribution that supports reproducibility.
minor comments (2)
- The variance analysis mentioned in the abstract would benefit from a dedicated subsection with explicit variance expressions or bounds to make the estimator's behavior easier to compare with standard policy gradients.
- Figure captions and axis labels in the empirical section should explicitly state the sample size N and weight vector w used in each experiment to allow direct verification of the fixed-N, fixed-w condition.
Simulated Author's Rebuttal
We thank the referee for the positive assessment of OrderGrad and the recommendation for minor revision. The provided summary accurately reflects the paper's focus on unbiased likelihood-ratio and reparameterization estimators for finite-sample L-statistic objectives via rank-based reward transformations.
Circularity Check
No significant circularity
full rationale
The paper's core claim is that OrderGrad yields an unbiased gradient estimator for any fixed N and fixed rank-weight vector w by applying the likelihood-ratio identity (or reparameterization) to the L-statistic L = sum w_k R_{(k)}. This is a direct, standard extension of the policy-gradient identity to a well-defined functional of the N i.i.d. samples; the unbiasedness holds by construction of the LR trick once N and w are held constant and independent of the data. No self-citation chain, fitted parameter renamed as prediction, or self-definitional step appears in the derivation. The assumption is stated explicitly in the claim itself, rendering the result self-contained against external benchmarks.
Assumptions & free parameters
assumptions (1)
- domain assumption Standard likelihood-ratio and reparameterization gradient estimators remain valid after the order-statistic reward transformation.
Cite this review
Pith. "Pith review of OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation." pith.science (2026). https://pith.science/paper/DMYZIDMH
@misc{pith2026260606096,
author = {Pith},
title = {Pith review of: OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/DMYZIDMH}},
note = {Machine review of arXiv:2606.06096}
}
read the original abstract
Policy-gradient methods usually optimize expected return, but many real world applications care about distributional properties of returns: tail risk, outlier robustness, or best-of-K discovery. We introduce OrderGrad, a family of likelihood-ratio and reparameterization gradient estimators for order-statistic objectives. OrderGrad optimizes finite-sample L-statistics, i.e., weighted averages of sorted rewards or costs, recovering objectives such as VaR, CVaR, trimmed means, medians, and top-m/best-of-K criteria by changing only the rank weights. For any fixed sample size and rank-weight vector, OrderGrad provides an unbiased gradient estimator for the corresponding order-statistic objective. The method is implemented as a simple reward transformation that can then be used in an otherwise standard policy-gradient or reparameterized update. We study the resulting estimator's variance behavior and evaluate it on tasks where mean optimization is mismatched to the deployment objective, including LLM math post-training and other tasks. OrderGrad provides a unified, plug-and-play route to risk-averse, robust, and exploratory learning. Code: https://github.com/paavo5/ordergrad
Figures
Figures from the paper (17 more)
Forward citations
Cited by 3 Pith papers
-
Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective
Rank-conditioned Horvitz–Thompson reuses all C(n,K) subsets of one Gumbel-Top-n pool for unbiased Plackett–Luce best-of-K value and score-function gradient, with an exact Max-specific DP collapse to a 1-D integral.
-
Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization
A stabilized rank-gap variant of Leader Reward reduced Best-of-8 TSP-100 cost in all three paired seeds (7.7944 vs 7.8136), but the gain is decoder-specific and not statistically confirmatory.
-
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
A per-category best-of-K RL reward, multi-axis max@K, shifts SD3.5-M perceived-appearance distributions toward uniform coverage (Fairness Score +0.23 to +0.36) without quality loss.
Reference graph
Works this paper leans on
-
[1]
Acerbi, C. (2002). Spectral measures of risk: A coherent representation of subjective risk aversion.Journal of Banking & Finance, 26(7):1505–1518
2002
-
[2]
and Tasche, D
Acerbi, C. and Tasche, D. (2002a). Expected shortfall: A natural coherent alternative to value at risk.Economic Notes, 31(2):379–388
-
[3]
and Tasche, D
Acerbi, C. and Tasche, D. (2002b). On the coherence of expected shortfall.Journal of Banking & Finance, 26(7):1487–1503
-
[4]
S., Courville, A
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M. (2021). Deep reinforcement learning at the edge of the statistical precipice.Advances in neural information processing systems, 34:29304–29320
2021
-
[5]
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S. (2024). Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pages 12248–12267
2024
-
[6]
C., Balakrishnan, N., and Nagaraja, H
Arnold, B. C., Balakrishnan, N., and Nagaraja, H. N. (1992).A First Course in Order Statistics. John Wiley & Sons
1992
-
[7]
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002). Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256
2002
- [8]
Show all 118 references
-
[9]
W., Budden, D., Dabney, W., Horgan, D., Dhruva, T
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., Dhruva, T. B., Muldal, A., Heess, N., and Lillicrap, T. P. (2018). Distributed distributional deterministic policy gradients. InInternational Conference on Learning Representations. 10
2018
-
[10]
G., Dabney, W., and Munos, R
Bellemare, M. G., Dabney, W., and Munos, R. (2017). A distributional perspective on rein- forcement learning. InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 449–458. PMLR
2017
-
[11]
G., Dabney, W., and Rowland, M
Bellemare, M. G., Dabney, W., and Rowland, M. (2023).Distributional Reinforcement Learning. The MIT Press, Cambridge, MA
2023
-
[12]
Bickel, P. J. and Lehmann, E. L. (1975). Descriptive statistics for nonparametric models. II. location.The Annals of Statistics, 3(5):1045–1069
1975
-
[13]
Bu, D., Huang, W., Han, A., Nitanda, A., Xue, B., Zhang, Q., Wong, H.-S., and Suzuki, T. (2025). Consistency is not always correct: Towards understanding the role of exploration in post-training reasoning.arXiv preprint arXiv:2511.07368
2025
-
[14]
Burda, Y ., Edwards, H., Storkey, A., and Klimov, O. (2018). Exploration by random network distillation.arXiv preprint arXiv:1810.12894
2018 arXiv
-
[15]
Cai, S., Gao, C., Zhang, Y ., Shi, W., Zhang, J., Bao, K., Wang, Q., and Feng, F. (2025). K-order ranking preference optimization for large language models. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T., editors,Findings of the Association for Computational Linguis...
2025
-
[16]
Cardoso, A. R. and Xu, H. (2019). Risk-averse stochastic convex bandit. InProceedings of the 22nd International Conference on Artificial Intelligence and Statistics, volume 89 ofProceedings of Machine Learning Research, pages 39–47
2019
-
[17]
T., Krishnamurthy, A., and Foster, D
Chen, F., Huang, A., Golowich, N., Malladi, S., Block, A., Ash, J. T., Krishnamurthy, A., and Foster, D. J. (2025a). The coverage principle: How pre-training enables post-training.arXiv preprint arXiv:2510.15020
-
[18]
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y ., Joseph, N., Brockman, G., et al. (2021). Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374
2021 arXiv
-
[19]
X., and Shi, G
Chen, Z., Qin, X., Wu, Y ., Ling, Y ., Ye, Q., Zhao, W. X., and Shi, G. (2025b). Pass@k training for adaptively balancing exploration and exploitation of large reasoning models.arXiv preprint arXiv:2508.10751
-
[20]
X., Zhang, Z., and Wei, F
Cheng, D., Huang, S., Zhu, X., Dai, B., Zhao, W. X., Zhang, Z., and Wei, F. (2025). Reasoning with exploration: An entropy perspective.arXiv preprint arXiv:2506.14758
2025 arXiv
-
[21]
Chow, Y ., Ghavamzadeh, M., Janson, L., and Pavone, M. (2018). Risk-constrained reinforcement learning with percentile risk criteria.Journal of Machine Learning Research, 18(167):1–51
2018
-
[22]
Chow, Y ., Tamar, A., Mannor, S., and Pavone, M. (2015). Risk-sensitive and robust decision- making: A CVaR optimization approach. InAdvances in Neural Information Processing Systems, volume 28, pages 1522–1530
2015
-
[23]
F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. InAdvances in Neural Information Processing Systems, volume 30
2017
-
[24]
Cui, G., Zhang, Y ., Chen, J., Yuan, L., Wang, Z., Zuo, Y ., Li, H., Fan, Y ., Chen, H., Chen, W., et al. (2025). The entropy mechanism of reinforcement learning for reasoning language models. arXiv preprint arXiv:2505.22617
2025 arXiv
-
[25]
Y ., Jegelka, S., and Krause, A
Curi, S., Levy, K. Y ., Jegelka, S., and Krause, A. (2020). Adaptive sampling for stochastic risk-averse learning. InAdvances in Neural Information Processing Systems 33, pages 1036–1047
2020
-
[26]
Dabney, W., Ostrovski, G., Silver, D., and Munos, R. (2018a). Implicit quantile networks for distributional reinforcement learning. InProceedings of the 35th International Conference on Machine Learning, volume 80 ofProceedings of Machine Learning Research, pages 1096–1105. PMLR. 11
-
[27]
G., and Munos, R
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R. (2018b). Distributional reinforce- ment learning with quantile regression. InProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, pages 2892–2901. AAAI Press
-
[28]
Dang, X., Baek, C., Wen, K., Kolter, Z., and Raghunathan, A. (2025). Weight ensembling improves reasoning in language models. InSecond Conference on Language Modeling
2025
-
[29]
Daniell, P. J. (1920). Observations weighted according to order.American Journal of Mathem- atics, 42(4):222–236
1920
-
[30]
Fan, Y ., Lyu, S., Ying, Y ., and Hu, B. (2017). Learning with average top-k loss. InAdvances in Neural Information Processing Systems 30
2017
-
[31]
Gao, J., Pan, L., Wang, Y ., Zhong, R., Lu, C., Cai, Q., Jiang, P., and Zhao, X. (2025). Navigate the unknown: Enhancing LLM reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621
2025
-
[32]
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. (2024). The Llama 3 herd of models.arXiv preprint arXiv:2407.21783
2024 arXiv
-
[33]
Guo, D. et al. (2025a). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning.Nature, 645(8081):633–638
-
[34]
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. (2025b). DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.arXiv preprint arXiv:2501.12948
-
[35]
W., Fried, D., and Welleck, S
He, A. W., Fried, D., and Welleck, S. (2025). Rewarding the unlikely: Lifting GRPO beyond distribution sharpening. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V ., editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processi...
2025
-
[36]
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021). Measuring mathematical problem solving with the MATH dataset.arXiv preprint arXiv:2103.03874
2021 arXiv
-
[37]
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. InProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence...
2018
-
[38]
Holland, M. J. and Haress, E. M. (2021). Learning with risk-averse feedback under potentially heavy tails. InProceedings of the 24th International Conference on Artificial Intelligence and Statistics, volume 130 ofProceedings of Machine Learning Research, pages 892–900
2021
-
[39]
Holland, M. J. and Haress, E. M. (2022). Spectral risk-based learning using unbounded losses. InProceedings of the 25th International Conference on Artificial Intelligence and Statistics, volume 151 ofProceedings of Machine Learning Research, pages 1871–1886
2022
-
[40]
Holland, M. J. and Tanabe, K. (2023). A survey of learning criteria going beyond the usual risk. Journal of Artificial Intelligence Research, 78:781–821
2023
-
[41]
Hu, S., Cai, X., Huang, Y ., Yao, Z., Zhang, L., Zhang, P., Deng, Y ., and Chen, K. (2025). Emergent slow thinking in LLMs as inverse tree freezing.arXiv preprint arXiv:2509.23629
2025 arXiv
-
[42]
Huber, P. J. and Ronchetti, E. M. (2009).Robust Statistics. John Wiley & Sons, 2 edition
2009
-
[43]
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. (2024). OpenAI o1 system card.arXiv preprint arXiv:2412.16720
2024 arXiv
-
[44]
Jiang, Y ., Li, Y ., Chen, G., Liu, D., Cheng, Y ., and Shao, J. (2025). Rethinking entropy regularization in large reasoning models.arXiv preprint arXiv:2509.25133. 12
2025
-
[45]
Khim, J., Leqi, L., Prasad, A., and Ravikumar, P. (2020). Uniform convergence of rank-weighted learning. InInternational conference on machine learning, pages 5254–5263. PMLR
2020
-
[46]
Kingma, D. P. and Welling, M. (2014). Auto-encoding variational Bayes. InInternational Conference on Learning Representations
2014
-
[47]
Koyamada, S., Okano, S., Nishimori, S., Murata, Y ., Habara, K., Kita, H., and Ishii, S. (2023a). pgx: Hardware-accelerated parallel game simulators for reinforcement learning.Advances in Neural Information Processing Systems, 36:45716–45743
-
[48]
Koyamada, S., Parmas, P., Kozuno, T., and Ishii, S. (2023b). Emergence of exploration in policy gradient reinforcement learning via resetting. OpenReview submission to ICLR 2023. https://openreview.net/forum?id=GKsNIC_mQRG
2023
-
[49]
Lambert, N., Morrison, J., Pyatkin, V ., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V ., Liu, A., Dziri, N., Lyu, X., Gu, Y ., Malik, S., Graf, V ., Hwang, J. D., Yang, J., Le Bras, R., Tafjord, O., Wilhelm, C., Soldaini, L., Smith, N. A., Wang, Y ., Dasigi, P., and Ha...
2025
-
[50]
L’Ecuyer, P. (1990). A unified view of the IPA, SF, and LR gradient estimation techniques. Management Science, 36(11):1364–1383
1990
-
[51]
Leqi, L., Huang, A., Lipton, Z., and Azizzadenesheli, K. (2022). Supervised learning with general risk functionals. InInternational Conference on Machine Learning, pages 12570–12592. PMLR
2022
-
[52]
J., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V
Lewkowycz, A., Andreassen, A. J., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V . V ., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y ., Neyshabur, B., Gur-Ari, G., and Misra, V . (2022). Solving quantitative reasoning problems with language models. InAdvances in ...
2022
-
[53]
Li, T., Zhang, Y ., Yu, P., Saha, S., Khashabi, D., Weston, J., Lanchantin, J., and Wang, T. (2025). Jointly reinforcing diversity and quality in language model generations.arXiv preprint arXiv:2509.02534
2025
-
[54]
Liang, Z., Lu, S., Yu, W., Panaganti, K., Zhou, Y ., Mi, H., and Yu, D. (2025). Can LLMs guide their own exploration? gradient-guided reinforcement learning for LLM reasoning.arXiv preprint arXiv:2512.15687
2025
-
[55]
S., and Lin, M
Liu, Z., Chen, C., Li, W., Qi, P., Pang, T., Du, C., Lee, W. S., and Lin, M. (2025). Understanding r1-zero-like training: A critical perspective. InConference on Language Modeling (COLM)
2025
-
[56]
and Mendelson, S
Lugosi, G. and Mendelson, S. (2021). Robust multivariate mean estimation: The optimality of trimmed mean.The Annals of Statistics, 49(1):393–410
2021
-
[57]
G., and Castro, P
Lyle, C., Bellemare, M. G., and Castro, P. S. (2019). A comparative analysis of expected and distributional reinforcement learning. InProceedings of the Thirty-Third AAAI Conference on Artificial Intelligence, volume 33, pages 4504–4511
2019
-
[58]
Matsutani, K., Takashiro, S., Minegishi, G., Kojima, T., Iwasawa, Y ., and Matsuo, Y . (2026). RL squeezes, SFT expands: A comparative study of reasoning LLMs. InThe Fourteenth International Conference on Learning Representations
2026
-
[59]
A., Paudice, A., and Pontil, M
Maurer, A., Parletta, D. A., Paudice, A., and Pontil, M. (2021). Robust unsupervised learning via L-statistic minimization. InInternational Conference on Machine Learning, pages 7524–7533. PMLR
2021
-
[60]
Mavrin, B., Zhang, S., Yao, H., Kong, L., Wu, K., and Yu, Y . (2019). Distributional reinforce- ment learning for efficient exploration. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 4424–443...
2019
-
[61]
and Rezende, D
Mnih, A. and Rezende, D. J. (2016). Variational inference for monte carlo objectives. In Proceedings of the 33rd International Conference on Machine Learning, volume 48 ofProceedings of Machine Learning Research, pages 2188–2196
2016
-
[62]
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A. (2020). Monte carlo gradient estimation in machine learning.Journal of Machine Learning Research, 21(132):1–62
2020
-
[63]
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T. (2010a). Nonpara- metric return distribution approximation for reinforcement learning. InProceedings of the 27th International Conference on Machine Learning, pages 799–806
-
[64]
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T. (2010b). Parametric return density estimation for reinforcement learning. InProceedings of the 26th Conference on Uncertainty in Artificial Intelligence, pages 368–375
-
[65]
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V ., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J. (2021). WebGPT: Browser-assisted question-answering with huma...
2021 arXiv
-
[66]
Nguyen-Tang, T., Gupta, S., and Venkatesh, S. (2021). Distributional reinforcement learning via moment matching. InProceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence, volume 35, pages 9144–9152
2021
-
[67]
Nishimori, S., Parmas, P., Koyamada, S., Kozuno, T., Kitamura, T., Ishii, S., and Matsuo, Y . (2026). Emergence of exploration in policy gradient reinforcement learning via retrying. In Proceedings of the International Conference on Machine Learning
2026
-
[68]
and Tamir, A
Ogryczak, W. and Tamir, A. (2003). Minimizing the sum of the k largest functions in linear time.Information Processing Letters, 85(3):117–122
2003
-
[69]
O’Neill, B. (2025). The distribution of order statistics under sampling without replacement. Journal of Statistical Theory and Applications, 24:663–698
2025
-
[70]
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L. (2019). Solving Rubik’s cube with a robot h...
2019 arXiv
-
[71]
OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Józefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W. (2020). Learning dexterous in-hand manipulation.The Internati...
2020
-
[72]
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
2022
-
[73]
E., Peters, J., and Doya, K
Parmas, P., Rasmussen, C. E., Peters, J., and Doya, K. (2018). PIPPS: Flexible model-based policy search robust to the curse of chaos. InProceedings of the 35th International Conference on Machine Learning, volume 80 ofProceedings of Machine Learning Research, pages 4065–4074
2018
-
[74]
and Seno, T
Parmas, P. and Seno, T. (2022). Proppo: A message passing framework for customizable and composable learning algorithms.Advances in Neural Information Processing Systems, 35:29152– 29165
2022
-
[75]
and Sugiyama, M
Parmas, P. and Sugiyama, M. (2021). A unified view of likelihood ratio and reparameterization gradients. InProceedings of the 24th International Conference on Artificial Intelligence and Statistics, volume 130 ofProceedings of Machine Learning Research, pages 4078–4086
2021
-
[76]
and Schaal, S
Peters, J. and Schaal, S. (2006). Policy gradient methods for robotics. InProceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. 14
2006
-
[77]
and Schaal, S
Peters, J. and Schaal, S. (2008). Reinforcement learning of motor skills with policy gradients. Neural Networks, 21(4):682–697
2008
-
[78]
J., Mohamed, S., and Wierstra, D
Rezende, D. J., Mohamed, S., and Wierstra, D. (2014). Stochastic backpropagation and approximate inference in deep generative models. InProceedings of the 31st International Conference on Machine Learning, volume 32 ofProceedings of Machine Learning Research, pages 1278–1286
2014
-
[79]
Rockafellar, R. T. and Uryasev, S. (2000). Optimization of conditional value-at-risk.Journal of Risk, 2:21–42
2000
-
[80]
Rockafellar, R. T. and Uryasev, S. (2002). Conditional value-at-risk for general loss distributions. Journal of Banking & Finance, 26(7):1443–1471
2002
-
[81]
G., Dabney, W., Munos, R., and Teh, Y
Rowland, M., Bellemare, M. G., Dabney, W., Munos, R., and Teh, Y . W. (2018). An analysis of categorical distributional reinforcement learning. InProceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics, volume 84 ofProceedings of Mach...
2018
-
[82]
G., and Dabney, W
Rowland, M., Dadashi, R., Kumar, S., Munos, R., Bellemare, M. G., and Dabney, W. (2019). Statistics and samples in distributional reinforcement learning. InProceedings of the 36th In- ternational Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Resea...
2019
-
[83]
Y ., Snell, C., Greer, J., Wu, I., Smith, V ., Simchowitz, M., and Kumar, A
Setlur, A., Yang, M. Y ., Snell, C., Greer, J., Wu, I., Smith, V ., Simchowitz, M., and Kumar, A. (2025). e3: Learning to explore enables extrapolation of test-time compute for llms.arXiv preprint arXiv:2506.09026
2025
-
[84]
J., Rushton, P., Singla, S., Parmar, M., Smith, K., Vanjani, Y ., Vaswani, A., Chaluvaraju, A., Hojel, A., Ma, A., Thomas, A., Polloreno, A., Tanwer, A., Sibai, B
Shah, D. J., Rushton, P., Singla, S., Parmar, M., Smith, K., Vanjani, Y ., Vaswani, A., Chaluvaraju, A., Hojel, A., Ma, A., Thomas, A., Polloreno, A., Tanwer, A., Sibai, B. D., Mansingka, D. S., Shivaprasad, D., Shah, I., Stratos, K., Nguyen, K., Callahan, M., Pust, M., Iyer, ...
2025
-
[85]
and Wexler, Y
Shalev-Shwartz, S. and Wexler, Y . (2016). Minimizing the maximal loss: How and why. In Proceedings of the 33rd International Conference on Machine Learning, pages 793–801
2016
-
[86]
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y ., Wu, Y ., et al. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300
2024 arXiv
-
[87]
Shapiro, A. (2013). On Kusuoka representation of law invariant risk measures.Mathematics of Operations Research, 38(1):142–152
2013
-
[88]
Shen, H. (2026). On entropy control in LLM-RL algorithms. InThe Fourteenth International Conference on Learning Representations
2026
-
[89]
and Sanghavi, S
Shen, Y . and Sanghavi, S. (2019). Learning with bad training data via iterative trimmed loss minimization. In Chaudhuri, K. and Salakhutdinov, R., editors,Proceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Researc...
2019
-
[90]
Song, Y ., Kempe, J., and Munos, R. (2025a). Outcome-based exploration for LLM reasoning. arXiv preprint arXiv:2509.06941
-
[91]
M., Foster, D., and Ghai, U
Song, Y ., Zhang, H., Eisenach, C., Kakade, S. M., Foster, D., and Ghai, U. (2025b). Mind the gap: Examining the self-improvement capabilities of large language models. InThe Thirteenth International Conference on Learning Representations
-
[92]
M., Lowe, R., V oss, C., Radford, A., Amodei, D., and Christiano, P
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., V oss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020). Learning to summarize with human feedback. InAdvances in Neural Information Processing Systems, volume 33, pages 3008–3021. Curran Associates, Inc. 15
2020
-
[93]
Sugiyama, M., Hachiya, H., Kashima, H., and Morimura, T. (2010). Least absolute policy iteration—a robust approach to value function approximation.IEICE Transactions on Information and Systems, E93-D(9):2555–2565
2010
-
[94]
Sui, Y ., Chuang, Y .-N., Wang, G., Zhang, J., Zhang, T., Yuan, J., Liu, H., Wen, A., Zhong, S., Zou, N., Chen, H., and Hu, X. (2025). Stop Overthinking: A survey on efficient reasoning for large language models.Transactions on Machine Learning Research. https://openreview. ne...
2025
-
[95]
Tang, Y ., Zheng, K., Synnaeve, G., and Munos, R. (2025). Optimizing language models for inference time objectives using reinforcement learning. InInternational Conference on Machine Learning, pages 59066–59085. PMLR
2025
-
[96]
J., Krishnamurthy, A., and Ash, J
Tuyls, J., Foster, D. J., Krishnamurthy, A., and Ash, J. T. (2025). Representation-based explora- tion for language models: From test-time to post-training.arXiv preprint arXiv:2510.11686
2025
-
[97]
and Karkhanis, D
Walder, C. and Karkhanis, D. T. (2025). Pass@k policy optimization: Solving harder reinforce- ment learning problems.Advances in Neural Information Processing Systems, 38:152416–152445
2025
-
[98]
A., Le, Q., Diao, E., Zhou, J., Ding, J., and Anwar, A
Wang, X., Du, J., Khan, A. A., Le, Q., Diao, E., Zhou, J., Ding, J., and Anwar, A. (2025a). Beyond expectations: Quantile-guided alignment for risk-calibrated language models. InAdvances in Neural Information Processing Systems
-
[99]
Wang, Z., Zhou, F., Li, X., and Liu, P. (2025b). OctoThinker: Mid-training incentivizes reinforcement learning scaling.arXiv preprint arXiv:2506.20512
-
[100]
Wen, X., Liu, Z., Zheng, S., Ye, S., Wu, Z., Wang, Y ., Xu, Z., Liang, X., Li, J., Miao, Z., Bian, J., and Yang, M. (2026). Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs. InThe Fourteenth International Conference on Learn...
2026
-
[101]
Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine Learning, 8:229–256
1992
-
[102]
Wu, F., Xuan, W., Lu, X., Liu, M., Dong, Y ., Harchaoui, Z., and Choi, Y . (2025). The invisible leash: Why RLVR may or may not escape its origin.arXiv preprint arXiv:2507.14843
2025
-
[103]
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. (2025a). Qwen3 technical report.arXiv preprint arXiv:2505.09388
-
[104]
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, ...
-
[105]
Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., Liu, D., Tu, J., Zhou, J., Lin, J., Lu, K., Xue, M., Lin, R., Liu, T., Ren, X., and Zhang, Z. (2024). Qwen2.5-Math technical report: Toward mathematical expert model via self-improvement.arXiv preprint arXiv:2409.12122
2024 arXiv
-
[106]
Yang, D., Zhao, L., Lin, Z., Qin, T., Bian, J., and Liu, T.-Y . (2019). Fully parameterized quantile function for distributional reinforcement learning. InAdvances in Neural Information Processing Systems 32
2019
-
[107]
and Tian, T
Young, K. and Tian, T. (2019). MinAtar: An atari-inspired testbed for thorough and reprodu- cible reinforcement learning experiments.arXiv preprint arXiv:1903.03176
2019 arXiv
-
[108]
Yu, Q., Zhang, Z., Zhu, R., Yuan, Y ., Zuo, X., Yue, Y ., Dai, W., Fan, T., Liu, G., Liu, L., et al. (2025). DAPO: An open-source LLM reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244
2025
-
[109]
E., Yu, P., and Xu, J
Yu, Z., Su, Z., Tao, L., Wang, H., Singh, A., Yu, H., Wang, J., Gao, H., Yuan, W., Weston, J. E., Yu, P., and Xu, J. (2026). RESTRAIN: From spurious votes to signals — self-training RL with self-penalization. InThe Fourteenth International Conference on Learning Representations. 16
2026
-
[110]
Yue, Y ., Chen, Z., Lu, R., Zhao, A., Wang, Z., Yue, Y ., Song, S., and Huang, G. (2025). Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? In The Thirty-ninth Annual Conference on Neural Information Processing Systems
2025
-
[111]
Zhang, C., Neubig, G., and Yue, X. (2025a). On the interplay of pre-training, mid-training, and rl on reasoning language models.arXiv preprint arXiv:2512.07783
-
[112]
Zhang, S., Yu, D., Feng, Y ., Jin, B., Wang, Z., Peebles, J., and Wang, Z. (2025b). Learning to reason as action abstractions with scalable mid-training rl.arXiv preprint arXiv:2509.25810
-
[113]
Zhao, R., Meterez, A., Kakade, S., Pehlevan, C., Jelassi, S., and Malach, E. (2025). Echo chamber: Rl post-training amplifies behaviors learned in pretraining. InSecond Conference on Language Modeling
2025
-
[114]
Zheng, T., Xing, T., Gu, Q., Liang, T., Qu, X., Zhou, X., Li, Y ., Wen, Z., Lin, C., Huang, W., et al. (2025). First return, entropy-eliciting explore.arXiv preprint arXiv:2507.07017
2025
-
[115]
Zhou, Y ., Liang, Z., Liu, H., Yu, W., Panaganti, K., Song, L., Yu, D., Zhang, X., Mi, H., and Yu, D. (2025). Evolving language models without labels: Majority drives selection, novelty promotes variation.arXiv preprint arXiv:2509.15194
2025
-
[116]
M., Stiennon, N., Wu, J., Brown, T
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019). Fine-tuning language models from human preferences.arXiv preprint arXiv:1909.08593. 17 Appendix A Extended Related Work 19 B Notation 20 C Exact known-( ¯R, p)fo...
2019 arXiv
-
[117]
[96] computed a representation-based novelty score from hidden states to boost exploration
employed a semantic diversity score with an external semantic comparator, and Tuyls et al. [96] computed a representation-based novelty score from hidden states to boost exploration. Liang et al
-
[118]
Setlur et al
leveraged reward model gradients to improve temperature sampling. Setlur et al. [83] promoted in-context exploration via skill asymmetries and negative gradients. B Notation This appendix gives detailed derivations behind the main text. We follow the organization: first an exa...
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.