REVIEW 2 major objections 4 minor 80 references
This paper claims that RLVR fine-tuning of billion-parameter language models admits non-vacuous PAC-Bayes generalization bounds when the weight update is compressed via TinyLoRA and on-policy distillation, with reported bounds exceeding bas
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 01:53 UTC pith:NBQNU4BA
load-bearing objection First non-vacuous RLVR bounds at scale, but the reported tightness relies on a subsample used for model selection; a small union-bound fix likely preserves the result. the 2 major comments →
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the generalization gap of RLVR can be bounded by a compression-based PAC-Bayes inequality. After reparameterizing autoregressive token sampling with Gumbel noise, the population reward takes the form R(h) = E_x E_ξ r(x, g(x, ξ; h)), so the gap splits into an extrinsic term over prompts and an intrinsic term over generation noise. The extrinsic term is penalized by the description length C(h) of the weight update under a universal prior; the intrinsic term shrinks as the number of generations per prompt grows. The paper argues that TinyLoRA plus quantization makes C(h) small enough (roughly 3,400–8,000 bits) that the resulting lower bound is non-vacuous, and
What carries the argument
Gumbel-max reparameterization: each sampled token is written as the argmax of (logits/τ) plus i.i.d. Gumbel noise, making a stochastic rollout a deterministic function g(x, ξ; h) of the prompt, fixed noise, and hypothesis. This converts RLVR's two sources of randomness — the prompt distribution and the model's own generation — into two explicit, hypothesis-independent distributions, which is what allows PAC-Bayes compression arguments to apply. The compression side is carried by TinyLoRA, a low-rank adapter built from the frozen weights' truncated SVD whose trainable part is a small vector (3,984 components here), quantized to a few bits; the paper's bound charges only the compressed bit-len
Load-bearing premise
The load-bearing premise is that the hypothesis h is independent of the m0-prompt subsample used to estimate the empirical reward (Corollary 1, Section 3.3); in Section 5.1/Appendix B.3 the 4,096 reward-estimation prompts come from the same training pool that produced h, so h is not independent, the in-sample reward estimate is biased high, and the reported bounds do not follow from the stated theorem.
What would settle it
Recompute the empirical reward on a genuinely held-out set of 4,096 prompts that were never used in teacher RLVR, distillation, or quantization selection, then apply Corollary 1 with the appropriate sample-size penalty; if the resulting lower bound falls below the base model's accuracy or becomes vacuous, the reported 9-51% margin is an artifact of in-sample reward estimation.
If this is right
- A domain-specific RLVR model can be certified before deployment: with high probability, its expected accuracy on unseen prompts from the training distribution is at least the computed lower bound.
- The intrinsic generation-noise penalty becomes negligible once n is a few dozen samples per prompt, so the main obstacle to tight bounds is the number of training prompts and the adapter's compressed size, not the stochasticity of decoding.
- Compression via distillation is an enabling step for generalization guarantees, not just a cost-saving step: direct RLVR with TinyLoRA and standard LoRA both give vacuous bounds in the paper's ablations.
- For out-of-distribution evaluation, if the hypothesis is chosen independently of a labeled target corpus, the compression term drops away entirely and only a sample-size Hoeffding term remains, giving non-vacuous cross-domain bounds in the paper's twelve ordered pairs.
Where Pith is reading between the lines
- The reported 6-11% gaps are computed with reward-estimation prompts drawn from the same pool used to train the student; a truly held-out reward-estimation set would test whether the bounds survive the independence condition in Corollary 1, and would likely loosen them.
- The paper's cross-domain asymmetry — math transfers to every target while SQL transfers negatively — suggests a possible ordering of reasoning skills by transferability, but that ordering is an interpretation, not a claim the paper proves.
- At the 4B scale the compression budget is met by TinyLoRA; extending the same argument to models tens or hundreds of times larger would require adapters with even fewer effective bits, and the paper itself leaves this as an open question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a PAC-Bayes compression framework for deriving generalization bounds for reinforcement learning with verifiable rewards (RLVR) in large language models. The theoretical core is Theorem 1, which reparameterizes stochastic token generation via Gumbel-max noise, decomposes the generalization gap into an extrinsic component over the prompt distribution and an intrinsic component over generation noise, and obtains a bound involving the compression size C(h) of the hypothesis. Corollary 1 extends this to subsampled reward estimation. To make the bounds non-vacuous, the authors propose Progressive RLVR, a pipeline that first trains an RLVR teacher, then distills on-policy into a TinyLoRA student, then quantizes the adapter. They report non-vacuous bounds on four domains (math, code, general knowledge, Text-to-SQL) and include detailed appendices describing the encoding scheme, hyperparameters, and OOD extensions. The central derivation is plausible, but the reported empirical bounds depend on a data-dependent model-selection procedure over quantization levels that violates the independence condition in Corollary 1 as stated.
Significance. If the reported bounds survive correction, this would be a notable contribution: it would give the first non-vacuous generalization bounds for billion-parameter RLVR fine-tuning, with a clean separation of prompt-level and generation-level stochasticity. The compression accounting in Appendix C is unusually explicit, and the inclusion of code/data links is good practice. The Gumbel-max reparameterization is a simple but effective adaptation of prior PAC-Bayes compression work. However, the validity of the headline numbers rests on Corollary 1's independence condition, and the paper's own Appendix C documents a grid search over quantization levels that violates that condition. The flaw is localized and likely fixable with a modest union bound, but the reported numbers are not rigorously justified as written.
major comments (2)
- [Appendix C, Eq. (1); Corollary 1] The reported bounds are computed after searching k in {2,3,4,5,6,7,8,16,32,64} and selecting the tightest bound on the same 4,096-point reward-estimation subsample used in the bound. Corollary 1 explicitly requires h to be chosen independently of that subsample, and its proof avoids a union bound over H precisely because h is fixed before the subsample is drawn. The 4-bit k-selection cost added to C(h) in Eq. (1) inflates only the full-sample term (the one with 1/sqrt(m)), not the subsampling term (the one with 1/sqrt(m0)). Thus the empirical reward for the selected k is a selection-biased estimate of the full-data reward, and the reported bounds do not follow from Corollary 1 as stated. A union bound over the 10 k values would increase the subsampling penalty by about a factor sqrt(log(30/delta)/log(3/delta)) ≈ 1.25 at delta=0.05, so the non-vacuous conclusion may survive, but the numbe
- [Appendix B.3 / Section 5.1] The evaluation protocol says the 4,096 problems are 'randomly drawn from the corresponding eval-task dataset,' while Corollary 1 requires the m0 points to be a subsample of the m training prompts. If the 4,096 points are actually a held-out set rather than a subsample of the training pool, the correct statement is Theorem 2 (OOD bound), not Corollary 1, and the compression term C(h) should be absent. The reported decomposition in Figure 4 with a 7-8% extrinsic penalty suggests the full training m is being used, but the text should state explicitly that the reward-estimation subsample is a subset of the training prompts. This ambiguity is load-bearing for interpreting what quantity is bounded.
minor comments (4)
- [Appendix A.2] In the first displayed equation of the proof, the intrinsic-error term writes r(x, g(x, ξ; h)) where it should be r(x_i, g(x_i, ξ; h)); the subscript i is missing. Also, ξ_j should be ξ_i,j in the full-sample estimator notation.
- [Section 1] Typo: 'we developProgressive RLVR' should read 'we develop Progressive RLVR'.
- [Appendix C] The paper should clarify whether the codebook entries are fit using the full training set or the 4,096-point reward-estimation subsample. If the latter, the same independence issue as the k-selection arises for the codebook itself.
- [Corollary 1 proof] The proof describes the subsampled indices as i.i.d. draws with replacement, while the statement says 'subsample drawn uniformly at random,' which often means without replacement. This is not a correctness problem because the with-replacement bound is conservative, but the mismatch should be noted.
Circularity Check
Data-dependent choice of quantization level k violates Corollary 1's independence condition, so the reported tight bounds are not valid as stated.
specific steps
-
fitted input called prediction
[Appendix C.2 (Eq. 1, k-selection); Corollary 1 (Section 3.3); Section 5.1 reward-estimation setup]
"We searched k∈{2,3,4,5,6,7,8,16,32,64} (a grid of 10 values) and report the tightest bound across this grid. ... C(h_k)=L_KT(h_k)+16k+4. ... For reward estimation (Corollary 1), we sample 4,096 data points and 64 noise vectors per data point. ... [Corollary 1:] h∈H be any hypothesis chosen independently of the subsample."
Corollary 1 applies only to a hypothesis chosen independently of the reward-estimation subsample. The experiments choose the quantization level k by evaluating the bound for each of 10 values on the same 4,096-prompt subsample used to compute the empirical reward, so the reported hypothesis (including k) is not fixed before the subsample is drawn. The added 4-bit k-selection cost only inflates the full-sample compression term C(h); the subsampling term Δ√(log(3/δ)/2m0) has no union bound over the 10 choices. The reported bound is the maximum of 10 data-dependent quantities, so its tightness is partly a selection artifact rather than a consequence of Corollary 1. The missing union bound is modest (~0.5 percentage points), so the non-vacuous claim may survive correction, but the stated deriv
full rationale
The core PAC-Bayes/Gumbel-max derivation (Theorem 1 and Corollary 1) is mathematically self-contained: it is a standard compression bound applied to a reparameterized stochastic reward, with the Solomonoff prior and Hoeffding arguments stated explicitly. The paper's self-citations (e.g., [28]) are motivational, not load-bearing. However, the experimental application of Corollary 1 is invalid as stated because the quantization level k is selected on the same 4,096-prompt subsample that supplies the empirical reward. Corollary 1 explicitly requires h to be chosen independently of the subsample, and the 4-bit k-selection cost added to C(h) only inflates the full-sample complexity term; the subsampling term has no union bound over the 10 grid values. The reported 'tightest bound across this grid' is therefore a maximum over ten data-dependent quantities, not a high-probability lower bound guaranteed by the stated theorem. This is a fitted-selection step presented as a valid prediction; the missing union bound is modest, so the central non-vacuous claim may survive repair, but the reported bounds are not justified by the paper's own derivation as written. Score 6 reflects partial circularity in the empirical bounds, not in the underlying theory.
Axiom & Free-Parameter Ledger
free parameters (2)
- Quantization level k =
5 (for all domains from Table 3)
- Codebook entries {c_1,...,c_k} =
Domain-specific fp16 values (80 bits total at k=5)
axioms (5)
- domain assumption The subsample of m0 prompts used for reward estimation is drawn independently of the hypothesis h (Corollary 1).
- domain assumption The reward is deterministic and bounded in [a, a+Δ] for all prompts and generations.
- domain assumption Prompts are i.i.d. from a fixed population P0, and deployment distribution matches training.
- standard math Token generation can be reparameterized via Gumbel-max with fixed, hypothesis-independent noise.
- standard math Kolmogorov complexity is upper-bounded by the explicit encoding length C(h) via a self-delimiting prefix code.
read the original abstract
While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting, and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To operationalize these bounds, we propose the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization. Progressive RLVR empirically retains 84-97% performance of standard LoRA fine-tuning while producing models that are 14,796x more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: mathematical problem-solving, programming, general-knowledge reasoning, and Text-to-SQL. Our bounds exceed the accuracy of the base model by 9-51% and lie within 6-11% of the accuracy of the fine-tuned models.
Figures
Reference graph
Works this paper leans on
-
[1]
On-policy distillation of language models: Learning from self-generated mistakes
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. On-policy distillation of language models: Learning from self-generated mistakes. InThe twelfth international conference on learning representations, 2024
2024
-
[2]
User-friendly introduction to pac-bayes bounds.Foundations and Trends in Machine Learning, 17(2):174–303, 2024
Pierre Alquier. User-friendly introduction to pac-bayes bounds.Foundations and Trends in Machine Learning, 17(2):174–303, 2024
2024
-
[3]
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang. Stronger generalization bounds for deep nets via a compression approach. InInternational conference on machine learning, pages 254–263. PMLR, 2018
2018
-
[4]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback.arXiv preprint arXiv:2204.05862, 2022
Pith/arXiv arXiv 2022
-
[5]
Klaudia Bałazy, Mohammadreza Banaei, Karl Aberer, and Jacek Tabor. Lora-xs: Low-rank adaptation with extremely small number of parameters.arXiv preprint arXiv:2405.17604, 2024
Pith/arXiv arXiv 2024
-
[6]
Rademacher and gaussian complexities: Risk bounds and structural results.Journal of machine learning research, 3(Nov):463–482, 2002
Peter L Bartlett and Shahar Mendelson. Rademacher and gaussian complexities: Risk bounds and structural results.Journal of machine learning research, 3(Nov):463–482, 2002
2002
-
[7]
Learnability and the vapnik-chervonenkis dimension.Journal of the ACM (JACM), 36(4):929–965, 1989
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth. Learnability and the vapnik-chervonenkis dimension.Journal of the ACM (JACM), 36(4):929–965, 1989
1989
-
[8]
Shiyi Cao, Sumanth Hegde, Dacheng Li, Tyler Griggs, Shu Liu, Eric Tang, Jiayi Pan, Xingyao Wang, Akshay Malik, Graham Neubig, et al. Skyrl-v0: Train real-world long-horizon agents via reinforcement learning.arXiv preprint arXiv:2502.02789, 2025. 10
Pith/arXiv arXiv 2025
-
[9]
Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al. Minimax-m1: Scaling test-time compute efficiently with lightning attention.arXiv preprint arXiv:2506.13585, 2025
Pith/arXiv arXiv 2025
-
[10]
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V Le, Sergey Levine, and Yi Ma. Sft memorizes, rl generalizes: A comparative study of foundation model post-training.arXiv preprint arXiv:2501.17161, 2025
Pith/arXiv arXiv 2025
-
[11]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv:1803.05457v1, 2018
Pith/arXiv arXiv 2018
-
[12]
Process reinforcement through implicit rewards
Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Yuchen Zhang, Jiacheng Chen, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, et al. Process reinforcement through implicit rewards. arXiv preprint arXiv:2502.01456, 2025
Pith/arXiv arXiv 2025
-
[13]
The design and operation of CloudLab
Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. The design and operation of CloudLab. InProceedings of the USENIX A...
2019
-
[14]
Soft adaptive policy optimization.arXiv preprint arXiv:2511.20347, 2025
Chang Gao, Chujie Zheng, Xiong-Hui Chen, Kai Dang, Shixuan Liu, Bowen Yu, An Yang, Shuai Bai, Jingren Zhou, and Junyang Lin. Soft adaptive policy optimization.arXiv preprint arXiv:2511.20347, 2025
Pith/arXiv arXiv 2025
-
[15]
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir. Size-independent sample complexity of neural networks. InConference on learning theory, pages 297–299. PMLR, 2018
2018
-
[16]
A tight excess risk bound via a unified pac-bayesian– rademacher–shtarkov–mdl complexity
Peter D Grünwald and Nishant A Mehta. A tight excess risk bound via a unified pac-bayesian– rademacher–shtarkov–mdl complexity. InAlgorithmic learning theory, pages 433–465. PMLR, 2019
2019
-
[17]
Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N
Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H. Cho...
2023
-
[18]
US Government Printing Office, 1954
Emil Julius Gumbel.Statistical theory of extreme values and some practical applications: a series of lectures, volume 33. US Government Printing Office, 1954
1954
-
[19]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025
Pith/arXiv arXiv 2025
-
[20]
Evalu- ation and mitigation of the limitations of large language models in clinical decision-making
Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, et al. Evalu- ation and mitigation of the limitations of large language models in clinical decision-making. Nature medicine, 30(9):2613–2622, 2024
2024
-
[21]
Song Han, Huizi Mao, and William J Dally. Deep compression: Compressing deep neural net- works with pruning, trained quantization and huffman coding.arXiv preprint arXiv:1510.00149, 2015
Pith/arXiv arXiv 2015
-
[22]
Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang. Parameter-efficient fine-tuning for large models: A comprehensive survey.arXiv preprint arXiv:2403.14608, 2024
Pith/arXiv arXiv 2024
-
[23]
Soufiane Hayou, Nikhil Ghosh, and Bin Yu. Plop: Precise lora placement for efficient finetuning of large models.arXiv preprint arXiv:2506.20629, 2025. 11
Pith/arXiv arXiv 2025
-
[24]
Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[25]
Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[26]
Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. Cuad: An expert-annotated nlp dataset for legal contract review.arXiv preprint arXiv:2103.06268, 2021
Pith/arXiv arXiv 2021
-
[27]
Nils Holzenberger and Benjamin Van Durme. Factoring statutory reasoning as language understanding challenges.arXiv preprint arXiv:2105.07903, 2021
Pith/arXiv arXiv 2021
-
[28]
Chuxuan Hu, Yuxuan Zhu, Antony Kellermann, Caleb Biddulph, Suppakit Waiwitlikhit, Jason Benn, and Daniel Kang. Breaking barriers: Do reinforcement post training gains transfer to unseen domains?arXiv preprint arXiv:2506.19733, 2025
arXiv 2025
-
[29]
Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022
2022
-
[30]
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. InAdvances in Neural Information Processing Systems, 2023
2023
-
[31]
Junseo Hwang, Wonguk Cho, and Taesup Kim. Pica: Parameter-efficient fine-tuning with column space projection.arXiv preprint arXiv:2505.20211, 2025
Pith/arXiv arXiv 2025
-
[32]
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016
Pith/arXiv arXiv 2016
-
[33]
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams.arXiv preprint arXiv:2009.13081, 2020
Pith/arXiv arXiv 2009
-
[34]
Three approaches to the quantitative definition ofinformation’.Problems of information transmission, 1(1):1–7, 1965
Andrei N Kolmogorov. Three approaches to the quantitative definition ofinformation’.Problems of information transmission, 1(1):1–7, 1965
1965
-
[35]
Yuta Koreeda and Christopher D Manning. Contractnli: A dataset for document-level natural language inference for contracts.arXiv preprint arXiv:2110.01799, 2021
Pith/arXiv arXiv 2021
-
[36]
The performance of universal encoding.IEEE Transactions on Information Theory, 27(2):199–207, 1981
Raphail Krichevsky and Victor Trofimov. The performance of universal encoding.IEEE Transactions on Information Theory, 27(2):199–207, 1981
1981
-
[37]
Omnisql: Synthesizing high-quality text-to-sql data at scale.arXiv preprint arXiv:2503.02240, 2025
Haoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang, Jing Zhang, Fuxin Jiang, Shuai Wang, Tieying Zhang, Jianjun Chen, Rui Shi, et al. Omnisql: Synthesizing high-quality text-to-sql data at scale.arXiv preprint arXiv:2503.02240, 2025
Pith/arXiv arXiv 2025
-
[38]
Numinamath
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu. Numinamath. [https://huggingface.co/AI-MO/NuminaMath-CoT](https://github.com/ project-numina/aimo-progress-prize/blob/main/report/nu...
2024
-
[39]
Uni-lora: One vector is all you need.arXiv preprint arXiv:2506.00799, 2025
Kaiyang Li, Shaobo Han, Qing Su, Wei Li, Zhipeng Cai, and Shihao Ji. Uni-lora: One vector is all you need.arXiv preprint arXiv:2506.00799, 2025
arXiv 2025
-
[40]
Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, and Tuo Zhao. Loftq: Lora-fine-tuning-aware quantization for large language models.arXiv preprint arXiv:2310.08659, 2023. 12
Pith/arXiv arXiv 2023
-
[41]
Claudette: an automated detector of potentially unfair clauses in online terms of service.Artificial Intelligence and Law, 27:117–139, 2019
Marco Lippi, Przemysław Pałka, Giuseppe Contissa, Francesca Lagioia, Hans-Wolfgang Mick- litz, Giovanni Sartor, and Paolo Torroni. Claudette: an automated detector of potentially unfair clauses in online terms of service.Artificial Intelligence and Law, 27:117–139, 2019
2019
-
[42]
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning.arXiv preprint arXiv:2007.08124, 2020
Pith/arXiv arXiv 2007
-
[43]
Skyrl-sql: Multi-turn sql data agents via rl
Shu Liu, Alan Zhu, Sumanth Hegde, Shiyi Cao, Shuo Yuan, Samion Suwito, Tyler Griggs, Matei Zaharia, Joseph E Gonzalez, and Ion Stoica. Skyrl-sql: Multi-turn sql data agents via rl. InFirst Workshop on Multi-Turn Interactions in Large Language Models
-
[44]
Understanding r1-zero-like training: A critical perspective.arXiv preprint arXiv:2503.20783, 2025
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. Understanding r1-zero-like training: A critical perspective.arXiv preprint arXiv:2503.20783, 2025
Pith/arXiv arXiv 2025
-
[45]
Pac-bayes compression bounds so tight that they can explain generalization
Sanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski, Micah Goldblum, and An- drew G Wilson. Pac-bayes compression bounds so tight that they can explain generalization. Advances in Neural Information Processing Systems, 35:31459–31473, 2022
2022
-
[46]
Non-vacuous generalization bounds for large language models.arXiv preprint arXiv:2312.17173, 2023
Sanae Lotfi, Marc Finzi, Yilun Kuang, Tim GJ Rudner, Micah Goldblum, and Andrew Gor- don Wilson. Non-vacuous generalization bounds for large language models.arXiv preprint arXiv:2312.17173, 2023
Pith/arXiv arXiv 2023
-
[47]
On-policy distillation.Thinking Machines Lab: Con- nectionism, 2025
Kevin Lu and Thinking Machines Lab. On-policy distillation.Thinking Machines Lab: Con- nectionism, 2025. doi: 10.64434/tml.20251026. https://thinkingmachines.ai/blog/on-policy- distillation
-
[48]
General- Reasoner: Advancing LLM reasoning across all domains
Xueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun MA, and Wenhu Chen. General- Reasoner: Advancing LLM reasoning across all domains. InThe Thirty-ninth Annual Con- ference on Neural Information Processing Systems, 2025. URLhttps://openreview.net/ forum?id=pBFVoll8Xa
2025
-
[49]
Learning to reason in 13 parameters.arXiv preprint arXiv:2602.04118, 2026
John X Morris, Niloofar Mireshghallah, Mark Ibrahim, and Saeed Mahloujifar. Learning to reason in 13 parameters.arXiv preprint arXiv:2602.04118, 2026
arXiv 2026
-
[50]
Uniform convergence may be unable to explain generalization in deep learning.Advances in neural information processing systems, 32, 2019
Vaishnavh Nagarajan and J Zico Kolter. Uniform convergence may be unable to explain generalization in deep learning.Advances in neural information processing systems, 32, 2019
2019
-
[51]
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro. Norm-based capacity control in neural networks. InConference on learning theory, pages 1376–1401. PMLR, 2015
2015
-
[52]
Phuc Minh Nguyen, Chinh D La, Duy MH Nguyen, Nitesh V Chawla, Binh T Nguyen, and Khoa D Doan. The reasoning boundary paradox: How reinforcement learning constrains language models.arXiv preprint arXiv:2510.02230, 2025
arXiv 2025
-
[53]
Question answering for privacy policies: Combining computational and legal perspectives
Abhilasha Ravichander, Alan W Black, Shomir Wilson, Thomas Norton, and Norman Sadeh. Question answering for privacy policies: Combining computational and legal perspectives. arXiv preprint arXiv:1911.00841, 2019
Pith/arXiv arXiv 1911
-
[54]
More pac-bayes bounds: From bounded losses, to losses with general tail behaviors, to anytime validity.Journal of Machine Learning Research, 25(110):1–43, 2024
Borja Rodriguez-Galvez, Ragnar Thobaben, and Mikael Skoglund. More pac-bayes bounds: From bounded losses, to losses with general tail behaviors, to anytime validity.Journal of Machine Learning Research, 25(110):1–43, 2024
2024
-
[55]
Lora without regret.Thinking Machines Lab: Connectionism, 2025
John Schulman and Thinking Machines Lab. Lora without regret.Thinking Machines Lab: Connectionism, 2025. doi: 10.64434/tml.20250929. https://thinkingmachines.ai/blog/lora/
-
[56]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv preprint arXiv:2402.03300, 2024
Pith/arXiv arXiv 2024
-
[57]
A formal theory of inductive inference
Ray J Solomonoff. A formal theory of inductive inference. part i.Information and control, 7(1): 1–22, 1964. 13
1964
-
[58]
Kimi k2.6: From code to creation, from one to many, April 2026
Kimi Team. Kimi k2.6: From code to creation, from one to many, April 2026. URL https: //www.kimi.com/ai-models/kimi-k2-6
2026
-
[59]
Kimi k2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025
Kimi Team, Yifan Bai, Yiping Bao, Y Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, et al. Kimi k2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025
Pith/arXiv arXiv 2025
-
[60]
Qwen3.5: Accelerating productivity with native multimodal agents, February
Qwen Team. Qwen3.5: Accelerating productivity with native multimodal agents, February
-
[61]
The power of rlvr: Training a leading sql rea- soning model on databricks, July 2025
The Databricks AI Research Team. The power of rlvr: Training a leading sql rea- soning model on databricks, July 2025. URL https://www.databricks.com/blog/ power-rlvr-training-leading-sql-reasoning-model-databricks
2025
-
[62]
On the generalization gap in reparameterizable reinforcement learning
Huan Wang, Stephan Zheng, Caiming Xiong, and Richard Socher. On the generalization gap in reparameterizable reinforcement learning. InInternational Conference on Machine Learning, pages 6648–6658. PMLR, 2019
2019
-
[63]
Parameter-efficient fine-tuning in large language models: a survey of methodologies.Artificial Intelligence Review, 58(8):227, 2025
Luping Wang, Sheng Chen, Linnan Jiang, Shu Pan, Runze Cai, Sen Yang, and Fei Yang. Parameter-efficient fine-tuning in large language models: a survey of methodologies.Artificial Intelligence Review, 58(8):227, 2025
2025
-
[64]
Steven H Wang, Antoine Scardigli, Leonard Tang, Wei Chen, Dimitry Levkin, Anya Chen, Spencer Ball, Thomas Woodside, Oliver Zhang, and Dan Hendrycks. Maud: An expert-annotated legal nlp dataset for merger agreement understanding.arXiv preprint arXiv:2301.00876, 2023
Pith/arXiv arXiv 2023
-
[65]
Openclaw-rl: Train any agent simply by talking.arXiv preprint arXiv:2603.10165, 2026
Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, and Ling Yang. Openclaw-rl: Train any agent simply by talking.arXiv preprint arXiv:2603.10165, 2026
Pith/arXiv arXiv 2026
-
[66]
Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I Wang. Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution.arXiv preprint arXiv:2502.18449, 2025
Pith/arXiv arXiv 2025
-
[67]
The creation and analysis of a website privacy policy corpus
Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N Cameron Russell, et al. The creation and analysis of a website privacy policy corpus. InProceedings of the 54th Annual Meeting of the Association for Computational Lingui...
2016
-
[68]
Parameter-efficient fine-tuning methods for pretrained language models: A critical review and assessment.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
Lingling Xu, Haoran Xie, S Joe Qin, Xiaohui Tao, and Fu Lee Wang. Parameter-efficient fine-tuning methods for pretrained language models: A critical review and assessment.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026
2026
-
[69]
Qa-lora: Quantization-aware low-rank adaptation of large language models
Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang, and Qi Tian. Qa-lora: Quantization-aware low-rank adaptation of large language models. InThe Twelfth International Conference on Learning Representations, 2023
2023
-
[70]
Kodcode: A diverse, challenging, and verifiable synthetic dataset for coding
Zhangchen Xu, Yang Liu, Yueqin Yin, Mingyuan Zhou, and Radha Poovendran. Kodcode: A diverse, challenging, and verifiable synthetic dataset for coding. 2025. URL https://arxiv. org/abs/2503.02951
Pith/arXiv arXiv 2025
-
[71]
Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[72]
Your efficient rl framework secretly brings you off-policy rl training, August 2025
Feng Yao, Liyuan Liu, Dinghuai Zhang, Chengyu Dong, Jingbo Shang, and Jianfeng Gao. Your efficient rl framework secretly brings you off-policy rl training, August 2025. URL https://fengyao.notion.site/off-policy-rl
2025
-
[73]
Trac: Tensor-train based across-layer compression for parameter-efficient fine-tuning
Bangguo Ye, Yuanwei Zhang, and Xiaoqun Zhang. Trac: Tensor-train based across-layer compression for parameter-efficient fine-tuning. InThe Fourteenth International Conference on Learning Representations, 2025. 14
2025
-
[74]
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. Dapo: An open-source llm reinforcement learning system at scale.arXiv preprint arXiv:2503.14476, 2025
Pith/arXiv arXiv 2025
-
[75]
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, and Gao Huang. Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? arXiv preprint arXiv:2504.13837, 2025
Pith/arXiv arXiv 2025
-
[76]
Understanding deep learning requires rethinking generalization.arXiv preprint arXiv:1611.03530, 2016
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization.arXiv preprint arXiv:1611.03530, 2016
Pith/arXiv arXiv 2016
-
[77]
Crossspectra: Exploiting cross-layer smoothness for parameter-efficient fine-tuning
Yifei Zhang, Hao Zhu, Junhao Dong, Haoran Shi, Ziqiao Meng, Piotr Koniusz, and Han Yu. Crossspectra: Exploiting cross-layer smoothness for parameter-efficient fine-tuning. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
2025
-
[78]
When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings
Lucia Zheng, Neel Guha, Brandon R Anderson, Peter Henderson, and Daniel E Ho. When does pretraining help? assessing self-supervised learning for law and the casehold dataset of 53,000+ legal holdings. InProceedings of the eighteenth international conference on artificial intelligence and law, pages 159–168, 2021
2021
-
[79]
Index (KT)
Sebastian Zimmeck, Peter Story, Daniel Smullen, Abhilasha Ravichander, Ziqi Wang, Joel R Reidenberg, N Cameron Russell, and Norman Sadeh. Maps: Scaling privacy compliance analysis to a million apps.Proc. Priv. Enhancing Tech., 2019:66, 2019. A Proof of Generalization Bounds A.1 Notation We summarize the notation in Table 1. A.2 Proof of Theorem 1 Proof.We...
2019
-
[2026]
URLhttps://qwen.ai/blog?id=qwen3.5
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.