REVIEW 4 major objections 5 minor 2 cited by
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper proves that in weak-to-strong generalization the strong student's error and calibration stay close to the weak teacher's, bounded by their disagreement, and that in a convex regression setting the student beats the teacher by…
desk verdict The paper's advertised classification and calibration bounds do not hold up — Theorem 4.2 is false as stated — but the Section 5 regression extension of Charikar et al. to KL divergence is genuine and worth preserving. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the KL divergence $d_P$ between model outputs, used both as the training objective and as the metric that separates the actors: the weakly-supervised strong model $F_{sw}$ is the projection of the weak teacher $F_w$ onto the convex set of strong-model functions $V_s = \{f \circ h_s : f \in F_s\}$. In the classification theorems the load is carried by the Donsker–Varadhan variational formula, which turns the KL divergence into a supremum over subgaussian functions, and by total variation with the Bretagnolle–Huber inequality for calibration. In the regression theorems the projection property plus a first-order Taylor expansion of KL near the projection shows the leading term is non-negative, which yields the improvement inequality.
What would settle it
Train a family of strong students on weak labels in a non-convex setting (any full neural-network fine-tuning), measure $d_P(F^*, F_{sw})$ and $d_P(F_{sw}, F_w)$, and test whether $d_P(F^*, F_{sw}) \le d_P(F^*, F_w) - d_P(F_{sw}, F_w)$; the paper's own epoch-ablation, where GPT-2-XL accuracy falls below the weak teacher's as training continues, is already a qualitative violation of that inequality and shows the assumptions, not the conclusion, are where the claim lives.
Extended reading notes
Core claim
The central discovery is a set of quantitative relations governing weak-to-strong generalization. For classification, Theorem 4.1 states $|d_P(F^*, F_{sw}) - d_P(F^*, F_w)| \le O(\sqrt{d_P(F_w, F_{sw})})$, so the strong model's error cannot stray far from the weak teacher's except through the disagreement that the training minimizes. Theorem 4.2 states $|\mathrm{MCE}(F_{sw}) - \mathrm{MCE}(F_w)| \le 2\sqrt{1 - \exp(-d_P(F_w, F_{sw}))}$, binding the student's calibration to the teacher's. For regression with reverse KL divergence, Theorem 5.1 yields $d_P(F^*, F_{sw}) \le d_P(F^*, F_w) - d_P(F_{sw}, F_w)$, which says the student beats the teacher by at least their disagreement; Theorem 5.2 relaxes realizability at the price of extra $O(\sqrt{\varepsilon})$ and $O(\sqrt{C_{F_s}/n} + \sqrt{\log(1/\delta)/n})$ terms.
Load-bearing premise
The regression guarantee assumes the strong model's hypothesis class is convex and already contains the ground-truth function, which turns weak-to-strong training into a Bregman projection; the proof and the inequality collapse without that assumption, and the open-set, full-network LLM fine-tuning that motivates the paper is neither convex nor guaranteed to realize the target.
Editorial extensions
If this is right
- The strong model inherits the weak teacher's miscalibration: when the student nearly mimics the teacher ($d_P(F_w, F_{sw}) \to 0$), the student's calibration error converges to the teacher's, so a poorly calibrated teacher yields a poorly calibrated student.
- Over-optimizing the W2SG objective erases the benefit: as training epochs grow, the student's accuracy and calibration approach the teacher's, and in the GPT-2 experiments the largest student's accuracy drops below the teacher's.
- In regression, the attainable gain is capped by the optimization objective itself: the improvement over the teacher is at most $d_P(F_w, F_{sw})$, the very disagreement the objective minimizes.
- Choosing a stronger, well-calibrated teacher improves both bounds: $d_P(F^*, F_w)$ and $\mathrm{MCE}(F_w)$ appear additively in the student's upper bounds.
- Non-realizability and finite samples add explicit costs $O(\sqrt{\varepsilon})$ and $O(\sqrt{C_{F_s}/n} + \sqrt{\log(1/\delta)/n})$, making representation quality and sample size first-order factors.
Reading between the lines
- A testable early-stopping rule falls out: monitor $d_P(F_w, F_{sw})$ during training and check when the empirical inequality $d_P(F^*, F_{sw}) + d_P(F_{sw}, F_w) \le d_P(F^*, F_w)$ begins to fail; the experiments suggest this happens as accuracy peaks and then declines.
- The same KL-based bounds should transfer to knowledge distillation with imperfect teachers, where teacher–student disagreement appears naturally as the objective.
- The $\gamma$ lower bound on output probabilities makes the constants in Theorem 4.1 grow like $1/\gamma$, so for near-deterministic predictions the universal bound becomes numerically vacuous even though it is formally true.
- Reverse KL gives a clean inequality while forward KL requires an extra weighted Itakura–Saito condition, hinting that the improvement guarantee depends on the direction of the Bregman projection, not on KL divergence generically.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies weak-to-strong generalization (W2SG) theoretically. In the classification setting, it claims universal bounds on the KL-divergence gap between the strong student and the ground truth in terms of the weak-teacher error and teacher–student disagreement (Theorem 4.1), and bounds on the marginal calibration error difference between strong and weak models (Theorem 4.2). It validates these claims with small-scale GPT-2 and Pythia experiments. In the regression setting, it extends Charikar et al. (2024) to reverse-KL output-distribution divergence, proving under convexity and realizability that the strong student's error is below the weak teacher's by at least their disagreement (Theorem 5.1), with a non-realizable finite-sample analogue (Theorem 5.2), supported by synthetic experiments.
Significance. If established, the classification bounds would provide a general explanation of how weak-model quality and the optimization objective jointly determine the strong model's generalization and calibration, and the regression result would be a nontrivial asymmetric-loss extension of Charikar et al. (2024). The paper also ships experiments and a clearly written limitations section. However, the two central classification theorems are not valid as stated: Theorem 4.1's proof is logically defective, and Theorem 4.2 is false, as shown by a simple binary counterexample that satisfies all assumptions. The regression theorems are more plausible but rest on convexity/realizability assumptions and on several proof details that are either under-specified or misreferenced. The overall contribution therefore does not currently meet the standard for publication.
major comments (4)
- [B.1 (Theorem 4.1)] The variational-argument proof is invalid because the test function g is not fixed across the two expectations in Eq. (12). In the third step, the paper defines g from a distribution Pg and then substitutes Pg = Psw to obtain E_{Psw}[g] = dP(F*,Fsw) and Pg = Pw to obtain E_{Pw}[g] = dP(F*,Fw). These two substitutions use different functions g_sw and g_w, since g depends on pg. Inequality (12) holds only for a single, fixed g, so the conclusion |dP(F*,Fsw) − dP(F*,Fw)| ≤ sqrt(2σ^2 dP(Fw,Fsw)) does not follow. This is a load-bearing defect: all of the classification generalization claims in Section 4.1 are consequences of this theorem.
- [B.3 (Theorem 4.2)] Theorem 4.2 is false as stated. The proof equates MCE(F) with E_X ||F(X) − F_b(X)||_1, where F_b(x) = P(Y | X=x), but Definition 4.1 conditions on the scalar [F(X)]_i, a coarser random variable than X. These two quantities are generally different, and the triangle-inequality step does not hold for the definitional MCE. A concrete counterexample: let X be Bernoulli(1/2), Y = X, Fw(0) = [0.4, 0.6], Fw(1) = [0.6, 0.4], and let Fs be the convex class of constant distributions, with outputs bounded below by γ = 0.4. The W2SG minimizer of dP(Fw, Fsw) over Fs is Fsw ≡ [0.5, 0.5]. Then MCE(Fw) = 1.2 and MCE(Fsw) = 0, while dP(Fw, Fsw) ≈ 0.020, so the right-hand side of Eq. (9) is approximately 0.28, which is far below 1.2. The theorem cannot be repaired by proof fixes alone; the statement itself must be revised.
- [5.1 (Theorem 5.1) and B.6 (Theorem 5.2)] There is an unresolved mismatch between the W2SG objective defined in Eq. (6), which uses dP(fw ◦ hw, f ◦ hs) (forward KL from the weak model), and the objective in Theorem 5.1, which minimizes dP(f ◦ hs, fw ◦ hw) (reverse KL). If Theorem 5.1 is intended as a result for a different loss, this needs to be stated explicitly and connected back to the W2SG framework. Moreover, in the proof of Theorem 5.2, Eq. (37) is justified by replacing F* with Fhat_sw in 'the final step of proof of Corollary B.1,' but Corollary B.1 is a forward-KL result requiring the additional weighted Itakura–Saito assumption, whereas Theorem 5.2 is stated for reverse KL; the correct reference should be Theorem 5.1. As written, the derivation of inequality (37) is not justified.
- [5.1 (Assumption 5.1 and normalization)] The regression setting treats model outputs as probability densities over X, but Assumption 5.1 requires convexity of the function class Fs before normalization. The probability-normalization step is nonlinear, so it is not automatic that the set of normalized output distributions is convex or that the Bregman projection argument applies to the actual optimization problem in Eq. (6)/(7). The paper should either prove that the normalized set inherits convexity, or state the convexity assumption directly on the output distributions. This is a technical but load-bearing point for the regression guarantees.
minor comments (5)
- [4.1] The discussion labels the two sides of Eq. (13) as 'lower bound' and 'upper bound,' but Eq. (13) is a single two-sided bound on |dP(F*,Fsw) − dP(F*,Fw)|; the labeling is misleading.
- [B.1] In the paragraph after Eq. (11), the text says 'we also define the probability distribution Psw(x) for Fw(x)'; this should presumably read 'for Fsw(x)'.
- [B.4] In the Taylor-expansion display after Eq. (17), the notation 'dKL(Fsw∥g)' is used where the preceding expression is dKL(g∥Fsw); this is a typo that makes the argument harder to follow.
- [4.1 / Theorem 4.1] The constant C1 = sqrt(2)/γ log(1/γ) is hidden in the O notation of Theorem 4.1. Since γ can be very small, the bound may be vacuous; the theorem statement should expose this dependence.
- [5.3] The synthetic experiments plot dP(F*,Fw) − dP(F*,Fsw) against dP(Fsw,Fw), which visually checks the reverse-KL identity, but they do not directly test the finite-sample claim in Theorem 5.2; the paper should clarify what exactly the experiments validate.
Circularity Check
No significant circularity: the theoretical bounds are derived from explicit assumptions and external prior results, and the calibration proof concern is a correctness issue rather than a circular reduction.
full rationale
The paper's main claims are not equivalent to their inputs by construction. Theorem 4.1 is proved for arbitrary Fw and Fsw via the Donsker-Varadhan variational formula and a subgaussian bound; it does not presuppose the optimization objective or a fitted value. Theorem 5.1 follows from the stated Convexity Assumption, Realizability, and the projection characterization of Fsw, using first-order Taylor expansions in the manner of Charikar et al. (2024), an external source; the inequality d_P(F*, Fsw) <= d_P(F*, Fw) - d_P(Fsw, Fw) is a mathematical consequence of that projection setup, not a restatement of an empirical fit. Theorem 5.2 adds standard uniform-convergence terms, again from stated assumptions. The experiments in Section 5.3 recompute the relevant d_P quantities after solving the same minimization problems; they illustrate the theorems rather than supply the theorems' content, so no fitted parameter is renamed as a prediction. The only self-citations, Yang et al. (2025), appear in related-work discussion and experimental hyperparameter choices, and are not load-bearing for any proof. The Limitations section candidly notes that the Section 5 assumptions may not match real LLM deployment, which is a scope caveat rather than circularity. A separate concern is that the proof of Theorem 4.2 (Section B.3) appears to replace Definition 4.1's conditioning on the scalar [F(x)]_i with conditioning on X via F_b(X); that is a potential correctness defect, not a circular reduction, and it does not affect the circularity score.
Assumptions & free parameters
assumptions (5)
- standard math Donsker-Varadhan variational formula and Hoeffding's lemma supply the subgaussian bound on the constructed test function g.
- domain assumption The model outputs are probability vectors bounded below by some gamma > 0.
- domain assumption Assumption 5.1: the strong model's task hypotheses Fs form a convex set.
- domain assumption Realizability: there exists fs in Fs with fs composed with hs equal to f* composed with h*.
- ad hoc to paper For forward KL results, the weighted Itakura-Saito divergence DWIS(F*||Fsw) <= 0 with weight Fsw - Fw must hold.
Cite this review
Pith. "Pith review of The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration." pith.science (2026). https://pith.science/paper/WABYMTGS
@misc{pith2026250201458,
author = {Pith},
title = {Pith review of: The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration},
year = {2026},
howpublished = {\url{https://pith.science/paper/WABYMTGS}},
note = {Machine review of arXiv:2502.01458}
}
read the original abstract
Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values. To deepen the understanding of this approach, we provide theoretical insights into its capabilities and limitations. First, in the classification setting, we establish upper and lower generalization error bounds for the strong model, identifying the primary limitations as stemming from the weak model's generalization error and the optimization objective itself. Additionally, we derive lower and upper bounds on the calibration error of the strong model. These theoretical bounds reveal two critical insights: (1) the weak model should demonstrate strong generalization performance and maintain well-calibrated predictions, and (2) the strong model's training process must strike a careful balance, as excessive optimization could undermine its generalization capability by over-relying on the weak supervision signals. Finally, in the regression setting, we extend the work of Charikar et al. (2024) to a loss function based on Kullback-Leibler (KL) divergence, offering guarantees that the strong student can outperform its weak teacher by at least the magnitude of their disagreement. We conduct sufficient experiments to validate our theory.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 2 Pith papers
-
On Weak-to-Strong Generalization and f-Divergence
Replacing cross-entropy with f-divergence losses in weak-to-strong generalization gives modest accuracy gains and improved label-noise tolerance, though the paper's theoretical equivalence result is constructed after ...
-
Contrastive Weak-to-strong Generalization
Contrastive decoding between pre- and post-alignment weak models generates better supervision samples, improving weak-to-strong generalization on AlpacaEval2 and Arena-Hard.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Aakriti Agrawal, Mucong Ding, Zora Che, Chenghao Deng, Anirudh Satheesh, John Langford, and Furong Huang. 2024. Ensemw2s: Can an ensemble of llms be leveraged to obtain a stronger llm? arXiv preprint arXiv:2410.04571
arXiv 2024
-
[3]
Gholamali Aminian, Mahed Abroshan, Mohammad Mahdi Khalili, Laura Toni, and Miguel Rodrigues. 2022. An information-theoretical approach to semi-supervised learning under covariate-shift. In International Conference on Artificial Intelligence and Statistics, pages 7433--7449
2022
-
[4]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, and 1 others. 2022 a . Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862
arXiv 2022
-
[5]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, and 1 others. 2022 b . Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073
arXiv 2022
-
[6]
Peter L Bartlett and Shahar Mendelson. 2002. Rademacher and gaussian complexities: Risk bounds and structural results. Journal of Machine Learning Research, 3:463--482
2002
-
[7]
Lucas Beyer, Xiaohua Zhai, Am \'e lie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. 2022. Knowledge distillation: A good teacher is patient and consistent. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10925--10934
2022
-
[8]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, and 1 others. 2023. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pages 2397--2430
2023
Show all 88 references
-
[9]
Yuheng Bu, Shaofeng Zou, and Venugopal V Veeravalli. 2020. Tightening mutual information-based bounds on generalization error. IEEE Journal on Selected Areas in Information Theory, 1(1):121--130
2020
-
[10]
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, and 1 others. 2023. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision. arXiv preprint arXiv...
2023 arXiv
-
[11]
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, J \'e r \'e my Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, and 1 others. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. T...
2023
-
[12]
Moses Charikar, Chirag Pabbaraju, and Kirankumar Shiragur. 2024. Quantifying the gain in weak-to-strong generalization. Advances in neural information processing systems
2024
-
[13]
Qi Chen, Changjian Shui, and Mario Marchand. 2021. Generalization bounds for meta-learning: An information-theoretic analysis. Advances in Neural Information Processing Systems, 34:25878--25890
2021
-
[14]
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji. 2023. A close look into the calibration of pre-trained language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1343--1367
2023
-
[15]
Lele Cheng, Xiangzeng Zhou, Liming Zhao, Dangwei Li, Hong Shang, Yun Zheng, Pan Pan, and Yinghui Xu. 2020. Weakly supervised learning with side information for noisy labeled images. In The European Conference on Computer Vision, pages 306--321
2020
-
[16]
P Chu and D Messerschmitt. 1982. A frequency weighted itakura-saito spectral distance measure. IEEE Transactions on Acoustics, Speech, and Signal Processing, 30(4):545--560
1982
-
[17]
Thomas M Cover. 1999. Elements of information theory. John Wiley & Sons
1999
-
[18]
Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. arXiv preprint arXiv:2003.07892
2020 arXiv
-
[19]
Inderjit S. Dhillon. 2007. Learning with bregman divergences. https://www.cs.utexas.edu/ inderjit/Talks/bregtut.pdf
2007
-
[20]
Monroe D Donsker and SR Srinivasa Varadhan. 1983. Asymptotic evaluation of certain markov process expectations for large time. iv. Communications on pure and applied mathematics, 36(2):183--212
1983
-
[21]
C \'e dric F \'e votte, Nancy Bertin, and Jean-Louis Durrieu. 2009. Nonnegative matrix factorization with the itakura-saito divergence: With application to music analysis. Neural computation, 21(3):793--830
2009
-
[22]
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024. A survey of confidence estimation and calibration in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Ling...
2024
-
[23]
Maurice G \"u nder, Nico Piatkowski, and Christian Bauckhage. 2022. Full kullback-leibler-divergence loss for hyperparameter-free label distribution learning. arXiv preprint arXiv:2209.02055
2022 arXiv
-
[24]
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International Conference on Machine Learning, pages 1321--1330
2017
-
[25]
Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2021. An overview of uncertainty calibration for text classification and the role of distillation. In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021), pages 289--306
2021
-
[26]
Jianyuan Guo, Hanting Chen, Chengcheng Wang, Kai Han, Chang Xu, and Yunhe Wang. 2024. Vision superalignment: Weak-to-strong generalization for vision foundation models. arXiv preprint arXiv:2402.03749
2024 arXiv
-
[27]
Yue Guo and Yi Yang. 2024. Improving weak-to-strong generalization with reliability-aware alignment. arXiv preprint arXiv:2406.19032
2024 arXiv
-
[28]
Fredrik Hellstr \"o m, Giuseppe Durisi, Benjamin Guedj, and Maxim Raginsky. 2023. Generalization bounds: Perspectives from information theory and pac-bayes. arXiv preprint arXiv:2309.04381
2023 arXiv
-
[29]
Geoffrey Hinton. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531
2015 arXiv
-
[30]
Jeremy Howard and Sebastian Ruder. 2018. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328--339
2018
-
[31]
Ehsan Imani and Martha White. 2018. Improving regression performance with distributional losses. In International conference on machine learning, pages 2157--2166
2018
-
[32]
Fumitada Itakura. 1968. Analysis synthesis telephony based on the maximum likelihood method. Reports of the 6-th Int. Cong. Acoust
1968
-
[33]
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, and 1 others. 2023. Ai alignment: A comprehensive survey. arXiv preprint arXiv:2310.19852
2023 arXiv
-
[34]
HyunJin Kim, Xiaoyuan Yi, Jing Yao, Jianxun Lian, Muhua Huang, Shitong Duan, JinYeong Bak, and Xing Xie. 2024. The road to artificial superintelligence: A comprehensive survey of superalignment. arXiv preprint arXiv:2412.16468
2024 arXiv
-
[35]
Yoshiaki Kitazawa. 2025. Bounds on \ l\_p\ errors in density ratio estimation via \ f\ -divergence loss functions. In The Thirteenth International Conference on Learning Representations
2025
-
[36]
Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon. 2018. Accurate uncertainties for deep learning using calibrated regression. In International conference on machine learning, pages 2796--2804
2018
-
[37]
Meelis Kull, Miquel Perello Nieto, Markus K \"a ngsepp, Telmo Silva Filho, Hao Song, and Peter Flach. 2019. Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration. Advances in neural information processing systems, 32
2019
-
[38]
Ananya Kumar, Percy S Liang, and Tengyu Ma. 2019. Verified uncertainty calibration. Advances in Neural Information Processing Systems, 32
2019
-
[39]
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. 2022. Fine-tuning can distort pretrained features and underperform out-of-distribution. In International Conference on Learning Representations
2022
-
[40]
Hunter Lang, David Sontag, and Aravindan Vijayaraghavan. 2024. Theoretical analysis of weak-to-strong generalization. Advances in neural information processing systems
2024
-
[41]
Michel Ledoux and Michel Talagrand. 2013. Probability in Banach Spaces: isoperimetry and processes. Springer Science & Business Media
2013
-
[42]
Sebastian Lee, Sebastian Goldt, and Andrew Saxe. 2021. Continual learning in the teacher-student setup: Impact of task similarity. In International Conference on Machine Learning, pages 6109--6119
2021
-
[43]
Shaojie Li, Bowei Zhu, and Yong Liu. 2024. Algorithmic stability unleashed: Generalization bounds with unbounded losses. In Forty-first International Conference on Machine Learning
2024
-
[44]
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, and 1 others. 2023. Holistic evaluation of language models. Transactions on Machine Learning Research
2023
-
[45]
Lydia T Liu, Max Simchowitz, and Moritz Hardt. 2019. The implicit fairness criterion of unconstrained learning. In International Conference on Machine Learning, pages 4051--4060
2019
-
[46]
Yuejiang Liu and Alexandre Alahi. 2024. Co-supervised learning: Improving weak-to-strong generalization with hierarchical mixture of experts. arXiv preprint arXiv:2402.15505
2024 arXiv
-
[47]
Tambet Matiisen, Avital Oliver, Taco Cohen, and John Schulman. 2019. Teacher--student curriculum learning. IEEE transactions on neural networks and learning systems, 31(9):3732--3740
2019
-
[48]
Alireza Mehrtash, William M Wells, Clare M Tempany, Purang Abolmaesumi, and Tina Kapur. 2020. Confidence calibration and predictive uncertainty estimation for deep medical image segmentation. IEEE transactions on medical imaging, 39(12):3868--3878
2020
-
[49]
Zhong Meng, Jinyu Li, Yong Zhao, and Yifan Gong. 2019. Conditional teacher-student learning. In IEEE International Conference on Acoustics, Speech and Signal Processing, pages 6445--6449
2019
-
[50]
Gabriel Meseguer-Brocal, Alice Cohen-Hadria, and Geoffroy Peeters. 2019. Dali: A large dataset of synchronized audio, lyrics and notes, automatically created using teacher-student machine learning paradigm. arXiv preprint arXiv:1906.10606
2019 arXiv
-
[51]
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI conference on artificial intelligence, pages 2901 -- 2907
2015
-
[52]
OpenAI. 2024. https://openai.com/index/introducing-superalignment/ Introducing superalignment
2024
-
[53]
Maxime Oquab, L \'e on Bottou, Ivan Laptev, and Josef Sivic. 2015. Is object localization for free?-weakly-supervised learning with convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 685--694
2015
-
[54]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing sys...
2022
-
[55]
Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari. 2017. Training object class detectors with click supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6374--6383
2017
-
[56]
Martin Pawelczyk, Lillian Sun, Zhenting Qi, Aounon Kumar, and Himabindu Lakkaraju. 2024. Generalizing trust: Weak-to-strong trustworthiness in language models. arXiv preprint arXiv:2501.00418
2024 arXiv
-
[57]
Ankit Pensia, Varun Jog, and Po-Ling Loh. 2018. Generalization error bounds for noisy, iterative algorithms. In 2018 IEEE International Symposium on Information Theory, pages 546--550
2018
-
[58]
Geoff Pleiss, Manish Raghavan, Felix Wu, Jon Kleinberg, and Kilian Q Weinberger. 2017. On fairness and calibration. Advances in neural information processing systems, 30
2017
-
[59]
Dani Prasetyawan and Nakamoto Takamichi. 2020. Sensory evaluation of odor approximation using nmf with kullback-leibler divergence and itakura-saito divergence in mass spectrum space. Journal of The Electrochemical Society, 167(16):167520
2020
-
[60]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, and 1 others. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9
2019
-
[61]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36
2024
-
[62]
Alexander Ratner, Stephen H Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher R \'e . 2020. Snorkel: rapid training data creation with weak supervision. The VLDB Journal, 29(2):709--730
2020
-
[63]
Rebecca Roelofs, Nicholas Cain, Jonathon Shlens, and Michael C Mozer. 2022. Mitigating bias in calibration error estimation. In International Conference on Artificial Intelligence and Statistics, pages 4036--4054
2022
-
[64]
Daniel Russo and James Zou. 2016. Controlling bias in adaptive data analysis using information theory. In Artificial Intelligence and Statistics, pages 1232--1240
2016
-
[65]
Jitao Sang, Yuhang Wang, Jing Zhang, Yanxu Zhu, Chao Kong, Junhong Ye, Shuyu Wei, and Jinlin Xiao. 2024. Improving weak-to-strong generalization with scalable oversight and ensemble learning. arXiv preprint arXiv:2402.00667
2024 arXiv
-
[66]
Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025
2023 arXiv
-
[67]
Rui Shu, Hung H Bui, Hirokazu Narui, and Stefano Ermon. 2018. A dirt-t approach to unsupervised domain adaptation. arXiv preprint arXiv:1802.08735
2018 arXiv
-
[68]
Seamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Ya'acov Ritov, Mikhail Yurochkin, and Yuekai Sun. 2024. A statistical framework for weak-to-strong generalization. arXiv preprint arXiv:2405.16236
2024 arXiv
-
[69]
Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee. 2022. Learning from noisy labels with deep neural networks: A survey. IEEE transactions on neural networks and learning systems, 34(11):8135--8153
2022
-
[70]
Huayi Tang and Yong Liu. 2023. Information-theoretic generalization bounds for transductive learning and its applications. arXiv preprint arXiv:2311.04561
2023 arXiv
-
[71]
Antti Tarvainen and Harri Valpola. 2017. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in neural information processing systems, 30
2017
-
[72]
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedin...
2023
-
[73]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and 1 others. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[74]
Dennis Ulmer, Jes Frellsen, and Christian Hardmeier. 2022. Exploring predictive uncertainty and calibration in nlp: A study on the impact of method & data scarcity. arXiv preprint arXiv:2210.15452
2022 arXiv
-
[75]
Jesper E Van Engelen and Holger H Hoos. 2020. A survey on semi-supervised learning. Machine learning, 109(2):373--440
2020
-
[76]
Ziqiao Wang and Yongyi Mao. 2023 a . Information-theoretic analysis of unsupervised domain adaptation. In International Conference on Learning Representations
2023
-
[77]
Ziqiao Wang and Yongyi Mao. 2023 b . Tighter information-theoretic generalization bounds from supersamples. In International Conference on Machine Learning
2023
-
[78]
David X Wu and Anant Sahai. 2025. Provable weak-to-strong generalization via benign overfitting. In International Conference on Learning Representations
2025
-
[79]
Aolin Xu and Maxim Raginsky. 2017. Information-theoretic analysis of generalization capability of learning algorithms. Advances in neural information processing systems, 30
2017
-
[80]
Wenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao, Yong Liu, Zhi Gong, Yankai Lin, and Ji-Rong Wen. 2025. Super (ficial)-alignment: Strong models may deceive weak models in weak-to-strong generalization. In International Conference on Learning Representations
2025
-
[81]
Xue Yang, Xiaojiang Yang, Jirui Yang, Qi Ming, Wentao Wang, Qi Tian, and Junchi Yan. 2021. Learning high-precision bounding box for rotated object detection via kullback-leibler divergence. In Advances in Neural Information Processing Systems
2021
-
[82]
Yuqing Yang, Yan Ma, and Pengfei Liu. 2024. Weak-to-strong reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8350--8367
2024
-
[83]
Ruimeng Ye, Yang Xiao, and Bo Hui. 2024. Weak-to-strong generalization beyond accuracy: a pilot study in safety, toxicity, and legal reasoning. arXiv preprint arXiv:2410.12621
2024 arXiv
-
[84]
Zhi-Hua Zhou. 2018. A brief introduction to weakly supervised learning. National science review, 5(1):44--53
2018
-
[85]
Chiwei Zhu, Benfeng Xu, Quan Wang, Yongdong Zhang, and Zhendong Mao. 2023. On the calibration of large language models and alignment. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9778--9795
2023
-
[86]
Wenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu, and Rui Wang. 2025. Weak-to-strong preference optimization: Stealing reward from weak aligned model. In International Conference on Learning Representations
2025
-
[87]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[88]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.