Pith. sign in

REVIEW 3 major objections 4 minor 51 references

Reliability-Safety Trade-off in AI Distillation: A Renormalization-Group Approach

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that repeated knowledge distillation is governed by a renormalization-group flow for a single 'hazard discrimination capability' $K$, with a tricritical phase diagram deciding whether safety is erased, transmitted, or…

desk verdict A clean formal core with a genuinely novel RG phase diagram, but the phase diagram's load-bearing inheritance rule needs justification and the empirical check is in-sample. read the letter →

arxiv 2608.08572 v1 pith:SM32CBRL submitted 2026-08-09 cond-mat.stat-mech

classification cond-mat.stat-mech PACS 05.10.Cc
keywords knowledgedistillationreliability-safetytrade-offhazarddiscriminationcapabilityrenormalizationgrouptricriticalpointrefusalcalibrationlargelanguagemodelsmultigenerational
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes that reliability and safety in AI models are two sides of one quantity, the hazard discrimination capability $K$, defined by $\operatorname{logit}(R)+\operatorname{logit}(S)=K$, where $R$ is the probability of answering an ordinary input and $S$ the probability of refusing a hazardous one. It models knowledge distillation as coarse-graining: the teacher's response bias acts as an effective field that reshapes the student's answer-versus-refusal free-energy landscape, and repeated distillation becomes an iterated renormalization-group-like transformation. The central claim is that this transformation has a tricritical phase diagram: depending on the student's per-generation capacity loss and the teacher-student alignment strength, the transmitted capability either flows to zero (erasure), reaches a stable nonzero value (persistent transmission), or is retained only when the original teacher's $K$ exceeds a threshold. If correct, this gives a quantitative criterion for whether safety-relevant behavior survives model compression and multigenerational distillation, and why very small students deteriorate. The predicted trade-off curve is checked against refusal-token threshold-sweep data, where $K$ stays approximately constant as the refusal threshold varies.

What carries the argument

The load-bearing object is the scalar $K=\operatorname{logit}(R)+\operatorname{logit}(S)$, the hazard discrimination capability, which fixes the entire reliability-safety frontier of a model. The argument is carried by a two-state Boltzmann machine whose macrostates are 'answer' and 'refusal'; integrating out hidden degrees of freedom leaves two free energies whose symmetric and antisymmetric combinations define $A$ and $K$. Distillation enters through the alignment $V=-\gamma y_T y_S$, which becomes an effective field $B_T(x)$ on the student, and repeated application with fixed loss fraction $a$ reduces the two-dimensional flow to a Ginzburg-Landau normal form $dK/dl\simeq \mu K + \gamma(C-2)K^3/[12(1+C)^2] + \gamma(C^2-13C+16)K^5/[960(1+C)^3]$ with $C=\cosh|\beta A_0|$ and $\mu=2\gamma/(1+C)-a$. The sign pattern of the cubic and quintic coefficients is what produces the one-, three-, and five-fixed-point regimes and the tricritical point.

What would settle it

In a multigenerational distillation run with fixed per-generation capacity reduction, measure $R$ and $S$ on held-out ordinary and hazardous inputs, extract $K_l=\operatorname{logit}(R_l)+\operatorname{logit}(S_l)$, and sweep the alignment strength $\gamma$: the theory predicts saturation to zero or to a nonzero fixed point whose onset scales as $|\gamma-\gamma_{\mathrm{PF}}|^{1/2}$ at a pitchfork boundary and $|\gamma-\gamma_{\mathrm{tri}}|^{1/4}$ at the tricritical point, so the absence of that saturation or of those scaling laws would refute the central claim.

Watch

Extended reading notes

Core claim

The central discovery is a trade-off identity for a Boltzmann-style two-state model of answering versus refusing, with inputs coarse-grained into ordinary and hazardous classes: $\operatorname{logit}(R)+\operatorname{logit}(S)=K$, with $K=2\beta Q$ called the hazard discrimination capability, while the orthogonal combination $\operatorname{logit}(S)-\operatorname{logit}(R)=2\beta A$ is set by the overall refusal tendency $A$. A single distillation step is shown to renormalize the student's effective free energy: after summing over the teacher's response, the teacher is compressed into the input-dependent field $B_T(x)=2\,\mathrm{arctanh}[m_T(x)\tanh\gamma]$, so the post-distillation capability is $K_{\mathrm{eff}}=K_0+B_T(1)-B_T(-1)$. Repeating this across generations with constant relative loss $a$ and constant pre-distillation tendency $A_0$, in the weak-alignment limit, yields the flow $dK/dl=-aK+4\gamma\sinh(K/2)/[\cosh(K/2)+\cosh(\beta A)]$. The fixed points of this flow organize into one-, three-, and five-fixed-point regimes separated by pitchfork and saddle-node bifurcations that meet at the tricritical point $(|\beta A_0|,a/\gamma)=(\mathrm{arcosh}\,2,2/3)$; near the pitchfork boundaries the nonzero fixed point onsets as $|\gamma-\gamma_{\mathrm{PF}}|^{1/2}$, and at the tricritical point as $|\gamma-\gamma_{\mathrm{tri}}|^{1/4}$.

Load-bearing premise

The phase diagram rests on the assumption that each generation loses the same fixed fraction $a$ of the teacher's discrimination capability before being taught, with a pre-distillation refusal tendency that is the same in every generation; the paper argues, but does not prove, that the phase structure survives for more general smooth inheritance rules.

Editorial extensions

If this is right

  • Changing a model's refusal threshold moves it along a curve of constant $K$; threshold tuning alone cannot push the attainable reliability-safety frontier outward.
  • A teacher that answers ordinary inputs and refuses hazardous ones raises the student's $K$ through the field contrast $B_T(1)-B_T(-1)$, so safety transmission is not set by task accuracy alone.
  • Repeated distillation saturates: the capability lands on $0$, on a stable nonzero fixed point, or on a threshold-dependent branch, so unbounded improvement or unbounded degradation of safety-relevant discrimination does not occur.
  • For small students the critical alignment strength is proportional to the per-generation model-size reduction $s$, meaning aggressive compression directly raises the bar for inheriting safety-relevant behavior.
  • The square-root and fourth-root onsets near the phase boundaries are concrete scaling predictions that a controlled multigenerational distillation experiment can test.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I would expect the same two-state, single-scalar logic to carry over to other inherited behavioral traits (sycophancy, overconfidence, refusal on benign inputs) whenever they can be encoded as a logit-sum invariant; each would then have its own 'K' and its own inheritance threshold.
  • A direct check the paper does not perform: measure $K_0(l+1)$ against $K_l$ across real distillation generations to test the constant-loss rule $a=\mathrm{const}$; strong curvature would shift the phase boundaries even if the tricritical skeleton survives.
  • The empirical support is a refusal-threshold sweep on two model configurations; a sharper test would vary distillation temperature, student width, and generation count together and compare the measured onset exponents with $1/2$ and $1/4$.
  • If the phase diagram is right, recursive training collapse is not merely a performance decline but an order-parameter transition in refusal behavior, visible as erasure or abrupt loss of hazard discrimination across generations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a two-state statistical-mechanics model of LLM reliability and safety during knowledge distillation. It defines reliability R and safety S from the response probabilities of a Boltzmann machine, derives the exact identity logit(R)+logit(S)=K with K=2βQ, and interprets K as a 'hazard discrimination capability'. A teacher-student alignment term V=-γ y_S y_T is then introduced, leading to an effective field and the rule K_eff=K0+B_T(1)-B_T(-1). For multigenerational distillation, the authors assume a constant per-generation HDC loss a through K0_{l+1}=(1-a)K_l, constant pre-distillation refusal tendency A0, weak alignment γ<<1, and a continuum limit in the generation index. This yields a two-dimensional RG-like flow for (K,A); adiabatic elimination of A gives a Ginzburg-Landau normal form. The fixed-point analysis produces one-, three-, and five-fixed-point regimes separated by pitchfork and saddle-node bifurcations that meet at a tricritical point (|βA0|,a/γ)=(arcosh 2,2/3), with square-root and fourth-root scaling laws. The empirical section digitizes refusal-token threshold sweeps from Ref. [43], computes K_i=logit(R_i)+logit(S_i) for each point, and reports that the K_i are approximately constant across thresholds.

Significance. The single-step content of the paper is clean and useful: Eq. (7) is derived exactly from the binary Boltzmann model, and SM §I gives a concrete microscopic example in which a is related to the fraction of retained hidden-layer eigenmodes. The bifurcation analysis around Eq. (16) is technically sound as a dynamical-systems exercise, and SM §II B correctly checks consistency of stability between the discrete and continuous generations. If the multigenerational assumptions were grounded in actual distillation dynamics, the phase diagram and the scaling predictions (18)-(19) would be a valuable minimal theory of behavioral inheritance. That grounding is presently missing: the only empirical test concerns the single-generation identity (7), not the RG flow, and the claimed robustness of the phase diagram to nonlinear inheritance rules is not established. The significance is therefore moderate: an elegant model with exact single-step results, whose central multi-generation predictions remain unvalidated.

major comments (3)
  1. [SM §III, Eq. (S70), Tables S2-S3] The claimed empirical support for Eq. (7) is partially circular. Each digitized (TPR,FPR) point is inserted into Eq. (S70), so K_i = logit(S_i)+logit(R_i) is defined by the data point itself; averaging the K_i and drawing the curve with K̄ does not provide a falsifiable goodness-of-fit test. The reported scatter σ_K=0.164 (main text Fig. 1) is not accompanied by error bars from digitization of Fig. 2 of Ref. [43], so it cannot distinguish a genuine threshold-independent K from digitization noise. Please provide a noise model with propagated uncertainties on the digitized R and S values, and compare the constant-K hypothesis against a null model in which K varies with the refusal threshold T.
  2. [Eq. (15), Eq. (16), SM §II C, Discussion] The phase diagram and the tricritical point are properties of the specific inheritance rule K0_{l+1}=(1-a)K_l with constant a and constant A0. The generalization claimed in SM §II C assumes that g(K) is odd with g2=g4=0 and that parameters exist with r=0, u=0, v<0 (SM Eq. S68); this is a condition on the effective flow, not a derivation from a distillation objective. Consider the smooth odd rule g(K)=(1-a)K-εK^3. The nonzero fixed-point equation becomes a/γ + (ε/γ)K^2 = f(K), with f defined in SM Eq. (S35). The saddle-node tangency condition is then 2(ε/γ)K = f'(K). For ε>0 this condition cannot be satisfied at K=0 when C>2, whereas in the ε=0 case the saddle-node line and the pitchfork line meet at (C,a/γ)=(2,2/3). Thus for any ε>0 the meeting point shifts or disappears, and the fourth-root scaling at (arcosh 2,2/3) is not a generic property of this family. The Discussion's robustness claim is therefore an overstatement. Either prove that a distillation objective produces inheritance functions with the required normal form, or explicitly restrict the central claims to the linear case and present the multigenerational phase diagram as an illustrative model prediction.
  3. [Multigenerational distillation, Fig. 4, Eqs. (18)-(19)] No multigenerational distillation data are presented. The only empirical test in SM §III concerns the single-generation trade-off relation (7); it does not test the flow (16), the constancy of a, or any of the scaling laws in Eqs. (18)-(19). Because Eq. (15) and A0_l≡A0 are imposed rather than derived from a distillation loss, Fig. 4 is currently a prediction of an unvalidated map. The paper should either provide multigenerational experiments or a calibration of a from an actual student-teacher training sequence, or explicitly label the phase diagram and scaling laws as a model-based conjecture rather than an empirically supported result.
minor comments (4)
  1. [Figure citations in main text] The sentence 'As shown in Fig. 2' in the section on the binary model refers to the threshold-sweep data, but the relevant figure in the main text is Fig. 1; please harmonize the figure numbering.
  2. [Equation numbering in main text] In the paragraphs following Eq. (16), the text refers to 'Eqs. (19)' and 'Eqs. (18)' when discussing the fixed-point configurations; correct these equation cross-references to the intended equations.
  3. [SM Section II A cross-reference] The SM text refers to 'Fig. 3 of the main text' for the pitchfork and saddle-node bifurcation lines, but the relevant main-text figure is the phase diagram in Fig. 4; please correct the cross-reference.
  4. [SM Section I B and Tables S2-S3] The SM contains a typo 'the relative HDC loss HDC al originates' and the main-text abstract says 'an renormalization-group like'; also, Tables S2-S3 state 'No averaging over the digitized vertices is performed,' while Section III B uses the averaged HDC K̄, so clarify that the tables list individual K_i values from which K̄ is later computed.

Circularity Check

1 steps flagged · score 4.0 of 10

Empirical trade-off check is partly self-definitional (K extracted with the tested equation), while the RG phase diagram is a self-contained consequence of stated assumptions.

  1. fitted input called prediction [Main text, 'Binary model for the reliability-safety trade-off' and Fig. 1; Supplemental Material Sec. III B, Eqs. (S69)-(S72), Tables S2-S3]
    "The results show that the values of (R, S) corresponding to different thresholds approximately satisfy Eq. (7), with the extracted HDC values clustering around ¯K = 3.471. ... 'Since the threshold T of the refusal token adjusts only the overall response tendency of the model ... we predict that varying T does not change the HDC parameter K. To verify this theoretical prediction, we digitized representative data points from Fig. 2 of Ref. [43] ... and calculated the corresponding K values at different thresholds T according to Eq. (S70).'"

    Eq. (S70) is exactly the relation being advertised as a prediction: logit(R)+logit(S)=K. Each threshold point's K_i is computed from that point's own (R_i,S_i) using Eq. (S70), so every data point automatically 'satisfies' Eq. (7) for its own K_i. Averaging those per-point K values and drawing the curve with K̄ does not convert the relation into an independent test; it only checks whether K_i is approximately constant across thresholds. Therefore the asserted consistency with Eq. (7) is true by construction and the empirical support reduces to a reparametrization of the data rather than a prediction of the trade-off curve. The RG phase diagram itself is not based on this step.

full rationale

The main RG derivation is not circular: Eq. (16) follows from explicit model assumptions (two-state logistic responses, teacher-student alignment V=-γ y_S y_T, and the linear inheritance rule K0_{l+1}=(1-a)K_l with constant a and A0). The fixed-point regimes, pitchfork and saddle-node boundaries, and tricritical point are mathematical consequences of that assumed flow, and the stated robustness conditions in SM Sec. II C are explicit rather than imported by self-citation. No load-bearing self-citation or imported uniqueness theorem is used. The only substantial circularity is in the experimental validation of the trade-off relation: K is extracted from the same data using the equation being tested, so the claim that the data 'satisfy Eq. (7)' is a definitional reparametrization, not an independent confirmation. The non-tautological part of that check is the approximate constancy of K across refusal thresholds, which is a genuine but weaker observation. Because this circularity attaches to a secondary consistency claim rather than to the central phase-diagram result, the score is moderate rather than severe.

Assumptions & free parameters 5 free parameters · 7 assumptions · 1 invented entities

The model contributes a reparametrization (R,S) -> K plus a postulated distillation dynamics. The main free parameters are the fitted HDC K and the unmeasured flow parameters a, γ, β, A0. The axioms are the two-state coarse-graining, the Boltzmann form, the alignment coupling, and the linear inheritance rule; the latter is the most load-bearing and is relaxed only under conditions stated in Section II.C of the supplement.

free parameters (5)
  • K (hazard discrimination capability) = 3.471 (no contrast), 4.471 (with contrast)
    Extracted from refusal-token data using Eq. (7) and averaged; the central trade-off curve is drawn from this fitted value.
  • a (per-generation relative HDC loss)
    Introduced in Eq. (15) to model student capacity loss; derived from a power-law spectrum model in the supplement but not measured, effectively free in the flow.
  • γ (teacher-student alignment strength)
    Coupling strength in Eq. (9); assumed weak (γ<<1), no value assigned; controls phase boundaries.
  • β (inverse temperature)
    Scales A and K; absorbed into βA0 and βA; no independent value; a modeling parameter.
  • A0 (pre-distillation refusal tendency)
    Assumed constant across generations in the flow; in the data comparison, the threshold T is mapped to changes in A.
assumptions (7)
  • standard math Response distribution is canonical, P(y,h|x) ∝ exp[-βE(y,h;x)] (Eq. 2).
    Assumes a thermal Boltzmann form with inverse temperature β; the entire reduced probability expression (Eq. 3) follows from this.
  • domain assumption Coarse-graining of inputs into ordinary/hazardous and outputs into answer/refusal is sufficient.
    Real LLM inputs and outputs are semantically rich; the two-state reduction is a modeling choice stated in the Introduction.
  • ad hoc to paper Teacher-student alignment takes the Ising-like form V = -γ y_S y_T (Eq. 9).
    Postulated to model distillation as a pairwise coupling; not derived from a specific training loss or algorithm.
  • ad hoc to paper Linear HDC inheritance K0_{l+1} = (1-a)K_l with constant a (Eq. 15).
    Central to the RG flow; the paper argues in Section II.C that smooth odd g(K) only renormalizes coefficients, but the proof is conditional.
  • ad hoc to paper Constant pre-distillation refusal tendency A0_l = A0 for all generations.
    Simplifies the flow to an autonomous system; the generalization to h(A) is discussed but not developed in the main text.
  • domain assumption Weak-alignment limit γ << 1 and continuum generation index.
    Used to obtain the approximate flow equations (16), the Ginzburg-Landau expansion (17), and the eigenvalue estimates in the supplement.
  • domain assumption Power-law spectrum u_q^T c_q / λ_q ~ q^{-p} in the supplemental microscopic model.
    Assumed to relate a to model-size reduction s; motivated by empirical spectra but not measured for the models in question.
invented entities (1)
  • Hazard discrimination capability K independent evidence
    purpose: Single scalar controlling the reliability-safety trade-off and its RG flow across distillation generations.
    Operationally defined as logit(R)+logit(S), so it is measurable from refusal/answer probabilities; however it is a reparametrization rather than an independently observable quantity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reliability-Safety Trade-off in AI Distillation: A Renormalization-Group Approach." pith.science (2026). https://pith.science/paper/SM32CBRL

@misc{pith2026260808572,
  author       = {Pith},
  title        = {Pith review of: Reliability-Safety Trade-off in AI Distillation: A Renormalization-Group Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SM32CBRL}},
  note         = {Machine review of arXiv:2608.08572}
}
read the original abstract

Knowledge distillation transfers more than task competence: it also transmits response propensities, refusal policies, error boundaries, and latent safety biases. We formulate this behavioral inheritance as a coarse-graining model grounded in statistical mechanics, in which the student's answer and refusal decisions define two macrostates, while the teacher induces an effective field that reshapes the student's free-energy landscape. The model yields a reliability-safety trade-off relation controlled by a single parameter K, which we term the hazard discrimination capability. The predicted trade-off is consistent with refusal-token data [arXiv: 2412.06748]. In knowledge distillation, a teacher with strong hazard discrimination improves the student's attainable reliability and safety, whereas poor discrimination limits the attainable trade-off. Repeated distillation acts as an iterated renormalization-group-like transformation, under which K follows a flow across generations. The flow exhibits a tricritical structure separating regimes of K loss, stable transmission, and threshold-dependent inheritance, and yields testable scaling predictions for multigenerational distillation.

Figures

Figures reproduced from arXiv: 2608.08572 by the authors.

Figure 1
Figure 1. FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 36 canonical work pages

  1. [43]

    N. Jain, A. Shrivastava, C. Zhu, D. Liu, A. Samuel, A. Panda, A. Kumar, M. Goldblum, and T. Goldstein, Refusal tokens: A simple way to calibrate refusals in large language models, arXiv preprint arXiv:2412.06748 (2025)

  2. [1]

    Bucila, R

    C. Bucila, R. Caruana, and A. Niculescu-Mizil, Model compression, Proc. ACM SIGKDD Int. Conf. Knowl. Dis- cov. Data Min. , 535 (2006)

  3. [2]

    Ba and R

    J. Ba and R. Caruana, Do deep nets really need to be 6 deep?, Adv. Neural Inf. Process. Syst. 27 (2014)

  4. [3]

    Hinton, O

    G. Hinton, O. Vinyals, and J. Dean, Distilling the knowl- edge in a neural network, arXiv:1503.02531 (2015)

  5. [4]

    J. Gou, B. Yu, S. J. Maybank, and D. Tao, Knowledge distillation: A survey, Int. J. Comput. Vis. 129, 1789 (2021)

  6. [5]

    A. K. Menon, A. S. Rawat, S. J. Reddi, S. Kim, and S. Kumar, A statistical perspective on distillation, Proc. Mach. Learn. Res. 139, 7632 (2021)

  7. [6]

    Stanton, P

    S. Stanton, P. Izmailov, P. Kirichenko, A. A. Alemi, and A. G. Wilson, Does knowledge distillation really work?, Adv. Neural Inf. Process. Syst. 34, 6906 (2021)

  8. [7]

    Kim and A

    Y. Kim and A. M. Rush, Sequence-level knowledge distil- lation, Proc. Conf. Empir. Methods Nat. Lang. Process. , 1317 (2016)

Show all 51 references
  1. [8]

    V. Sanh, L. Debut, J. Chaumond, and T. Wolf, Distil- bert, a distilled version of bert: Smaller, faster, cheaper and lighter, arXiv:1910.01108 (2019)

  2. [9]

    Jiao et al

    X. Jiao et al. , Tinybert: Distilling bert for natural language understanding, Findings Assoc. Comput. Lin- guist.: EMNLP , 4163 (2020)

  3. [10]

    Wang et al

    W. Wang et al. , Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transform- ers, Adv. Neural Inf. Process. Syst. 33, 5776 (2020)

  4. [11]

    Xu et al

    X. Xu et al. , A survey on knowledge distillation of large language models, arXiv:2402.13116 (2024)

  5. [12]

    J. Ko, S. Kim, T. Chen, and S.-Y. Yun, Distillm: To- wards streamlined distillation for large language models, arXiv:2402.03898 (2024)

  6. [13]

    Hsieh et al

    C.-Y. Hsieh et al. , Distilling step-by-step! outperform- ing larger language models with less training data and smaller model sizes, Findings Assoc. Comput. Linguist.: ACL , 8003 (2023)

  7. [14]

    L. H. Li et al. , Symbolic chain-of-thought distillation: Small models can also think step-by-step, Proc. Annu. Meet. Assoc. Comput. Linguist. , 2665 (2023)

  8. [15]

    Busbridge et al., Distillation scaling laws, Proc

    D. Busbridge et al., Distillation scaling laws, Proc. Mach. Learn. Res. 267, 5977 (2025)

  9. [16]

    Askell et al

    A. Askell et al. , A general language assistant as a labo- ratory for alignment, arXiv:2112.00861 (2021)

  10. [17]

    Bai et al

    Y. Bai et al. , Training a helpful and harmless assis- tant with reinforcement learning from human feedback, arXiv:2204.05862 (2022)

  11. [18]

    Ouyang et al

    L. Ouyang et al. , Training language models to follow in- structions with human feedback, Adv. Neural Inf. Pro- cess. Syst. 35, 27730 (2022)

  12. [19]

    Bai et al

    Y. Bai et al. , Constitutional ai: Harmlessness from ai feedback, arXiv:2212.08073 (2022)

  13. [20]

    Rafailov et al

    R. Rafailov et al. , Direct preference optimization: Your language model is secretly a reward model, Adv. Neural Inf. Process. Syst. 36 (2023)

  14. [21]

    Casper et al

    S. Casper et al. , Open problems and fundamental limi- tations of reinforcement learning from human feedback, Trans. Mach. Learn. Res. , bx24KpJ4Eb (2023)

  15. [22]

    I. J. Goodfellow, J. Shlens, and C. Szegedy, Explaining and harnessing adversarial examples, arXiv:1412.6572 (2014)

  16. [23]

    Amodei et al

    D. Amodei et al. , Concrete problems in ai safety, arXiv:1606.06565 (2016)

  17. [24]

    Hendrycks, N

    D. Hendrycks, N. Carlini, J. Schulman, and J. Stein- hardt, Unsolved problems in ml safety, arXiv:2109.13916 (2021)

  18. [25]

    Weidinger et al

    L. Weidinger et al. , Ethical and social risks of harm from language models, arXiv:2112.04359 (2021)

  19. [26]

    Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, New York, 2019)

    S. Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, New York, 2019)

  20. [27]

    Shumailov et al

    I. Shumailov et al. , Ai models collapse when trained on recursively generated data, Nature 631, 755 (2024)

  21. [28]

    Cloud et al

    A. Cloud et al. , Language models transmit behavioural traits through hidden signals in data, Nature 652, 615 (2026)

  22. [29]

    L. P. Kadanoff, Scaling laws for ising models near tc, Physics Phys. Fiz. 2, 263 (1966)

  23. [30]

    K. G. Wilson, The renormalization group: Critical phe- nomena and the kondo problem, Rev. Mod. Phys. 47, 773 (1975)

  24. [31]

    Tishby, F

    N. Tishby, F. C. Pereira, and W. Bialek, The information bottleneck method, Proc. Annu. Allerton Conf. Com- mun. Control Comput. , 368 (1999)

  25. [32]

    Tishby and N

    N. Tishby and N. Zaslavsky, Deep learning and the in- formation bottleneck principle, Proc. IEEE Inf. Theory Workshop , 1 (2015)

  26. [33]

    Koch-Janusz and Z

    M. Koch-Janusz and Z. Ringel, Mutual information, neu- ral networks and the renormalization group, Nat. Phys. 14, 578 (2018)

  27. [34]

    Gordon, A

    A. Gordon, A. Banerjee, M. Koch-Janusz, and Z. Ringel, Relevance in the renormalization group and in informa- tion theory, Phys. Rev. Lett. 126, 240601 (2021)

  28. [35]

    C. E. Shannon, A mathematical theory of communica- tion, Bell Syst. Tech. J. 27, 379 (1948)

  29. [36]

    E. T. Jaynes, Information theory and statistical mechan- ics, Phys. Rev. 106, 620 (1957)

  30. [37]

    Mezard and A

    M. Mezard and A. Montanari, Information, Physics, and Computation (Oxford University Press, Oxford, 2009)

  31. [38]

    Zdeborova and F

    L. Zdeborova and F. Krzakala, Statistical physics of in- ference: Thresholds and algorithms, Adv. Phys. 65, 453 (2016)

  32. [39]

    Y.-H. Ma, D. Xu, H. Dong, and C.-P. Sun, Optimal op- erating protocol to achieve efficiency at maximum power of heat engines, Phys. Rev. E 98, 022133 (2018)

  33. [40]

    Y.-H. Ma, D. Xu, H. Dong, and C.-P. Sun, Universal constraint for efficiency and power of a low-dissipation heat engine, Phys. Rev. E 98, 042112 (2018)

  34. [41]

    Zhai, F.-M

    R.-X. Zhai, F.-M. Cui, Y.-H. Ma, C. P. Sun, and H. Dong, Experimental test of power-efficiency trade-off in a finite- time carnot cycle, Phys. Rev. E 107, L042101 (2023)

  35. [42]

    R. X. Zhai, X. Yue, and C. P. Sun, Exact bound of power- efficiency trade-off in finite-time thermodynamic cycles, Phys. Rev. E 113, L012103 (2026)

  36. [44]

    S. I. Mirzadeh, M. Farajtabar, A. Li, N. Levine, A. Mat- sukawa, and H. Ghasemzadeh, Improved knowledge dis- tillation via teacher assistant, Proc. AAAI Conf. Artif. Intell. 34, 5191 (2020)

  37. [45]

    Zhang, Q

    C. Zhang, Q. Li, D. Song, Z. Ye, Y. Gao, and Y. Hu, Towards the law of capacity gap in distilling language models, Proc. Annu. Meet. Assoc. Comput. Linguist. , 22504 (2025)

  38. [46]

    K. K. Agrawal, A. K. Mondal, A. Ghosh, and B. A. Richards, α-ReQ: Assessing Representation Quality in Self-Supervised Learning by Measuring Eigenspectrum Decay, Adv. Neural Inf. Process. Syst. 35, 17626 (2022)

  39. [47]

    Papyan, X

    V. Papyan, X. Y. Han, and D. L. Donoho, Prevalence of neural collapse during the terminal phase of deep learning 7 training, Proc. Natl. Acad. Sci. U.S.A. 117, 24652 (2020)

  40. [48]

    Reliability–Safety Trade-off in AI Distillation: A Renormalization-Group Approach

    A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda, Refusal in language models is mediated by a single direction, Adv. Neural Inf. Process. Syst. 37, 136037 (2024). 8 Supplemental Material for “Reliability–Safety Trade-off in AI Distillation: A Re...

  41. [49]

    In this region, f (K) in Eq

    |βA0| < arcosh 2 We first discuss the region of relatively weak overall refusal tendency, jβA0j < arcosh 2. In this region, f (K) in Eq. ( S35) decreases monotonically for K > 0 from f (0) = sech 2(βA0/2) to f (1) = 0 . Therefore, a/γ and f (K) can have at most one intersectio...

  42. [50]

    In this region, the function f (K) in Eq

    |βA0| > arcosh 2 We now discuss the region of relatively strong overall response tendency, jβA0j > arcosh 2. In this region, the function f (K) in Eq. ( S35) first increases and then decreases for K > 0. Let the maximum of f (K) occur at Km > 0, satisfying f ′(Km) = 0 . This g...

  43. [51]

    answerable contrast queries

    |βA0| = arcosh 2 We now discuss the critical region jβA0j = arcosh 2 . In this case, f (K) in Eq. ( S35) decreases monotonically for K > 0 from f (0) = 2 /3 to f (1) = 0 , giving the same fixed-point structure as in Sec. II A 1. Hence, the parameter a/γ can be divided into the...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.