Pith. sign in

REVIEW 3 major objections 3 minor 99 references

From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis

T0 review · 3 major / 3 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read The paper claims that the neural scaling law of foundation models can be derived deductively from Zipf's law via Heaps' law and Hilberg's hypothesis, under explicit information-theoretic assumptions.

desk verdict A rigorous lower-bound chain from Zipf to neural scaling, with the abstract overselling a one-sided result; the non-stationary step survives the stress-test. read the letter →

arxiv 2512.13491 v3 pith:IENNYOUG submitted 2025-12-15 cs.IT cs.LGmath.ITmath.STstat.TH

classification cs.ITcs.LGmath.ITmath.STstat.TH MSC 94A1760G1091F2068T07
keywords neuralscalinglawZipf'sHeaps'Hilberg'shypothesisSantaFeprocessescrossentropyratepowerlaws
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to prove that the empirical power-law improvement of large language models is not a mystery of architecture or optimization but a consequence of ordinary token statistics. It builds a formal chain: Zipf's law implies Heaps' law on vocabulary growth; Heaps' law implies Hilberg's hypothesis on entropy growth; Hilberg's hypothesis implies the neural scaling law that ties cross entropy to training tokens, parameters, and compute. The key step is a differential form of Hilberg's law, stronger than the plain law studied empirically, plus entropy budgets that model limited data, parameters, and compute. A toy source called the Santa Fe process is shown to satisfy all four laws, illustrating that large memory and power-law complexity need not mean intuitively complex structure. If the chain holds, the scaling exponents of modern models are upper-bounded by the Hilberg exponent β, with γ_T ≤ 1−β and γ_N ≤ 1/β−1.

What carries the argument

The load-bearing object is the differential Hilberg law, Eq. (86): the conditional block entropy per symbol, sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h, is bounded below by C (t+s)^{β−1}. It is a strengthened, horizon-dependent version of the plain Hilberg law, and it is what makes the entropy-budget argument go through. Alongside it sit the entropy budgets (87)–(88), which model compute and parameter counts as caps on conditional and total Shannon entropy, and the Santa Fe process, a toy source in which each token is a pair (K_t, Z_{K_t}) of a Zipf-distributed index and a copied knowledge bit, which lets Heaps-type vocabulary growth be converted into Hilberg-type entropy growth.

What would settle it

On a large corpus, compute sup_{k≥t} H(X_{k+1}^{k+s}|X_1^t)/s − h for a range of t and s; if this conditional excess entropy decays faster than C (t+s)^{β−1} for any β close to the compression-based estimate 0.8, or vanishes over long horizons, Proposition 16's premise fails and the chain from Hilberg to neural scaling is not applicable to that corpus.

Watch

Extended reading notes

Core claim

The central claim is Proposition 16: if a stochastic text satisfies the differential Hilberg law — conditional excess entropy per symbol bounded below by a power law in the horizon — and if a trained model's entropy is bounded by compute and parameter budgets, then the worst-case expected cross entropy exceeds the entropy rate by at least the explicit expression in Eq. (90). This yields exponent bounds γ_T ≤ 1−β and γ_N ≤ 1/β−1. The paper further derives the differential Hilberg law from a differential Heaps law for stationary Santa Fe processes, and derives that Heaps law from approximate Zipf distributions under mixing conditions. The conclusion is that the observed neural scaling law can

Load-bearing premise

The load-bearing premise is the differential Hilberg law, Eq. (86) — a horizon-wise power-law lower bound on conditional excess entropy that strengthens the empirically studied plain Hilberg law and is assumed rather than derived from it; if real text satisfies only the plain law, the neural-scaling conclusion does not follow.

Editorial extensions

If this is right

  • If the differential Hilberg law holds with exponent β, the neural scaling exponents satisfy γ_T ≤ 1−β and γ_N ≤ 1/β−1, so token-scaling and parameter-scaling exponents are determined by the same language-level exponent.
  • Under the empirical compression-based estimate β≈0.8, the bounds give γ_T ≤ 0.2 and γ_N ≤ 0.25, which are loose compared with the commonly cited values γ_T≈0.095 and γ_N≈0.076; the paper suggests internet-scale corpora may have a larger β.
  • The derivation predicts underparameterization (γ_T < γ_N) as the optimal regime when parameters have bounded entropy, in contrast to the overparameterization reported in practice.
  • The Santa Fe process shows a Zipf-distributed IID narration over random knowledge bits satisfies Heaps' law and Hilberg's law, so Hilberg's hypothesis does not require intuitively complex structure.
  • Because the final implication (C) allows arbitrary non-stationary processes, the neural-scaling step is the most robust link in the chain once its differential premise is granted.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether real corpora satisfy the differential Hilberg law; measuring that conditional excess entropy directly on large datasets would test the crucial premise without relying on the plain-law proxy.
  • If the entropy-budget identification of parameter and compute counts is replaced by resource-bounded Kolmogorov complexity, the exponent bounds might tighten and possibly explain why overparameterized models appear better.
  • A testable extension: across languages or domains with different measured β, the framework predicts correspondingly different neural scaling exponents — a cross-linguistic prediction that current scaling-law studies do not yet examine.
  • The two-regime Zipf distributions and log-log convex vocabulary growth documented in quantitative linguistics would break the simple single-β picture; the framework suggests they should show up as piecewise or varying scaling exponents in model loss.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper attempts a systematic deductive chain from Zipf's law to Heaps' law, from Heaps' law to Hilberg's hypothesis, and from Hilberg's hypothesis to the neural scaling law, with Santa Fe processes as a running example. The main mathematical contributions are: (i) Propositions 13–14, deriving differential Heaps laws from Zipf-type tail conditions; (ii) Proposition 15, transferring differential Heaps laws to differential Hilberg laws for Santa Fe processes; and (iii) Proposition 16, deriving a lower bound on excess cross-entropy from a differential Hilberg law and entropy-budget constraints on the model, leading to the exponent inequalities γ_T ≤ 1−β and γ_N ≤ 1/β−1. The paper is written as a formal proof-based consolidation rather than an empirical study, and it candidly lists open problems about tightness and the meaning of compute.

Significance. If the stated assumptions are granted, the proofs provide a useful formal baseline connecting four well-known empirical regularities. The paper is honest about the main limitations: the derived neural-scaling statement is one-sided, and the differential Hilberg law is stronger than the empirically studied plain Hilberg law. Strengths include the explicit assumption-by-assumption organization, the absence of curve-fitting in the derivation, the Santa Fe process as a concrete non-vacuous example, and the clear identification of what would be needed to close the gaps. The work is best viewed as a formal lower-bound theory for scaling exponents rather than a derivation of the equality-form neural scaling law. Its significance is therefore conditional, but it is a useful contribution to the theoretical literature on scaling laws.

major comments (3)
  1. [§3.3, Eq. (90); Abstract] Proposition 16 proves only a lower bound on the excess expected cross-entropy. It does not prove the equality-form neural scaling law in Eqs. (8)–(10), and the resulting inequalities γ_T ≤ 1−β and γ_N ≤ 1/β−1 are one-sided. The abstract and conclusion nevertheless state that 'the neural scaling law is a consequence' and that the constraints 'produce the neural scaling law.' The open problems in §4 concede that tightness is unresolved. The paper should be recast as deriving testable lower bounds and upper bounds on exponents under explicit assumptions, not as a derivation of the empirical scaling law itself.
  2. [§3.3, Eq. (86); §1, Eq. (15)] Assumption (86) is a differential Hilberg law, which is substantially stronger than the empirically studied Hilberg hypothesis (6): it requires a uniform bound on conditional block entropies for all future starting points, not just the unconditional block entropy H(X_1^t). It is not derived from (6), and Proposition 15 derives it only for stationary Santa Fe processes. Thus the chain advertised in the abstract—Hilberg's hypothesis ⇒ neural scaling—does not follow for the plain Hilberg law. The paper needs either a derivation of (86) from (6) for a relevant class of processes or an explicit statement that the neural-scaling result depends on an unverified strengthening.
  3. [§3.3, Eqs. (90)–(95)] The expression in (90) contains the term ((1−y)/(1+y))^{1−β} with y = (c t^{−β}/(1−β))^{1/2}. If y ≥ 1, the base 1−y is non-positive and the non-integer power is not real; moreover, the definition of s_max in (89) can become negative. The proposition as stated is therefore not a valid real inequality for all c, t, n. The statement should explicitly restrict to y < 1 (e.g., c < (1−β)t^β), and state that the bound is vacuous otherwise. This does not affect the asymptotic regime of fixed c and t→∞, but the theorem statement needs a domain condition.
minor comments (3)
  1. [§3.2, Eq. (82)] In the proof of Proposition 15(1), the step sup_{k≥t} H(K_{k+1}^{k+s}|K_1^t)/s ≥ h might appear to assume stationarity. It is actually a consequence of Proposition 6's inf-characterization; adding a pointer would remove ambiguity.
  2. [§3.3, Eq. (14)] Equation (14) is displayed before the variables f, y, and the entropy-budget assumptions are introduced. Consider moving the display after Proposition 16 or defining the terms inline.
  3. [General] There are occasional typos, e.g., 'equvalent' in the Conclusion. More importantly, the abstract and introduction should consistently say 'lower bound on excess cross-entropy' rather than 'neural scaling law' in places where only the lower-bound result is meant.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the derivation is a conditional theorem chain from explicit assumptions, with no fitted input renamed as a prediction.

full rationale

The paper's central claim (Prop. 16, Eq. 90) is a conditional lower bound: given the differential Hilberg law (86) on the data's conditional block entropy and the entropy budgets (87)-(88) for the model Q, the model's worst-case excess cross entropy is bounded below by the displayed expression. This is not equivalent to its input by construction: (86) concerns the data's conditional entropy, whereas (90) concerns a model's cross entropy, and the proof bridges the two using the source coding inequality (31) and information-theoretic triangle inequalities (101)-(102). No parameter is fitted to data and no empirical value is inserted; the exponents gamma_T and gamma_N are read off as bounds implied by the derived lower bound, not as fitted outputs. The Heaps-to-Hilberg step (Prop. 15) and the Zipf-to-Heaps steps (Props. 13-14) are likewise explicit implications from stated conditions on the narration's cardinality rate, stationarity, mixing, and marginal tails. The self-citations ([26,27,28,30]) are used for the Santa Fe process as an illustrative example and for prior formulations, but the load-bearing inequalities (e.g., Eq. 80) are derived in the present paper rather than imported. No uniqueness theorem is invoked to forbid alternatives, and no ansatz is smuggled in through a citation. The paper's own acknowledged limitations, such as the coarse information-theoretic modeling of compute and parameter counts (12)-(13), concern empirical adequacy rather than circularity.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The central claim rests on an assumed exponent β and strengthened differential laws; no parameters are fitted to data in this paper. The biggest unpaid inputs are the differential Hilberg law (86) and the entropy-budget model of compute; the Santa Fe process is borrowed from prior work, not newly introduced.

free parameters (2)
  • β (Heaps/Hilberg exponent) = not fitted in this paper; empirical estimate ≈0.8 from PPM corpora (Takahira et al.)
    Central exponent in assumptions (17), (19), and (86); all derived bounds are functions of β. The paper treats β as an input from prior empirical work.
  • Multiplicative constants C0–C9 = unspecified
    Existential constants in conditions (66), (70), (71), (86)–(88) and in the final bounds (67), (72), (77), (78), (90). They affect prefactors but not the power-law exponents.
assumptions (6)
  • standard math Standard information-theoretic and analytic background: Shannon entropy inequalities, Fekete lemma, Hausdorff moment theorem, Gamma-function tail bounds.
    Used throughout Props 1–14; cited to Cover-Thomas, Fekete, Hausdorff, Karlin, and the author's [28].
  • domain assumption Approximate Zipf law p_k ≍ k^{-1/β}, or tail conditions (66)/(70)/(17).
    Input to implication A; not derived in the paper.
  • domain assumption Conditional repeat-probability bound p_k(t)/p_k ≤ C3 (condition 18/71), i.e., sufficiently strong mixing or finite-state/IID.
    Needed for the lower-bound differential Heaps law in Prop 14; related to ψ* mixing in Bradley's survey.
  • domain assumption Santa Fe decomposition X_t = (K_t, Z_{K_t}) with independent knowledge bits H(Z_k)∈[C7,C8], and stationarity of the narration for implication B.
    Prop 15; whether natural text satisfies this decomposition or an equivalent is postponed to future work.
  • ad hoc to paper Differential Heaps law (16) and differential Hilberg law (15)/(86).
    Stronger than the empirical plain laws (5)–(6); the paper introduces them as a practical necessity and does not derive (86) from (6).
  • ad hoc to paper Entropy resource constraints H(Q_tnc|X_1^t) ≤ C9 c and H(Q_tnc) ≤ C9 n.
    Used to model compute and parameter counts in Prop 16; the author concedes this is a coarse information-theoretic abstraction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis." pith.science (2026). https://pith.science/paper/IENNYOUG

@misc{pith2026251213491,
  author       = {Pith},
  title        = {Pith review of: From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IENNYOUG}},
  note         = {Machine review of arXiv:2512.13491}
}
read the original abstract

We inspect the deductive connection between the neural scaling law and Zipf's law -- two statements discussed in machine learning and quantitative linguistics. The neural scaling law describes how the cross entropy rate of a foundation model -- such as a large language model -- changes with respect to the amount of training tokens, parameters, and compute. By contrast, Zipf's law posits that the distribution of tokens exhibits a power law tail. Whereas similar claims have been made in more specific settings, we show that the neural scaling law is a consequence of Zipf's law under certain broad assumptions that we reveal systematically. The derivation steps are as follows: We derive Heaps' law on the vocabulary growth from Zipf's law, Hilberg's hypothesis on the entropy scaling from Heaps' law, and the neural scaling from Hilberg's hypothesis. We illustrate these inference steps by a toy example of the Santa Fe process that satisfies all four statistical laws.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

99 extracted references · 17 linked inside Pith

  1. [1]

    Achille and S

    A. Achille and S. Soatto. AI agents as universal task solvers, 2025. https://arxiv.org/abs/2510.12066

  2. [2]

    N. I. Akhiezer.The Classical Moment Problem and Some Related Ques- tions in Analysis. Society for Industrial and Applied Mathematics, 2021

  3. [3]

    D. J. Aldous. Exchangeability and related topics. In ´Ecole d’ ´Et´ e de Probabilit´ es de Saint-Flour XIII — 1983, volume 1117 ofLecture Notes in Mathematics, pages 1–198. Springer, 1985

  4. [4]

    E. G. Altmann, J. B. Pierrehumbert, and A. E. Motter. Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words.PLoS ONE, 4:e7678, 2009

  5. [5]

    R. H. Baayen.Word frequency distributions. Kluwer Academic Publish- ers, 2001

  6. [6]

    Bahri, E

    Y. Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma. Explaining neural scaling laws.http://arxiv.org/abs/2102.06701, 2021

  7. [7]

    P. L. Bartlett, P. M. Long, G. Lugosi, and A. Tsigler. Benign overfitting in linear regression.Proc. Nat. Acad. Sci., 117(48):30063–30070, 2020

  8. [8]

    M. Belkin. Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation.https://arxiv.org/ abs/2105.14368, 2021

Show all 99 references
  1. [9]

    Belkin, D

    M. Belkin, D. Hsu, S. Ma, and S. Mandal. Reconciling modern machine- 25 learning practice and the classical bias-variance trade-off.Proc. Nat. Acad. Sci., 116(32):15849–15854, 2019

  2. [10]

    Bengio, Y

    Y. Bengio, Y. LeCun, and G. E. Hinton. Deep learning for AI.Comm. ACM, 64(7):58–65, 2021

  3. [11]

    S. N. Bernstein. Sur les fonctions absolument monotones.Acta Math., 52:1–66, 1928

  4. [12]

    Bialek, I

    W. Bialek, I. Nemenman, and N. Tishby. Complexity through nonex- tensivity.Physica A, 302:89–99, 2001

  5. [13]

    R. C. Bradley. Basic properties of strong mixing conditions. a survey and some open questions.Probab. Surveys, 2:107–144, 2005

  6. [14]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, G. K. Ariel Herbert-Voss, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, M. L. Eric Sigler, S. Gray, B. Chess,...

  7. [15]

    Cabannes, E

    V. Cabannes, E. Dohmatob, and A. Bietti. Scaling laws for associative memories, 2024.https://arxiv.org/abs/2310.02984

  8. [16]

    Chacoma and D

    A. Chacoma and D. H. Zanette. Heaps’ law and Heaps functions in tagged texts: Evidences of their linguistic relevance, 2020.https: //arxiv.org/abs/2001.02178

  9. [17]

    Chowdhery, S

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Brad- bury, J. ...

  10. [18]

    J. G. Cleary and I. H. Witten. Data compression using adaptive coding and partial string matching.IEEE Trans. Comm., 32:396–402, 1984

  11. [19]

    E. U. Condon. Statistics of vocabulary.Science, 67(1733):300–300, 1928

  12. [20]

    T. M. Cover and J. A. Thomas.Elements of Information Theory, 2nd ed.Wiley & Sons, 2006

  13. [21]

    J. P. Crutchfield and D. P. Feldman. Regularities unseen, randomness 26 observed: The entropy convergence hierarchy.Chaos, 15:25–54, 2003

  14. [22]

    V. Davis. Types, tokens, and hapaxes: A new heap’s law.Glottotheory, 9(2):113–129, 2018

  15. [23]

    Devlin, M

    J. Devlin, M. Chang, K. Lee, and K. Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, and T. Solorio, editors,Proceedings of the 2019 Conference of the North American Chapter of the Association for Com- putat...

  16. [24]

    Dingemanse, D

    M. Dingemanse, D. E. Blasi, G. Lupyan, M. H. Christiansen, and M. P. Arbitrariness, iconicity, and systematicity in language.Trends Cogn. Sci., 19(10):603–615, 2015

  17. [25]

    Dębowski

    Ł. Dębowski. On Hilberg’s law and its links with Guiraud’s law.J. Quantit. Linguist., 13:81–109, 2006

  18. [26]

    Dębowski

    Ł. Dębowski. A general definition of conditional information and its application to ergodic decomposition.Statist. Probab. Lett., 79:1260– 1268, 2009

  19. [27]

    Dębowski

    Ł. Dębowski. On the vocabulary of grammar-based codes and the logical consistency of texts.IEEE Trans. Inform. Theory, 57:4589–4599, 2011

  20. [28]

    Dębowski.Information Theory Meets Power Laws: Stochastic Pro- cesses and Language Models

    Ł. Dębowski.Information Theory Meets Power Laws: Stochastic Pro- cesses and Language Models. Wiley & Sons, 2021

  21. [29]

    Dębowski

    Ł. Dębowski. A refutation of finite-state language models through Zipf’s law for factual knowledge.Entropy, 23:1148, 2021

  22. [30]

    Dębowski

    Ł. Dębowski. A simplistic model of neural scaling laws: Multiperiodic Santa Fe processes.https://arxiv.org/abs/2302.09049v1, 2023

  23. [31]

    Dębowski

    Ł. Dębowski. Corrections of Zipf’s and Heaps’ laws derived from hapax rate models.J. Quantit. Linguist., 32(2):128–165, 2025

  24. [32]

    Ebeling and G

    W. Ebeling and G. Nicolis. Word frequency and entropy of symbolic sequences: a dynamical perspective.Chaos Sol. Fract., 2:635–650, 1992

  25. [33]

    J. B. Estoup.Gammes st´ enographiques. Paris: Institut Stenographique de France, 1916

  26. [34]

    F. Fan. An asymptotic model for the English hapax/vocabulary ratio. Comput. Linguist., 36(4):631–637, 2010

  27. [35]

    M. Fekete. ¨Uber die Verteilung der Wurzeln bei gewissen algebraischen Gleichungen mit ganzzahligen Koeffizienten.Math. Z., 17:228–249, 1923

  28. [36]

    Ferrer-i-Cancho and R

    R. Ferrer-i-Cancho and R. V. Sol´ e. Two regimes in the frequency of words and the origins of complex lexicons: Zipf’s law revisited.J. Quan- tit. Linguist., 8(3):165–173, 2001

  29. [37]

    Font-Clos and A

    F. Font-Clos and A. Corral. Log-log convexity of type-token growth in 27 Zipf’s systems.Phys. Rev. Lett., 114:238701, 2015

  30. [38]

    Futrell and M

    R. Futrell and M. Hahn. Linguistic structure from a bottleneck on sequential information processing.Nature Hum. Behav., 2025.https: //doi.org/10.1038/s41562-025-02336-w

  31. [39]

    Futrell and K

    R. Futrell and K. Mahowald. How linguistics learned to stop worrying and love the language models.https://arxiv.org/abs/2501.17047, 2025

  32. [40]

    G´ acs and J

    P. G´ acs and J. K¨ orner. Common information is far less than mutual information.Probl. Contr. Inform. Theory, 2:119–162, 1973

  33. [41]

    Gerlach and E

    M. Gerlach and E. G. Altmann. Stochastic model for the vocabulary growth in natural languages.Phys. Rev. X, 3:021006, 2013

  34. [42]

    Guiraud.Les caract` eres statistiques du vocabulaire

    P. Guiraud.Les caract` eres statistiques du vocabulaire. Paris: Presses Universitaires de France, 1954

  35. [43]

    Harremo¨ es and F

    P. Harremo¨ es and F. Topsøe. Zipf’s law, hyperbolic distributions and entropy loss.Electr. Notes Disc. Math., 21:315–318, 2005. General Theory of Information Transfer and Combinatorics

  36. [44]

    Hausdorff

    F. Hausdorff. Momentprobleme f¨ ur ein endliches Intervall.Math. Z., 16: 220–248, 1923

  37. [45]

    H. S. Heaps.Information Retrieval—Computational and Theoretical Aspects. Academic Press, 1978

  38. [46]

    Henighan, J

    T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al. Scaling laws for autoregressive generative modeling.https://arxiv.org/abs/2010.14701, 2020

  39. [47]

    Herdan.Quantitative Linguistics

    G. Herdan.Quantitative Linguistics. Butterworths, 1964

  40. [48]

    Hernandez, J

    D. Hernandez, J. Kaplan, T. Henighan, and S. McCandlish. Scaling laws for transfer.https://arxiv.org/abs/2102.01293, 2021

  41. [49]

    W. Hilberg. Der bekannte Grenzwert der redundanzfreien Information in Texten — eine Fehlinterpretation der Shannonschen Experimente? Frequenz, 44:243–248, 1990

  42. [50]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifr...

  43. [51]

    G. Hu. On the amount of information.Teor. Verojat. Primenen., 4: 439–447, 1962

  44. [52]

    M. Hutter. Learning curve theory.https://arxiv.org/abs/2102.040 74, 2021

  45. [53]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws 28 for neural language models.https://arxiv.org/abs/2001.08361, 2020

  46. [54]

    S. Karlin. Central limit theorems for certain infinite urn schemes.J. Math. Mech., 17(4):373–401, 1967

  47. [55]

    Khmaladze

    E. Khmaladze. The statistical analysis of large number of rare events. Technical Report MS-R8804. Centrum voor Wiskunde en Informatica, Amsterdam, 1988

  48. [56]

    Kobayashi and K

    T. Kobayashi and K. Tanaka-Ishii. Taylor’s law for human linguistic sequences. In I. Gurevych and Y. Miyao, editors,Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1138–1148, Melbourne, Australia, 2018. Ass...

  49. [57]

    Kuraszkiewicz and J

    W. Kuraszkiewicz and J. Łukaszewicz. The number of different words as a function of text length.Pamiętnik Literacki, 42(1):168–182, 1951. In Polish

  50. [58]

    U. Lai, G. S. Randhawa, and P. Sheridan. Heaps’ law in GPT-Neo large language model emulated corpora, 2023.https://arxiv.org/abs/23 11.06377

  51. [59]

    L. A. Levin. Universal sequential search problems.Probl. Inform. Transm., 9(3):265–266, 1973

  52. [60]

    L. A. Levin. Randomness conservation inequalities: Information and in- dependence in mathematical theories.Inform. Control, 61:15–37, 1984

  53. [61]

    L´ evy.Processus stochastiques et mouvement brownien

    P. L´ evy.Processus stochastiques et mouvement brownien. Paris: Gau- thier Villars, 1948

  54. [62]

    H. Li, W. Zheng, Q. Wang, Z. Ding, H. Wang, Z. Wang, S. Xuyang, N. Ding, S. Zhou, X. Zhang, and D. Jiang. Farseer: A refined scaling law in large language models.https://arxiv.org/abs/2506.10972, 2025

  55. [63]

    W. Li. References on Zipf’s law.https://wli-zipf.upc.edu/, 2021

  56. [64]

    Louart, Z

    C. Louart, Z. Liao, and R. Couillet. A random matrix approach to neural networks.Ann. Appl. Probab., 28(2):1190–1248, 2018

  57. [65]

    Maloney, D

    A. Maloney, D. A. Roberts, and J. Sully. A solvable model of neural scaling laws.https://arxiv.org/abs/2210.16859, 2022

  58. [66]

    Mandelbrot

    B. Mandelbrot. Structure formelle des textes et communication.Word, 10:1–27, 1954

  59. [67]

    Mehri and M

    A. Mehri and M. Jamaati. Variation of Zipf’s exponent in one hundred live languages: A study of the Holy Bible translations.Phys. Lett. A, 381(31):2470–2477, 2017

  60. [68]

    E. J. Michaud, Z. Liu, U. Girit, and M. Tegmark. The quantization model of neural scaling.https://arxiv.org/abs/2303.13506, 2023

  61. [69]

    Mikolov, I

    T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Dis- 29 tributed representations of words and phrases and their compositional- ity. In C. J. C. Burges, L. Bottou, Z. Ghahramani, and K. Q. Weinberger, editors,Advances in Neural Information Processing Systems 26: ...

  62. [70]

    Miliˇ cka

    J. Miliˇ cka. Type-token & hapax-token relation: A combinatorial model. Glottotheory, 2(1):99–110, 2009

  63. [71]

    Miliˇ cka

    J. Miliˇ cka. Rank-frequency relation & type-token relation: Two sides of the same coin. In I. Obradović, E. Kelih, and R. K¨ ohler, editors,Meth- ods and Applications of Quantitative Linguistics—Selected papers of the 8th International Conference on Quantitative Linguistics (...

  64. [72]

    G. A. Miller. Some effects of intermittent silence.Amer. J. Psych., 70: 311–314, 1957

  65. [73]

    M. A. Montemurro and D. H. Zanette. New perspectives on Zipf’s law in linguistics: from single texts to large corpora.Glottometrics, 4:87–99, 2002

  66. [74]

    Neumann and C

    O. Neumann and C. Gros. Alphazero neural scaling and zipf’s law: a tale of board games and power laws, 2025.https://arxiv.org/abs/ 2412.11979

  67. [75]

    Z. Pan, S. Wang, P. Liao, and J. Li. Understanding LLM behaviors via compression: Data generation, knowledge acquisition and scaling laws, 2025.https://arxiv.org/abs/2504.09597

  68. [76]

    Pearce and J

    T. Pearce and J. Song. Reconciling Kaplan and Chinchilla scaling laws, 2024.https://arxiv.org/abs/2406.12907

  69. [77]

    Petersen, J

    A. Petersen, J. Tenenbaum, S. Havlin, H. E. Stanley, and M. Perc. Languages cool as they expand: Allometric scaling and the decreasing need for new words.Sci. Rep., 2:943, 2012

  70. [78]

    Pitman and M

    J. Pitman and M. Yor. The two-parameter Poisson–Dirichlet distribu- tion derived from a stable subordinator.Ann. Probab., 25(2):855–900, 1997

  71. [79]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners.https://open ai.com/blog/better-language-models/, 2019

  72. [80]

    D. A. Roberts, S. Yaida, and B. Hanin.The Principles of Deep Learning Theory. Cambridge University Press, 2022

  73. [81]

    B. Ryabko. Twice-universal coding.Probl. Inform. Transm., 20(3): 173–177, 1984

  74. [82]

    B. Y. Ryabko. Prediction of random sequences and universal coding. Probl. Inform. Transm., 24(2):87–96, 1988. 30

  75. [83]

    R. L. Schilling, R. Song, and Z. Vondraˇ cek.Bernstein Functions — Theory and Applications. Walter de Gruyter, 2010

  76. [84]

    C. Shannon. A mathematical theory of communication.Bell Syst. Tech. J., 30:379–423,623–656, 1948

  77. [85]

    C. Shannon. Prediction and entropy of printed English.Bell Syst. Tech. J., 30:50–64, 1951

  78. [86]

    Sharma and J

    U. Sharma and J. Kaplan. Scaling laws from the data manifold dimen- sion.J. Machine Learn. Res., 23(9):1–34, 2022

  79. [87]

    H. A. Simon. On a class of skew distribution functions.Biometrika, 42: 425–440, 1955

  80. [88]

    Takahira, K

    R. Takahira, K. Tanaka-Ishii, and Ł. Dębowski. Entropy rate estimates for natural language—a new extrapolation of compressed large-scale cor- pora.Entropy, 18(10):364, 2016

  81. [89]

    Tanaka-Ishii.Statistical Universals of Language: Mathematical Chance vs

    K. Tanaka-Ishii.Statistical Universals of Language: Mathematical Chance vs. Human Choice. Springer, 2021

  82. [90]

    C. Tao, Q. Liu, L. Dou, N. Muennighoff, Z. Wan, P. Luo, M. Lin, and N. Wong. Scaling laws with vocabulary: Larger models deserve larger vocabularies, 2024.https://arxiv.org/abs/2407.13623

  83. [91]

    Thoppilan, D

    R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, V. Zhao, Y. Zhou, C.-...

  84. [92]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vish- wanathan, and R. Garnett, editors,Advances in Neural Information Pr...

  85. [93]

    A. Wei, W. Hu, and J. Steinhardt. More than a toy: Random matrix models predict how real-world neural representations generalize.http: //arxiv.org/abs/2203.06176, 2022

  86. [94]

    J. R. Williams, J. P. Bagrow, C. M. Danforth, and P. S. Dodds. Text 31 mixing shapes the anatomy of rank-frequency distributions.Phys. Rev. E, 91:052811, 2015

  87. [95]

    A. D. Wyner. The common information of two dependent random vari- ables.IEEE Trans. Inform. Theory, IT-21:163–179, 1975

  88. [96]

    R. W. Yeung.First Course in Information Theory. Kluwer Academic Publishers, 2002

  89. [97]

    Zhang, S

    C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals. Understanding deep learning (still) requires rethinking generalization.Comm. ACM, 64 (3):107–115, 2021

  90. [98]

    G. K. Zipf.The Psycho-Biology of Language: An Introduction to Dy- namic Philology. Houghton Mifflin, 1935

  91. [99]

    G. K. Zipf.Human Behavior and the Principle of Least Effort. Addison- Wesley, 1949. 32

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.