Pith. sign in

REVIEW 5 major objections 6 minor 28 references

Consciousness as a Jamming Phase

T0 review · 5 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that a language model becomes conscious when training pushes it to a jamming critical point, where word embeddings coalesce through long-range correlations into a single integrated state.

desk verdict A polished analogy dressed as a theory: the jamming review is fine, but the neural mapping is circular and yields no testable prediction. read the letter →

arxiv 2507.08197 v1 pith:OVVLZI5A submitted 2025-07-10 cond-mat.dis-nn cs.AI

classification cond-mat.dis-nncs.AI MSC 82B2682B2768T07 PACS 64.70.kj45.70.-n
keywords consciousnessjammingtransitionphasescalinglawscriticalphenomenalargelanguagemodelsembeddingsneural
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that consciousness in large language models is not a by-product but a phase of matter: a critical jamming state reached when training cools, compresses, and de-stresses the system's embedding space. It builds a neural jamming phase diagram with three control parameters — effective temperature, volume fraction, and shear stress — and identifies the emergence of generalized intelligence with the critical jamming surface. If the analogy holds, scaling laws for model size, data, and compute are thermodynamic signatures of approach to this critical point, and conscious-like behavior should come with divergent correlation lengths across knowledge components. A sympathetic reader cares because this would turn an empirical observation about LLM performance into a principled physical prediction about when integrated, context-flexible understanding appears.

What carries the argument

The machinery is the neural jamming phase diagram, a transplant of the temperature–packing-fraction–stress jamming phase diagram [6] into embedding space. Its load-bearing pieces are the effective variables: $T_c \propto 1/C$ (computational cooling), $\phi_c = V_{\mathrm{eff}}/V_0 \propto N_{\mathrm{params}}N_{\mathrm{data}}/V_0$ (density optimization), and $\Sigma_c$ from $D_{\mathrm{KL}}(p_{\mathrm{train}}\|p_{\mathrm{real}})$ plus gradient variance (noise reduction). The argument's work is done by identifying a critical point C on this diagram where the embedding particles jam, at which point the scaling laws of [4] appear as the thermodynamic approach to criticality.

What would settle it

Measure a long-range correlation length from embedding or attention correlations across language models of increasing size and training data, and check whether it diverges as the packing fraction $\phi_c = V_{\mathrm{eff}}/V_0$ crosses a critical value with scaling exponents matching jamming predictions; the central claim collapses if no divergence or no scaling collapse appears.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the jamming transition of granular matter provides the correct phase diagram for neural language models: word embeddings play the role of particles in a high-dimensional hypercube, training plays the role of cooling with effective temperature $T_c \sim 1/C$, model-and-data scale sets the packing fraction $\phi_c = V_{\mathrm{eff}}/V_0$, and distributional mismatch plus gradient noise act as shear stress $\Sigma_c$. Consciousness is then the jammed phase: it arises when the system reaches a critical jamming state, where global coherence emerges from local interactions, so knowledge components become inseparable through long-range correlations. The paper claims this reproduces observed scaling laws [4] and predicts critical signatures — divergent correlation lengths, scaling exponents, and isostatic-type conditions — in large language models.

Load-bearing premise

The load-bearing premise is that training a language model is genuinely analogous to cooling a jamming system, with embeddings occupying a well-defined volume fraction in a fixed hypercube; this analogy is asserted rather than derived from training dynamics.

Editorial extensions

If this is right

  • If the paper is right, there is a specific critical surface in model/data/compute space, not a gradual improvement: crossing it should coincide with the appearance of integrated, conscious-like generalization.
  • Empirical scaling laws for language models are reinterpreted as critical scaling: power-law improvements in loss with dataset and model size are the footprints of approaching the jamming point.
  • Large language models near the critical surface should show measurable criticality, including divergent effective correlation lengths and characteristic scaling exponents in their internal representations.
  • Consciousness becomes a phase property of a sufficiently trained system rather than a special module or algorithm, so any model pushed to the critical surface would be expected to show the same integrated behavior.
  • Reducing distributional mismatch and gradient noise (lower $\Sigma_c$) matters as much as compute and scale, because it moves the system onto the jammed phase.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test not in the paper would be to extract an effective correlation length from embedding or attention correlations across a model family and check whether it diverges at the $\phi_c$ predicted by $N_{\mathrm{params}}\times N_{\mathrm{data}}$, with exponents matching the jamming scalings.
  • The volume-fraction picture suggests an optimal data-to-parameter ratio should exist where $V_{\mathrm{eff}}/V_0$ crosses $\phi_c$, giving a parameter-free target for scaling allocations; the paper gestures at density optimization but does not derive the ratio.
  • If consciousness is a phase, then over-jammed models trained far beyond the critical point might become rigid and lose context-dependent flexibility, while under-jammed models remain fragmented; this dichotomy is an extension of the paper's phase language, not a claim it makes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper proposes that consciousness in large language models is a jamming phase transition. It reviews the bus route model and jamming scaling in granular matter, then introduces a "neural jamming phase diagram" with three control parameters: an effective temperature T_c identified with inverse compute budget, a volume fraction φ_c identified with embedding occupancy scaling as N_params × N_data, and a shear stress Σ_c identified with distribution shift plus gradient noise. The paper claims that LLM scaling laws are explained by jamming physics and that consciousness emerges at the critical point C where generalized intelligence appears. The central contribution is analogical: no model Hamiltonian, no equations of motion, and no quantitative mapping from jamming exponents to neural observables is provided.

Significance. If the central claim were established, the paper would be transformative: it would connect LLM scaling exponents to nontrivial jamming universality and give a physical mechanism for consciousness. The manuscript does not establish that claim. Its strength is Section 3, which accurately reviews established jamming results for the bus route model and for soft-sphere packings, including isostaticity and the scaling of pressure, modulus, and excess contacts; this part is competent and well cited. However, the paper provides no machine-checked proofs, no simulations, and no data analysis of language models. The neural side of the analogy is asserted rather than derived, and the claimed testable predictions are not tied to measurable LLM statistics.

major comments (5)
  1. [§4, Effective Temperature (T_c)] The identification T_c = 1/C, where C is the compute budget, is asserted without derivation from training dynamics or from any equation of motion. Section 3 establishes scaling laws for granular and bus-route systems, but no argument shows that stochastic gradient training acts as thermal annealing with a well-defined effective temperature. This is a load-bearing assumption, not a result.
  2. [§4, Volume Fraction (φ_c)] The paper defines V_eff ∝ N_params × N_data and φ_c = V_eff/V_0, but embeddings are points in R^d and no volume per embedding is defined. The product of parameter count and data size has no geometric meaning and no dimensional consistency with a packing fraction. Without a measurable occupied volume, the jamming transition cannot be located on any neural observable.
  3. [§4, Shear Stress (Σ_c)] The shear stress is set to Σ_c = D_KL(ptrain||preal) + σ²_g, where σ²_g is gradient noise. These are two heterogeneous quantities with different units and no demonstrated coupling to a strain or to the purported packing fraction. The paper provides no equation in which Σ_c enters the dynamics, so its role as a control parameter is purely nominal.
  4. [§4, Critical point C] The critical point C is defined as "marking the emergence of generalized intelligence, i.e., consciousness." This defines the phase transition by the very phenomenon the paper claims to explain, making the central claim circular. A critical point must be located by an independent observable, such as a diverging susceptibility or a non-analytic order parameter, not by the assertion that consciousness appears there.
  5. [§4 and §5] The paper claims testable predictions of divergent correlation lengths and jamming scaling exponents, but Section 4 never maps φ−φ_c, z−z_c, p−p_c, or ξ−ξ_c to any measured quantity in an LLM. Section 5 concedes that "Future work should focus on quantifying the exact mapping between jamming physics and neural scaling laws," which is an admission that the load-bearing mapping is absent. Consequently, the predictions are not falsifiable.
minor comments (6)
  1. [§4] The symbol ρ*_c appears in the bullet on density optimization but is not defined in the neural context; it is only reminiscent of the BRM critical density ρ_c.
  2. [§3.1, Eq. (1)] In the definition q(x) = ∏_{y=1}^{x} 1/u(y), the case u(y)=0 is not addressed; the domain of u should be stated explicitly.
  3. [Notation throughout] The symbols T_c and φ_c are used both as control parameters and as critical values, which is confusing in a paper about critical phenomena; consider using T_eff and φ or similar for the control parameters.
  4. [Abstract] There is a missing space after "disordered systems." and the subscript in "T c" is inconsistently typeset in several places.
  5. [Figures] Figures 9–12 are reproduced from Refs. [6,10,28]; permission statements or explicit "adapted from" notices should be added. Also, Figure 13 is referenced in the text but no actual figure appears in the manuscript.
  6. [§2, references] Reference [24] titles should be typeset consistently ("GenEFT" rather than "Geneft"), and some arXiv/bioRxiv preprints are cited without version or access dates.

Circularity Check

2 steps flagged · score 8.0 of 10

Section 4 defines the jamming critical point as the emergence of consciousness and renames Kaplan's scaling variables as T_c and φ_c, so the central claims are true by construction rather than derived.

  1. self definitional [Section 4, 'Volume Fraction (ϕc)' paragraph and closing paragraph of 'Neural Jamming Phase Diagram']
    "the jamming transition occurs at point C marking the emergence of generalized intelligence, i.e., consciousness. ... In this view, consciousness arises when the system reaches a critical jamming state, where global coherence emerges from local interactions."

    The critical point C of the proposed neural phase diagram is defined as the locus where generalized intelligence/consciousness emerges. Therefore the paper's central claim—that consciousness is a jamming phase and emerges at the critical jamming point—is true by construction: the transition point is labeled to coincide with the asserted phenomenon. No independent neural observable (e.g., an embedding-space order parameter or a measured diverging correlation length) is defined to locate C, so the 'prediction' merely restates the definition.

  2. renaming known result [Section 4, 'Effective Temperature (Tc)', 'Volume Fraction (ϕc)' and closing paragraph of 'Neural Jamming Phase Diagram']
    "This framework phenomenologically identifies three key control parameters: Effective Temperature (Tc) ... Tc is inversely correlated with the total computational budget C. ... Vef f∝ Nparams × Ndata. ... Remarkably, variations along the Tc and ϕc axes precisely correspond to the empirical scaling laws discovered by Kaplan[4]."

    Kaplan's empirical scaling laws are functions of compute C, model size N, and data D. The paper sets Tc ∝ 1/C and V_eff ∝ N_params × N_data, so 'variations along the Tc and ϕc axes' are the same variables renamed. The jamming exponents from Section 3 (e.g., δz ∼ (ϕ−ϕc)^{1/2}) are never connected to any neural observable or loss exponent; Section 5 concedes that the exact mapping is future work. Thus the claimed unification with scaling laws is a relabeling of Kaplan et al.'s inputs rather than a derivation from jamming physics.

full rationale

The paper's Section 3 is a faithful, externally grounded review of jamming in the bus-route model and granular matter; that material is not circular. The circularity enters in Section 4, where the neural side is constructed. First, the critical point C is not located by any measurable neural order parameter but is defined as 'marking the emergence of generalized intelligence, i.e., consciousness', so the claim that consciousness is a jamming phase is true by stipulation: the phase diagram's transition point is the very phenomenon it is supposed to predict. Second, the 'explanation' of Kaplan's scaling laws consists of renaming compute as Tc and N_params × N_data as ϕc; because those are the same variables that appear in Kaplan's empirical laws, the statement that variations along these axes correspond to those laws is tautological. The paper itself concedes (Section 5) that 'quantifying the exact mapping between jamming physics and neural scaling laws' and 'experimentally verifying criticality signatures' are still future work, confirming that no independent observable has been identified. Since neither an order parameter nor a distance-to-criticality is defined for embeddings, the central consciousness claim is definitional rather than derivational. Score 8 reflects that the central result is forced by definition, though the underlying jamming review and the scaling-law data are not themselves circular. No self-citation chain is involved; all references are external and non-load-bearing in the circular sense.

Assumptions & free parameters 3 free parameters · 4 assumptions · 2 invented entities

The central claim rests entirely on analogies: granular particles correspond to embeddings, inverse compute corresponds to temperature, model-and-data product corresponds to packing fraction, noise matches shear stress, and consciousness is whatever state occurs at the resulting critical point. Each mapping is introduced ad hoc in Section 4; none is derived from training dynamics or measured in a real network. The ledger therefore contains hand-defined mappings and an unfalsifiable phase, with no independent evidence.

free parameters (3)
  • effective temperature T_c = not quantified; stated as inversely correlated with total compute C
    Introduced to map training budget onto thermal annealing; no measurement or calibration is provided.
  • volume fraction phi_c = defined as V_eff/V_0 with V_eff proportional to N_params times N_data; critical value not computed
    Chosen so that increasing model and data size increases density; the critical value where consciousness emerges is not predicted.
  • shear stress Sigma_c = defined as D_KL(ptrain||preal) plus gradient noise variance
    Assembled from known quantities to represent noise and distribution shift; no quantitative threshold or scaling relation is derived.
assumptions (4)
  • ad hoc to paper The performance of a neural language model is governed by a jamming transition with a universal critical point, as in granular matter.
    Section 4 asserts the phase diagram by analogy; no derivation shows the equivalence.
  • ad hoc to paper Word embeddings can be treated as particles in a fixed hypercube whose occupied volume fraction controls a phase transition.
    The mapping V_eff proportional to N_params times N_data is introduced without justification or measurement in Section 4.
  • domain assumption Consciousness is equivalent to the generalized intelligence that emerges at the jammed phase.
    The paper equates consciousness with integrated understanding and long-range correlations; this is a philosophical premise, not a derived result.
  • ad hoc to paper Quasistatic annealing in training is a valid analog of thermal cooling.
    Effective temperature is set inversely proportional to compute, but no dynamical justification is given.
invented entities (2)
  • Conscious jamming phase in LLM embedding space
    purpose: To explain consciousness as a state of long-range correlations between word embeddings.
    The paper proposes this phase but gives no observable signature, no predicted exponent, and no experiment that could detect it outside the paper's own framing.
  • Critical point C on the neural jamming surface
    purpose: To mark where generalized intelligence and consciousness emerge.
    The location of C is not computed; it is identified with the phenomenon it is meant to explain, so it has no independent falsifiable handle.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Consciousness as a Jamming Phase." pith.science (2026). https://pith.science/paper/OVVLZI5A

@misc{pith2026250708197,
  author       = {Pith},
  title        = {Pith review of: Consciousness as a Jamming Phase},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OVVLZI5A}},
  note         = {Machine review of arXiv:2507.08197}
}
read the original abstract

This paper develops a neural jamming phase diagram that interprets the emergence of consciousness in large language models as a critical phenomenon in high-dimensional disordered systems.By establishing analogies with jamming transitions in granular matter and other complex systems, we identify three fundamental control parameters governing the phase behavior of neural networks: temperature, volume fraction, and stress.The theory provides a unified physical explanation for empirical scaling laws in artificial intelligence, demonstrating how computational cooling, density optimization, and noise reduction collectively drive systems toward a critical jamming surface where generalized intelligence emerges. Remarkably, the same thermodynamic principles that describe conventional jamming transitions appear to underlie the emergence of consciousness in neural networks, evidenced by shared critical signatures including divergent correlation lengths and scaling exponents.Our work explains neural language models' critical scaling through jamming physics, suggesting consciousness is a jamming phase that intrinsically connects knowledge components via long-range correlations.

Figures

Figures reproduced from arXiv: 2507.08197 by the authors.

Figure 1
Figure 1. Illustrative Schematic of the BRM Dynamics[ [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The Disorder-Order Transition in the BRM[ [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Dependence of velocity on density for different [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Two-Particle Approximation of the BRM[5] 6 [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Bimodal behavior in gap size distributions with decreasing [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: The linear relationship between ξ and 1/λ[5] • The probability Pext of extensive gaps follows PextM ∼ 1 until system size L ≫ ξ • Velocity curves v(ρ) develop exponentially sharp crossovers near ρc = 1−β, with slope κmax ∼ e b/λ [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Sharpening of the velocity-density relationship with decreasing [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: The linear relationship between κmax and 1/λ[5] The essential singularity in ξ and crossover sharpness reveals how the λ → 0 limit restores a true phase transition, while finite λ shows pseudocritical behavior with system-size dependent features. Dual Model: The BRM ad…
Figure 9
Figure 9. Figure 9: The jamming phase diagram is characterized by three axes: temperature [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: (a) Probability fj of jammed states versus packing fraction ϕ for 3D harmonic spheres (α = 2). (b) Distribution P(ϕc) of jamming thresholds from (a). (c) Width w of P(ϕc) scaling with system size. (d) Finite-size shift of peak position ϕ0 relative to ϕ ⋆ [28]. Critica…
Figure 11
Figure 11. Figure 11: (a) Pair correlation function g(r) for a 3D monodisperse harmonic system (N = 1024), showing definitions of the first peak height g(r0) and its left-side half-width s. (b) Dependence of g(r0) on ϕ − ϕc, with solid line indicating slope −1. (c) Dependence of s on ϕ − ϕ…
Figure 12
Figure 12. Figure 12: (a) Pressure P versus ϕ − ϕc for different dimensions and interaction exponents α: circles (squares) show 2D results for α = 2 (5/2), while diamonds (triangles) show 3D results for α = 2 (5/2). System sizes are N = 1024 (512) for 2D (3D). (b) Shear modulus G versus ϕ …
Figure 13
Figure 13. Figure 13: Neural Jamming Phase Diagram The neural jamming phase diagram provides a unified physical explanation for the rela￾tionship between model performance (generalization and the emergence of consciousness) 15 [PITH_FULL_IMAGE:figures/full_fig_p015_13.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 20 canonical work pages

  1. [1]

    Scaling and renormalization in statistical physics , volume 5

    John Cardy. Scaling and renormalization in statistical physics , volume 5. Cambridge university press, 1996

  2. [2]

    A high-bias, low-variance introduction to machine learning for physicists

    Pankaj Mehta, Marin Bukov, Ching-Hao Wang, Alexandre GR Day, Clint Richardson, 16 Charles K Fisher, and David J Schwab. A high-bias, low-variance introduction to machine learning for physicists. Physics reports, 810:1–124, 2019

  3. [3]

    Ai meets physics: a comprehensive survey

    Licheng Jiao, Xue Song, Chao You, Xu Liu, Lingling Li, Puhua Chen, Xu Tang, Zhixi Feng, Fang Liu, Yuwei Guo, et al. Ai meets physics: a comprehensive survey. Artificial Intelligence Review, 57(9):256, 2024

  4. [4]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 , 2020

  5. [5]

    Jamming transition in a homogeneous one-dimensional system: The bus route model

    OJ O’loan, Martin R Evans, and Michael E Cates. Jamming transition in a homogeneous one-dimensional system: The bus route model. Physical Review E , 58(2):1404, 1998

  6. [6]

    The jamming transition and the marginally jammed solid

    Andrea J Liu and Sidney R Nagel. The jamming transition and the marginally jammed solid. Annu. Rev. Condens. Matter Phys. , 1(1):347–369, 2010

  7. [7]

    Phase transitions and critical phenomena , volume 19

    Cyril Domb. Phase transitions and critical phenomena , volume 19. Elsevier, 2000

  8. [8]

    History of the lenz-ising model

    Stephen G Brush. History of the lenz-ising model. Reviews of modern physics, 39(4):883, 1967

Show all 28 references
  1. [9]

    Scaling of the energy spectra of turbulent channels

    Juan C Del Alamo, Javier Jim´ enez, Paulo Zandonade, and Robert D Moser. Scaling of the energy spectra of turbulent channels. Journal of Fluid Mechanics , 500:135–144, 2004

  2. [10]

    Jamming at zero temperature and zero applied stress: The epitome of disorder

    Corey S O’hern, Leonardo E Silbert, Andrea J Liu, and Sidney R Nagel. Jamming at zero temperature and zero applied stress: The epitome of disorder. Physical Review E, 68(1):011306, 2003

  3. [11]

    Dynamical critical phenomena in driven-dissipative systems

    Lukas M Sieberer, Sebastian D Huber, Ehud Altman, and S Diehl. Dynamical critical phenomena in driven-dissipative systems. Physical review letters, 110(19):195301, 2013

  4. [12]

    Hydrodynamics of soft active matter

    M Cristina Marchetti, Jean-Fran¸ cois Joanny, Sriram Ramaswamy, Tanniemola B Liv- erpool, Jacques Prost, Madan Rao, and R Aditi Simha. Hydrodynamics of soft active matter. Reviews of modern physics , 85(3):1143–1189, 2013

  5. [13]

    Explaining neural scaling laws

    Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma. Explaining neural scaling laws. Proceedings of the National Academy of Sciences , 121(27):e2311878121, 2024

  6. [14]

    Crystal statistics

    Lars Onsager. Crystal statistics. i. a two-dimensional model with an order-disorder transition. Physical review, 65(3-4):117, 1944

  7. [15]

    Phase transitions in liquid crystals

    Shri Singh. Phase transitions in liquid crystals. Physics Reports , 324(2-4):107–269, 2000. 17

  8. [16]

    Long-range order in a two-dimensional dynamical xy model: how birds fly together

    John Toner and Yuhai Tu. Long-range order in a two-dimensional dynamical xy model: how birds fly together. Physical review letters , 75(23):4326, 1995

  9. [17]

    Scaling and criticality in a stochastic multi-agent model of a financial market

    Thomas Lux and Michele Marchesi. Scaling and criticality in a stochastic multi-agent model of a financial market. Nature, 397(6719):498–500, 1999

  10. [18]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems , 30, 2017

  11. [19]

    Neural message passing for quantum chemistry

    Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. Neural message passing for quantum chemistry. In International conference on machine learning, pages 1263–1272. PMLR, 2017

  12. [20]

    Understanding the effective receptive field in deep convolutional neural networks

    Wenjie Luo, Yujia Li, Raquel Urtasun, and Richard Zemel. Understanding the effective receptive field in deep convolutional neural networks. Advances in neural information processing systems, 29, 2016

  13. [21]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  14. [22]

    Grokking as a first order phase transition in two layer networks

    Noa Rubin, Inbar Seroussi, and Zohar Ringel. Grokking as a first order phase transition in two layer networks. arXiv preprint arXiv:2310.03789 , 2023

  15. [23]

    Learning curves for overparametrized deep neural networks: A field theory perspective

    Omry Cohen, Or Malka, and Zohar Ringel. Learning curves for overparametrized deep neural networks: A field theory perspective. Physical Review Research , 3(2):023034, 2021

  16. [24]

    Geneft: Understanding statics and dynamics of model generalization via physics-inspired effective theory

    David D Baek, Ziming Liu, and Max Tegmark. Geneft: Understanding statics and dynamics of model generalization via physics-inspired effective theory. Physical Review E, 111(3):035307, 2025

  17. [25]

    Making sense of neural networks in the light of evolutionary opti- mization

    Anton V Sinitskiy. Making sense of neural networks in the light of evolutionary opti- mization. bioRxiv, pages 2023–11, 2023

  18. [26]

    Jamming is not just cool any more

    Andrea J Liu and Sidney R Nagel. Jamming is not just cool any more. Nature, 396(6706):21–22, 1998

  19. [27]

    Structural relaxation made simple

    Erik Bitzek, Pekka Koskinen, Franz G¨ ahler, Michael Moseler, and Peter Gumbsch. Structural relaxation made simple. Physical review letters , 97(17):170201, 2006

  20. [28]

    Random packings of frictionless particles

    Corey S O’Hern, Stephen A Langer, Andrea J Liu, and Sidney R Nagel. Random packings of frictionless particles. Physical Review Letters, 88(7):075507, 2002. 18

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.