REVIEW 2 major objections 6 minor 50 references
Exact solutions show how concept directions align during training, not just after it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 06:18 UTC pith:YFIZE6ZB
load-bearing objection Exact abstraction trajectories under 2FS linear nets are the real contribution; the GELU-ablation application is a useful but under-proved leap from the infinite-width attenuation law. the 2 major comments →
How are linear representations learned? Exact solutions to the dynamics of abstraction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under a minimal linear network with two-factor symmetry and optimal readout, abstraction has an exact closed-form trajectory. Its terminal value is set only by the geometric mean of input and target inverse signal-to-noise ratios; depth and initialization separately control layerwise growth and peak overshoot. In infinite-width nonlinear nets the same geometry is reshaped by the nonlinearity, and both erf and leaky ReLU attenuate abstraction so that feature abstraction never exceeds preactivation abstraction.
What carries the argument
Abstraction score α: the cosine similarity of a concept direction measured in two contexts, rewritten as a function of the signal-to-noise ratio of two eigenmodes (S versus SC) of a five-entry two-factor-symmetric kernel. Exact scalar ODEs for those modes yield the terminal law, depth interpolation, and attenuation factor.
Load-bearing premise
The whole exact theory needs the data, targets, and initial features to obey a strong two-factor symmetry so that every kernel collapses to five numbers and the same five modes; without that symmetry the closed-form trajectory disappears.
What would settle it
Train a deep ReLU or transformer stack on data that clearly violates two-factor symmetry and check whether terminal abstraction still tracks the geometric-mean formula, whether deeper layers are more abstract, and whether ablating the local nonlinearity still raises feature abstraction and probe transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a dynamical theory of “abstraction” (cosine alignment of context-specific concept vectors) under the linear representation hypothesis. In a two-layer linear network with variable-projected readout and two-factor symmetry (2FS) kernels, the authors reduce gradient flow to scalar mode ODEs and obtain an exact implicit trajectory for abstraction (Theorem 3). From this they derive a terminal law α∞ determined by the geometric mean of input and target inverse-SNRs (Theorem 4), overshoot and initialization-scale control of peak abstraction (Theorem 5, Proposition 6), and depth laws for layerwise interpolation and terminal abstraction under balanced deep linear networks (Theorem 7). They extend to infinite-width two-layer erf and leaky-ReLU networks, prove an attenuation law |αK| ≤ |αQ| for features vs preactivations (Theorem 8 / C.2), and characterize how ReLU terminal abstraction depends more on input than target geometry. Empirical checks include ResNets on 3dshapes, local GELU ablation in DINOv3 and Gemma 4, and macaque V4 vs IT.
Significance. If the linear results hold under the stated axioms, they give a rare closed-form account of how linear concept geometry evolves during training—not only at convergence—linking data/target geometry, depth, and rich/lazy initialization to abstraction trajectories. That is a clear advance over asymptotic perfect-abstraction results. The attenuation law and ReLU vs erf phase diagrams offer a mechanistic explanation for known nonlinearity effects on abstract geometry. Strengths include exact solutions (Riccati reduction, eigenvalue ODEs, Theorems 3–7), careful infinite-width nonlinear maps and proofs (Appendix C), and falsifiable qualitative predictions tested on ResNets, open transformers, and neural data. The applied GELU-ablation result is modest but practically relevant for probing/steering if the effect is robust.
major comments (2)
- §5 and abstract claim evidence for the attenuation law and improved probe generalization via local GELU ablation (Fig. 5). Theorem 8 / Theorem C.2 only prove |αK| ≤ |αQ| for infinite-width Gaussian preactivations with exact 2FS kernels under erf or leaky ReLU (Eq. 15; §C.5). Residual streams in DINOv3/Gemma are finite-width, multi-layer, non-Gaussian, and not 2FS; GELU is not among the verified nonlinearities. The manuscript should explicitly separate the theorem’s hypotheses from the empirical intervention, avoid language that presents Fig. 5 as a consequence of Theorem 8, and add controls (e.g., ablating other MLP components, random linear maps, or non-concept baselines) so the ~1pp probe gain is not over-attributed to the attenuation mechanism.
- Assumption 2 (2FS; §2.2, §A.3) is load-bearing for every exact scalar ODE and closed-form law (Proposition 2, Theorems 3–7). It is imposed via group invariance rather than derived for realistic data. The ResNet/3dshapes checks (§3.3, Fig. 3) are qualitative and still use a 2×2 factor design. The paper should state more sharply which predictions are expected to survive without 2FS (e.g., qualitative depth/init trends) versus which are 2FS-specific (exact α∞ formula, arctanh layerwise interpolation), and ideally report a controlled symmetry-breaking experiment (e.g., unbalanced class sizes or non-orthogonal Fourier components) to bound sensitivity.
minor comments (6)
- Assump. 1 (variable-projected readout) is better motivated than a frozen readout, and §A.4/Fig. 7 help, but early-training discrepancies should be flagged when interpreting non-monotonic α(t) near initialization.
- Setting 2 (layerwise balancing) for depth results (§3.2, §B.8) is standard but strong; a short remark on how unbalanced init or SGD noise would perturb Eq. (13)–(14) would help readers.
- Proposition 1’s Gaussian score approximation is used to link α to probe transfer; state when the approximation fails (heavy-tailed residual streams, multi-token concepts).
- Fig. 1E / 2FS entry notation (ad, a2, a1s, a1c, a0) is dense; a small table mapping entries to modes S/SC would improve readability.
- Related work on Word2Vec analogies and LRH is good; a one-sentence contrast with Jiang et al. [4] on dynamics vs asymptotics in the main text (not only appendix) would help orientation.
- Gemma results: report confidence intervals or paired tests for the 17/18 probe improvements; French–Spanish decrease should be discussed briefly.
Circularity Check
No significant circularity: terminal, depth, and attenuation laws are derived from gradient-flow ODEs under stated assumptions, not forced by definition or fitted targets.
full rationale
The paper's load-bearing analytic claims (Theorems 3–5, 7–8; Propositions 2, 6, 10–11) are obtained by reducing gradient flow under Assumps. 1–2 (and Settings 1–3) to scalar mode ODEs, then solving or analyzing those ODEs. Terminal α∞ follows from late-time √t growth of λS and λSC (Theorem B.1 / Theorem 4); depth laws from balanced gain factorization in arctanh-SNR space (Theorem 7); attenuation from NNGP maps with nonnegative power-series coefficients (Theorem C.2). None of these equal their inputs by construction: α is defined as cosine similarity of concept vectors, then shown to obey dynamics whose fixed points and bounds are nontrivial functions of data geometry, depth, init scale, and nonlinearity. Variable projection and 2FS are explicit modeling assumptions that enable closed form, not renamings of the claimed trajectory. Empirical applications (GELU ablation, macaque hierarchy) are independent tests, not fitted-then-predicted quantities. Self-citations to linear-network dynamics literature supply techniques, not uniqueness theorems that force the abstraction laws. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (3)
- ridge γ / readout regularization
- initialization scale κ (or σ_w²)
- nonlinearity parameters β (erf), ω (leaky ReLU)
axioms (6)
- standard math Gradient flow on MSE with ridge on readout only; features Z = WX (linear) or infinite-width Gaussian preactivations.
- domain assumption Assumption 1: readout always at ridge optimum Wr*(Z) (variable projection / infinite readout learning-rate limit).
- domain assumption Assumption 2 (2FS): Σx, Σy, Q(0) invariant under (Sn)^4 ⋊ (Z2)^2, yielding five-entry kernels and modes I,S,C,SC,G.
- domain assumption Setting 2: layerwise balanced gains per mode so λ_m^(ℓ) = λ_m^(0) u_m^{2ℓ}.
- domain assumption Infinite-width NNGP: feature kernel is E[ϕ(zi)ϕ(zj)] for z ~ N(0,Q); 2FS preserved under erf/L-ReLU.
- ad hoc to paper Setting 1 / Setting 3 regimes for overshoot and ReLU terminal analysis (signal-dominant; centered signal-balanced).
invented entities (3)
-
Abstraction score α (cosine of context-specific concept vectors)
independent evidence
-
2FS eigenmodes and inverse-SNR ν = λ_SC/λ_S
no independent evidence
-
Effective target kernel M(Q) and nonlinear gain R_ϕ
no independent evidence
read the original abstract
In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of how they emerge $\textit{during}$ training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training - a process we call "abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.
Figures
Reference graph
Works this paper leans on
-
[1]
Linguistic Regularities in Continuous Space Word Representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. Linguistic Regularities in Continuous Space Word Representations. In Lucy Vanderwende, Hal Daumé III, and Katrin Kirchhoff, editors, Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 746–751, Atlanta, Georgia,...
2013
-
[2]
Emergent Linear Representations in World Models of Self-Supervised Sequence Models, September 2023
Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent Linear Representations in World Models of Self-Supervised Sequence Models, September 2023. arXiv preprint arXiv:2309.00941
Pith/arXiv arXiv 2023
-
[3]
The Linear Representation Hypothesis and the Geometry of Large Language Models, July 2024
Kiho Park, Yo Joong Choe, and Victor Veitch. The Linear Representation Hypothesis and the Geometry of Large Language Models, July 2024. arXiv preprint arXiv:2311.03658
Pith/arXiv arXiv 2024
-
[4]
On the Origins of Linear Representations in Large Language Models
Yibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam, and Victor Veitch. On the Origins of Linear Representations in Large Language Models. InProceedings of the 41st International Conference on Machine Learning, pages 21879–21911. PMLR, July 2024
2024
-
[5]
The Geometry of Categorical and Hier- archical Concepts in Large Language Models, February 2025
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch. The Geometry of Categorical and Hier- archical Concepts in Large Language Models, February 2025. arXiv preprint arXiv:2406.01506
Pith/arXiv arXiv 2025
-
[6]
Benna, Mattia Rigotti, Jérôme Munuera, Stefano Fusi, and C
Silvia Bernardi, Marcus K. Benna, Mattia Rigotti, Jérôme Munuera, Stefano Fusi, and C. Daniel Salzman. The Geometry of Abstraction in the Hippocampus and Prefrontal Cortex.Cell, 183(4):954–967.e21, November 2020
2020
-
[7]
Rodgers, Randy M
Ramon Nogueira, Chris C. Rodgers, Randy M. Bruno, and Stefano Fusi. The geometry of cortical representations of touch in rodents.Nature Neuroscience, 26(2):239–250, February 2023
2023
-
[8]
Shin, Wenbo Tang, and Shantanu P
Justin D. Shin, Wenbo Tang, and Shantanu P. Jadhav. Protocol for geometric transformation of cognitive maps for generalization across hippocampal-prefrontal circuits.STAR Protocols, 4(3):102513, September 2023
2023
-
[9]
Neural representational geometries reflect behavioral differences in monkeys and recurrent neural networks.Nature Communications, 15(1):6479, August 2024
Valeria Fascianelli, Aldo Battista, Fabio Stefanini, Satoshi Tsujimoto, Aldo Genovesio, and Stefano Fusi. Neural representational geometries reflect behavioral differences in monkeys and recurrent neural networks.Nature Communications, 15(1):6479, August 2024
2024
-
[10]
Courellis, Juri Minxha, Araceli R
Hristos S. Courellis, Juri Minxha, Araceli R. Cardenas, Daniel L. Kimmel, Chrystal M. Reed, Taufik A. Valiante, C. Daniel Salzman, Adam N. Mamelak, Stefano Fusi, and Ueli Rutishauser. Abstract representations emerge in human hippocampal neurons during inference.Nature, 632(8026):841–849, August 2024
2024
-
[11]
Boyle, Lorenzo Posani, Sarah Irfan, Steven A
Lara M. Boyle, Lorenzo Posani, Sarah Irfan, Steven A. Siegelbaum, and Stefano Fusi. Tuned geometries of hippocampal representations meet the computational demands of social memory. Neuron, 112(8):1358–1371.e9, April 2024
2024
-
[12]
Karyna Mishchanchuk, Gabrielle Gregoriou, Albert Qü, Alizée Kastler, Quentin J. M. Huys, Linda Wilbrecht, and Andrew F. MacAskill. Hidden state inference requires abstract contextual representations in the ventral hippocampus.Science, 386(6724):926–932, November 2024
2024
-
[13]
Schoonover, Andrew J
Pia-Kelsey O’Neill, Lorenzo Posani, Jozsef Meszaros, Phebe Warren, Carl E. Schoonover, Andrew J. P. Fink, Stefano Fusi, and C. Daniel Salzman. The representational geometry of emotional states in basolateral amygdala, April 2024. bioRxiv preprint 2023.09.23.558668
2024
-
[14]
Finding Neurons in a Haystack: Case Studies with Sparse Probing, June 2023
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas. Finding Neurons in a Haystack: Case Studies with Sparse Probing, June 2023. arXiv preprint arXiv:2305.01610
Pith/arXiv arXiv 2023
-
[15]
De- tecting Strategic Deception with Linear Probes
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, and Marius Hobbhahn. De- tecting Strategic Deception with Linear Probes. InForty-Second International Conference on Machine Learning, June 2025
2025
-
[16]
Vazquez, Ulisse Mini, and Monte MacDiarmid
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering Language Models With Activation Engineering, October
-
[17]
arXiv preprint arXiv:2308.10248
-
[18]
Inference- Time Intervention: Eliciting Truthful Answers from a Language Model, June 2024
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference- Time Intervention: Eliciting Truthful Answers from a Language Model, June 2024. arXiv preprint arXiv:2306.03341. 11
Pith/arXiv arXiv 2024
-
[19]
Takuya Ito, Tim Klinger, Douglas H. Schultz, John D. Murray, Michael W. Cole, and Mattia Rigotti. Compositional generalization through abstract representations in human and artificial neural networks, September 2022. arXiv preprint arXiv:2209.07431
Pith/arXiv arXiv 2022
-
[20]
Jeffrey Johnston, and Stefano Fusi
Bin Wang, W. Jeffrey Johnston, and Stefano Fusi. A mathematical theory for understand- ing when abstract representations emerge in neural networks, March 2026. arXiv preprint arXiv:2510.09816
arXiv 2026
-
[21]
Matteo Alleman, Jack W. Lindsey, and Stefano Fusi. Task structure and nonlinearity jointly determine learned representational geometry, January 2024. arXiv preprint arXiv:2401.13558
Pith/arXiv arXiv 2024
-
[22]
Disentangling by Factorising
Hyunjik Kim and Andriy Mnih. Disentangling by Factorising. InProceedings of the 35th International Conference on Machine Learning, pages 2649–2658. PMLR, July 2018
2018
-
[23]
Jeffrey Johnston and Stefano Fusi
W. Jeffrey Johnston and Stefano Fusi. Abstract representations emerge naturally in neural networks trained to perform multiple tasks.Nature Communications, 14(1):1040, February 2023
2023
-
[24]
Mickiewicz, James L
Hanlin Zhu, Melissa Franch, Elizabeth A. Mickiewicz, James L. Belanger, Rhiannon L. Cowan, Kalman A. Katlowitz, Ana G. Chavez, Assia Chericoni, Danika Paulo, Xinyuan Yan, Shervin Rahimpour, Ben Shofty, Eleonora Bartoli, Jay A. Hennig, Nicole R. Provenza, Elliot H. Smith, Steven T. Piantadosi, Benjamin Y . Hayden, and Sameer A. Sheth. A geometric foundatio...
2026
-
[25]
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the non- linear dynamics of learning in deep linear neural networks, February 2014. arXiv preprint arXiv:1312.6120
Pith/arXiv arXiv 2014
-
[26]
Andrew K. Lampinen and Surya Ganguli. An analytic theory of generalization dynamics and transfer learning in deep linear networks, January 2019. arXiv preprint arXiv:1809.10374
Pith/arXiv arXiv 2019
-
[27]
Daniel Kunin, Allan Raventós, Clémentine Dominé, Feng Chen, David Klindt, Andrew Saxe, and Surya Ganguli. Get rich quick: Exact solutions reveal how unbalanced initializations promote rapid feature learning.Advances in Neural Information Processing Systems, 37:81157– 81203, December 2024
2024
-
[28]
Yoonsoo Nam, Seok Hyeong Lee, Clementine C. J. Domine, Yeachan Park, Charles London, Wonyl Choi, Niclas Goring, and Seungjai Lee. Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking), May 2025. arXiv preprint arXiv:2502.21009
Pith/arXiv arXiv 2025
-
[29]
Michaud, Berkan Ottlik, and Joseph Turnbull
Jamie Simon, Daniel Kunin, Alexander Atanasov, Enric Boix-Adserà, Blake Bordelon, Jeremy Cohen, Nikhil Ghosh, Florentin Guth, Arthur Jacot, Mason Kamb, Dhruva Karkada, Eric J. Michaud, Berkan Ottlik, and Joseph Turnbull. There Will Be a Scientific Theory of Deep Learning, April 2026. arXiv preprint arXiv:2604.21691
Pith/arXiv arXiv 2026
-
[30]
Exact learning dynam- ics of deep linear networks with prior knowledge.Advances in Neural Information Processing Systems, 35:6615–6629, December 2022
Lukas Braun, Clémentine Dominé, James Fitzgerald, and Andrew Saxe. Exact learning dynam- ics of deep linear networks with prior knowledge.Advances in Neural Information Processing Systems, 35:6615–6629, December 2022
2022
-
[31]
Clémentine C. J. Dominé, Nicolas Anguita, Alexandra M. Proca, Lukas Braun, Daniel Kunin, Pedro A. M. Mediano, and Andrew M. Saxe. From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks, March 2025. arXiv preprint arXiv:2409.14623
Pith/arXiv arXiv 2025
-
[32]
Korchinski, Dhruva Karkada, Yasaman Bahri, and Matthieu Wyart
Daniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, and Matthieu Wyart. On the Emergence of Linear Analogies in Word Embeddings, October 2025. arXiv preprint arXiv:2505.18651
arXiv 2025
-
[33]
David Saad and Sara A. Solla. On-line learning in soft committee machines.Physical Review E, 52(4):4225–4243, October 1995
1995
-
[34]
Du, Wei Hu, Zhiyuan Li, and Ruosong Wang
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang. Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks, May
-
[35]
arXiv preprint arXiv:1901.08584
Pith/arXiv arXiv 1901
-
[36]
Raphaël Barboni, Gabriel Peyré, and François-Xavier Vialard. Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime, July 2025. arXiv preprint arXiv:2504.18208
Pith/arXiv arXiv 2025
-
[37]
Leveraging the two timescale regime to demonstrate convergence of neural networks, October 2023
Pierre Marion and Raphaël Berthier. Leveraging the two timescale regime to demonstrate convergence of neural networks, October 2023. arXiv preprint arXiv:2304.09576. 12
Pith/arXiv arXiv 2023
-
[38]
A Convergence Analysis of Gradi- ent Descent for Deep Linear Neural Networks, October 2019
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu. A Convergence Analysis of Gradi- ent Descent for Deep Linear Neural Networks, October 2019. arXiv preprint arXiv:1810.02281
Pith/arXiv arXiv 2019
-
[39]
Kernel Methods for Deep Learning
Youngmin Cho and Lawrence Saul. Kernel Methods for Deep Learning. InAdvances in Neural Information Processing Systems, volume 22. Curran Associates, Inc., 2009
2009
-
[40]
Computing with Infinite Networks
Christopher Williams. Computing with Infinite Networks. InAdvances in Neural Information Processing Systems, volume 9. MIT Press, 1996
1996
-
[41]
Oriane Siméoni, Huy V . V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julie...
Pith/arXiv arXiv 2025
-
[42]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Anto...
Pith/arXiv arXiv 2024
-
[43]
Majaj, Ha Hong, Ethan A
Najib J. Majaj, Ha Hong, Ethan A. Solomon, and James J. DiCarlo. Simple Learned Weighted Sums of Inferior Temporal Neuronal Firing Rates Accurately Predict Human Core Object Recognition Performance.Journal of Neuroscience, 35(39):13402–13418, September 2015
2015
-
[44]
Analogies Explained: Towards Understanding Word Embeddings, May 2019
Carl Allen and Timothy Hospedales. Analogies Explained: Towards Understanding Word Embeddings, May 2019. arXiv preprint arXiv:1901.09813
Pith/arXiv arXiv 2019
-
[45]
Simon, Yasaman Bahri, and Michael R
Dhruva Karkada, James B. Simon, Yasaman Bahri, and Michael R. DeWeese. Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models, February 2025
2025
-
[46]
Korchinski, Andres Nava, Matthieu Wyart, and Yasaman Bahri
Dhruva Karkada, Daniel J. Korchinski, Andres Nava, Matthieu Wyart, and Yasaman Bahri. Symmetry in language statistics shapes the geometry of model representations, February 2026
2026
-
[47]
Separable nonlinear least squares: The variable projection method and its applications.Inverse Problems, 19(2):R1, February 2003
Gene Golub and Victor Pereyra. Separable nonlinear least squares: The variable projection method and its applications.Inverse Problems, 19(2):R1, February 2003
2003
-
[48]
Training Two-Layered Feedforward Networks With Variable Projection Method.IEEE Transactions on Neural Networks, 19(2):371–375, February 2008
Cheol-Taek Kim and Ju-Jang Lee. Training Two-Layered Feedforward Networks With Variable Projection Method.IEEE Transactions on Neural Networks, 19(2):371–375, February 2008
2008
-
[49]
abstract
Jack W Lindsey and Elias B Issa. Factorized visual representations in the primate visual system and deep neural networks.eLife, 13:RP91685, July 2024. 13 Appendix Appendix Contents A Supplementary discussion 15 A.1 Further related work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 A.2 Relationship between abstraction and linear repr...
2024
-
[50]
Each concept is instantiated by 80 ordered word pairs comprising common words
as well as an additional 14 English–X translation directions (where X is Arabic, Chinese, Dutch, German, Indonesian, Italian, Japanese, Korean, Persian, Polish, Portuguese, Russian, Spanish, and Turkish). Each concept is instantiated by 80 ordered word pairs comprising common words. See Table 1 for the English–Spanish word pairs. English Spanish English S...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.