Pith. sign in

REVIEW 2 major objections 4 minor 122 references

The Role of Rigor in Artificial Intelligence

T0 review · 2 major / 4 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Modern AI advances mainly by operational rigor—benchmarks and deployment reliability—while conceptual clarity and scientific understanding lag, explaining both its speed and its uncertainties.

desk verdict Clean conceptual synthesis that names why deep learning runs on operational rigor; useful organizing lens, not a tested theory. read the letter →

arxiv 2607.03634 v1 pith:TBZXODEC submitted 2026-05-19 cs.AI

classification cs.AI
keywords rigorinAIconceptualepistemicoperationaldeeplearningbenchmarksintelligencealignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that artificial intelligence has produced striking capabilities without the conceptual foundations or scientific understanding that usually precede reliable technology in mature fields. It organizes the field’s problems under three kinds of rigor: conceptual rigor, which clarifies contested notions such as intelligence and understanding; epistemic rigor, which requires reproducibility, predictability, and explainability; and operational rigor, which evaluates and steers systems through benchmarks, post-training, and safety procedures. The central claim is that the distinctive path of modern deep learning comes from how these forms of rigor interact across successive paradigms, leaving operational rigor dominant. That dominance lets performance improve rapidly through metric-driven iteration even when theory is thin, while also leaving persistent gaps in generalization, robustness, and alignment. A reader who cares about whether AI can become a mature science and trustworthy technology will find a map of where each form of rigor is strong, where it is weak, and what must still be developed.

What carries the argument

A three-part framework of rigor—conceptual (clear foundational concepts and paradigms), epistemic (reproducibility, predictability, explainability), and operational (benchmarks, reliability, and safety practices)—used as the diagnostic lens for AI’s history, methods, and future bottlenecks.

What would settle it

If a future AI paradigm (or a careful historical re-analysis of current deep learning) showed that lasting capability gains required simultaneous advances in conceptual and epistemic rigor rather than operational metric-chasing, or if an alternative rigor taxonomy better predicted progress and failure modes, the central claim would be undercut.

Watch

Extended reading notes

Core claim

The distinctive trajectory of AI arises from how conceptual, epistemic, and operational rigor interact across paradigms, resulting in the primacy of operational rigor in modern deep learning; that primacy explains both the field’s rapid capability gains and its lasting uncertainties, and it clarifies what is required to turn AI into a mature science and reliable technology.

Load-bearing premise

The analysis rests on the premise that this three-way split of rigor is the right and sufficiently complete way to diagnose AI’s scientific status, rather than some other taxonomy of standards.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper argues that AI's distinctive trajectory—rapid capability gains with persistent conceptual and scientific uncertainty—arises from the interaction of three forms of rigor: conceptual (clarity of foundational terms and paradigms), epistemic (reproducibility, predictability, explainability), and operational (benchmarks, reliability, and safety). It applies this framework to contested notions of intelligence and understanding, the empirical character of deep learning, the strengths and pathologies of benchmarks, and the historical succession of symbolic, classical statistical-learning, and connectionist/deep-learning paradigms. The central claim is that modern deep learning has elevated operational rigor above the other two, which both explains progress and clarifies the obstacles to maturing AI as a science and reliable technology.

Significance. If accepted as a useful analytic lens, the paper supplies a coherent organizing vocabulary for a multidisciplinary field whose progress is often discussed in fragmented or polemical terms. The historical sketch of paradigms and the treatment of benchmarks, reproducibility distinctions, and alignment are standard but well-integrated; the explicit contrast between AGI (favoring operational rigor) and alignment (requiring conceptual and epistemic rigor) is a clear contribution. The work is philosophical and taxonomic rather than theorematic or empirical; its value lies in clarifying structure and priorities rather than in new derivations or falsifiable predictions. Strengths include careful citation of the literature and a measured tone that avoids both hype and pure critique.

major comments (2)
  1. The three-part taxonomy is introduced by stipulation in the Introduction and then applied throughout; its exhaustiveness and superiority relative to alternatives (e.g., Olteanu et al. [1] or classical philosophy-of-science categories) are not independently argued or tested. Because the central claim—that the distinctive trajectory of AI arises from how these forms interact, with operational rigor primary under deep learning—depends on this partition, the manuscript should either (a) defend the partition more explicitly against nearby alternatives or (b) state more clearly that the framework is provisional and heuristic rather than uniquely privileged.
  2. Section 4.2's historical narrative (symbolic → classical statistical learning → connectionism/deep learning) is standard and well-cited, but the claim that the current primacy of operational rigor is "historically contingent rather than inevitable" remains under-supported. A brief discussion of what would count as evidence that a future paradigm rebalanced the three forms (or of counter-examples already present) would strengthen the load-bearing historical claim.
minor comments (4)
  1. The footnote distinguishing the present three-part scheme from Olteanu et al. [1] is useful but brief; a short paragraph in the Introduction or Conclusion comparing the two taxonomies would help readers locate the contribution.
  2. Section 2.3's discussion of explainability vs. interpretability is careful, yet the claim that deep learning "defies effective hierarchical abstraction" could be sharpened with one or two concrete examples of failed localization (e.g., attribution of a particular failure mode to data vs. architecture vs. optimization).
  3. Occasional informal phrasing ("alchemy," "jagged intelligence") is already hedged, but ensuring each such term is immediately tied to a citation or definition would further reduce ambiguity.
  4. References to scaling laws and infinite-width theories are accurate; a brief note on known caveats (already alluded to via Hooker [48]) would keep the predictability discussion balanced.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the three-part rigor taxonomy is a stipulated analytic lens applied to known history, not a derivation that reduces to its inputs by construction.

full rationale

This is a philosophical/analytic position paper, not a quantitative derivation. The central claim—that AI's distinctive trajectory (rapid capability growth with lagging conceptual and scientific understanding) arises from the interaction of conceptual, epistemic, and operational rigor, with operational rigor becoming primary under deep learning—is an organizing thesis introduced by stipulation in the Introduction and then used to re-describe well-known historical paradigms (symbolic AI, classical statistical learning, connectionism/deep learning), the empirical character of modern deep learning, and the roles of benchmarks and alignment. There are no equations, fitted parameters, or 'predictions' that reduce by construction to inputs. The only mild definitional element is the author's introduction of the three categories themselves; once accepted as a provisional lens, the subsequent historical and diagnostic claims do not loop back to force those categories. Citations are overwhelmingly to independent sources (Turing, McCarthy, Legg & Hutter, Kaplan et al. scaling laws, ImageNet, Goodhart's law literature, alignment papers, etc.); the single self-citation is the author's own prior NeurIPS paper on n-gram statistics of transformers, which is used only as an illustrative example of training-data regurgitation and is not load-bearing for the framework or the primacy-of-operational-rigor thesis. No uniqueness theorem is imported from the author's prior work, no ansatz is smuggled via self-citation, and no known empirical pattern is merely renamed as a new result. Score 1 reflects only the ordinary philosophical practice of defining terms and then applying them; the paper is self-contained as an interpretive essay and exhibits no circular reduction of the kind the analyzer is charged to detect.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The paper is a conceptual analysis, not a formal derivation. Its load-bearing commitments are definitional and historical rather than free parameters or new physical entities. The main axioms are the three-part taxonomy itself and the claim that deep learning's structure privileges operational over epistemic rigor. No numerical free parameters appear. Invented entities are limited to the named rigor categories, which function as analytic tools rather than postulated mechanisms with independent empirical handles.

assumptions (4)
  • ad hoc to paper Rigor in AI is usefully partitioned into conceptual, epistemic, and operational forms that interact across paradigms.
    Introduced by stipulation in the introduction; alternative partitions (e.g. Olteanu et al. six-part) exist and are not refuted.
  • domain assumption Modern deep learning is characterized by a tight feedback loop in which the same metrics used for evaluation are also optimization targets.
    Stated in §4.1; widely accepted but not independently measured in the paper.
  • domain assumption AI produces the artifacts it studies, so conceptual and epistemic inquiry always chase a moving target.
    §4.1; structural claim used to explain uneven progress.
  • domain assumption Earlier AI paradigms (symbolic, classical statistical learning) embodied different balances of the three rigor forms than deep learning.
    §4.2 historical narrative; standard but selective.
invented entities (1)
  • Three-part rigor framework (conceptual / epistemic / operational)
    purpose: Organize analysis of AI's scientific and technological status and explain primacy of operational rigor.
    Analytic categories introduced by the paper; no independent empirical prediction beyond re-description of known history.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Role of Rigor in Artificial Intelligence." pith.science (2026). https://pith.science/paper/TBZXODEC

@misc{pith2026260703634,
  author       = {Pith},
  title        = {Pith review of: The Role of Rigor in Artificial Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TBZXODEC}},
  note         = {Machine review of arXiv:2607.03634}
}
read the original abstract

Artificial intelligence (AI) has achieved extraordinary capabilities despite lacking many of the conceptual and scientific foundations associated with mature disciplines. Unlike traditional sciences, where reliable technology typically emerges from theoretical understanding, modern AI has progressed largely through performance-driven iteration and "alchemical" experimentation. This tension motivates a systematic analysis of AI through the lens of rigor. We introduce a three-part framework consisting of conceptual rigor (clarifying foundational concepts), epistemic rigor (establishing scientific understanding), and operational rigor (ensuring reliable performance and deployment). Using this framework, we analyze competing conceptions of intelligence and understanding, the strengths and limitations of the empirical approach to deep learning, the power and pitfalls of benchmarks, and the obstacles to theory development posed by modern AI systems. We argue that the distinctive trajectory of AI arises from how forms of rigor interact across paradigms, resulting in the primacy of operational rigor in modern deep learning. This perspective helps explain both AI's rapid advances and its persistent uncertainties, while clarifying the challenges involved in transforming AI into a mature science and reliable technology.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

122 extracted references · 10 linked inside Pith

  1. [1]

    Rigor in AI: Doing rigorous AI work requires a broader, responsible AI-informed conception of rigor

    Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang, Fernando Diaz, Flavio Calmon, Margaret Mitchell, Michael Ekstrand, Reuben Binns, and Solon Barocas. Rigor in AI: Doing rigorous AI work requires a broader, responsible AI-informed conception of rigor. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems Position Pa...

  2. [2]

    Alan M. Turing. Computing machinery and intelligence.Mind, 59(236):433–460, 1950

  3. [3]

    A proposal for the Dartmouth summer research project on Artificial Intelligence, 1955

    John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. A proposal for the Dartmouth summer research project on Artificial Intelligence, 1955. Proposal for the 1956 Dart- mouth workshop. URL:http://jmc.stanford.edu/articles/dartmouth/dartmouth.pdf

  4. [4]

    Luciano Floridi and Anna C. Nobre. Anthropomorphising machines and computerising minds: The crosswiring of languages between artificial intelligence and brain & cognitive sciences.Cen- tre for Digital Ethics (CEDE) Research Paper, 2024. Available at SSRN:https://ssrn.com/ abstract=4738331

  5. [5]

    Pei Wang.Non-Axiomatic Reasoning System: Exploring the Essence of Intelligence. Ph.d. thesis, Indiana University, Bloomington, IN, USA, 1995

  6. [6]

    A collection of definitions of intelligence

    Shane Legg and Marcus Hutter. A collection of definitions of intelligence. InAdvances in Artificial General Intelligence: Concepts, Architectures and Algorithms, pages 17–24, Amsterdam, The Netherlands, 2007. IOS Press. 13

  7. [7]

    On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019

    Fran¸ cois Chollet. On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019

  8. [8]

    Sparks of artificial general intelligence: Early experiments with GPT-4, 2023

    S´ ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with GPT-4, 2023

Show all 122 references
  1. [9]

    Godfather of AI

    60 Minutes. “Godfather of AI”: Geoffrey Hinton – the 60 minutes interview, Oct 2023. YouTube video, 13:12. URL:https://www.youtube.com/watch?v=qrvK_KuIeJk

  2. [10]

    Godfather of AI

    The Royal Institution. Will AI outsmart human intelligence? - with “Godfather of AI” Ge- offrey Hinton, Jul 2025. YouTube video, 47:15. URL:https://www.youtube.com/watch?v= IkdziSLYzHw

  3. [11]

    A conversation with Yann LeCun AI: Lifeline or landmine?, Feb

    World Governments Summit. A conversation with Yann LeCun AI: Lifeline or landmine?, Feb

  4. [12]

    URL:https://www.youtube.com/watch?v=rf9jgZYAni8

    YouTube video, 24:51. URL:https://www.youtube.com/watch?v=rf9jgZYAni8

  5. [13]

    Krakauer, John W

    David C. Krakauer, John W. Krakauer, and Melanie Mitchell. Large language models and emergence: A complex systems perspective, 2025

  6. [14]

    Krakauer

    Melanie Mitchell and David C. Krakauer. The Debate over Understanding in AI’s Large Language Models.Proceedings of the National Academy of Sciences, 120(13):e2215907120, March 2023

  7. [15]

    Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, Fran¸ cois Candelon, and Karim R

    Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, Fran¸ cois Candelon, and Karim R. Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowl...

  8. [16]

    Jagged Intelligence, 2024

    Andrej Karpathy. Jagged Intelligence, 2024. Tweet on X (formerly Twitter), July 2024. URL: https://x.com/karpathy/status/1816531576228053133

  9. [17]

    Why 9.11 is larger than 9.9

    OpenAI Community. Why 9.11 is larger than 9.9. . . . . . incredible. Online forum post, Jul 2024. Discussion on comparing numeric values; community.openai.com thread 869824. URL:https: //community.openai.com/t/why-9-11-is-larger-than-9-9-incredible/869824

  10. [18]

    a is b” fail to learn “b is a

    Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reversal curse: LLMs trained on “a is b” fail to learn “b is a”. InThe Twelfth International Conference on Learning Representations, 2024

  11. [19]

    Hutchinson, London, 1949

    Gilbert Ryle.The Concept of Mind. Hutchinson, London, 1949

  12. [20]

    Fodor.The Language of Thought

    Jerry A. Fodor.The Language of Thought. Harvard University Press, 1975

  13. [21]

    Psychological predicates

    Hilary Putnam. Psychological predicates. In W. H. Capitan and D. D. Merrill, editors,Art, Mind, and Religion, pages 37–48. University of Pittsburgh Press, 1967

  14. [22]

    The symbol grounding problem.Physica D: Nonlinear Phenomena, 42(1–3):335– 346, 1990

    Stevan Harnad. The symbol grounding problem.Physica D: Nonlinear Phenomena, 42(1–3):335– 346, 1990

  15. [23]

    From System 1 Deep Learning to System 2 Deep Learning

    Yoshua Bengio. From System 1 Deep Learning to System 2 Deep Learning. NeurIPS 2019 keynote talk on YouTube, 2019. Accessed 2026-01-29. URL:https://www.youtube.com/watch? v=FtUbMG3rlFs

  16. [24]

    Othello-GPT

    Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Vi´ egas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. InInternational Conference on Learning Representations (ICLR) 2023, 2023. Also known as...

  17. [25]

    Genie 3: A new frontier for world mod- els

    Jack Parker-Holder and Shlomi Fruchter. Genie 3: A new frontier for world mod- els. DeepMind Blog, Aug 2025. Available at:https://deepmind.google/blog/ genie-3-a-new-frontier-for-world-models/. 14

  18. [26]

    A Path Towards Autonomous Machine Intelligence

    Yann LeCun. A Path Towards Autonomous Machine Intelligence. Working paper, OpenReview,

  19. [27]

    Version 0.9.2, June 27, 2022

  20. [28]

    Meta’s AI chief Yann LeCun on AGI, open-source, and AI risk.TIME, February

    Charlotte Alter. Meta’s AI chief Yann LeCun on AGI, open-source, and AI risk.TIME, February

  21. [29]

    Interview with Yann LeCun

  22. [30]

    Meaning without reference in large language models

    Steven Piantadosi and Felix Hill. Meaning without reference in large language models. In NeurIPS 2022 Workshop on Neuro Causal and Symbolic AI (nCSI), 2022

  23. [31]

    Nastase, Martin Chodorow, Mengru Wu, and Ping Li

    Qianwen Xu, Yixuan Peng, Samuel A. Nastase, Martin Chodorow, Mengru Wu, and Ping Li. Large language models without grounding recover non-sensorimotor but not sensorimotor fea- tures of human concepts.Nature Human Behaviour, 9(9):1871–1886, 2025

  24. [32]

    Bender and Alexander Koller

    Emily M. Bender and Alexander Koller. Climbing towards NLU: On meaning, form, and under- standing in the age of data. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online, July 2020. Association for Computational Li...

  25. [33]

    Universal intelligence: A definition of machine intelligence

    Shane Legg and Marcus Hutter. Universal intelligence: A definition of machine intelligence. Minds and Machines, 17(4):391–444, 2007

  26. [34]

    Trends in the dollar training cost of machine learning systems

    Ben Cottier. Trends in the dollar training cost of machine learning systems. Epoch AI Blog, 2023. Published January 31, 2023; accessed 2026-01-29. URL:https://epoch.ai/blog/ trends-in-the-dollar-training-cost-of-machine-learning-systems

  27. [35]

    How much power will frontier AI training demand in 2030? Epoch AI Blog, 2025

    Josh You and David Owen. How much power will frontier AI training demand in 2030? Epoch AI Blog, 2025. Published August 11, 2025; accessed 2026-01-29. URL:https://epoch.ai/blog/ power-demands-of-frontier-ai-training

  28. [36]

    The AI index 2023 annual report

    Nestor Maslej, Loredana Fattorini, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Helen Ngo, Juan Carlos Niebles, Vanessa Parli, Yoav Shoham, Rus- sell Wald, Jack Clark, and Raymond Perrault. The AI index 2023 annual report. Re- port, Institute...

  29. [37]

    Unreproducible research is reproducible

    Xavier Bouthillier, C´ esar Laurent, and Pascal Vincent. Unreproducible research is reproducible. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors,Proceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research,...

  30. [38]

    Deep reinforcement learning that matters

    Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. InProceedings of the 32nd AAAI Conference on Artificial Intelligence, AAAI’18, pages 3207–3214. AAAI Press, 2018

  31. [39]

    Are GANs created equal? A large-scale study.arXiv preprint arXiv:1711.10337, 2017

    Mario Lucic, Karol Kurach, Marcin Michalski, Sylvain Gelly, and Olivier Bousquet. Are GANs created equal? A large-scale study.arXiv preprint arXiv:1711.10337, 2017

  32. [40]

    Random search and reproducibility for neural architecture search

    Liam Li and Ameet Talwalkar. Random search and reproducibility for neural architecture search. In Ryan P. Adams and Vibhav Gogate, editors,Proceedings of The 35th Uncertainty in Artificial Intelligence Conference, volume 115 ofProceedings of Machine Learning Research, pages 36...

  33. [41]

    Julian D

    Moritz Herrmann, F. Julian D. Lange, Katharina Eggensperger, Giuseppe Casalicchio, Marcel Wever, Matthias Feurer, David R¨ ugamer, Eyke H¨ ullermeier, Anne-Laure Boulesteix, and Bernd Bischl. Position: why we must rethink empirical research in machine learning. InProceedings o...

  34. [42]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020. 15

  35. [43]

    Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell

    James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, An- drei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic for- getti...

  36. [44]

    Extracting training data from large language models

    Nicholas Carlini, Florian Tram` er, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, ´Ulfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In30th USENIX Security Symposium (...

  37. [45]

    Understanding transformers via n-gram statistics

    Timothy Nguyen. Understanding transformers via n-gram statistics. InAdvances in Neural Information Processing Systems 37 (NeurIPS 2024). Curran Associates, Inc., 2024

  38. [46]

    Neural tangent kernel: Convergence and generalization in neural networks

    Arthur Jacot, Franck Gabriel, and Cl´ ement Hongler. Neural tangent kernel: Convergence and generalization in neural networks. InAdvances in Neural Information Processing Systems, pages 8580–8589, Red Hook, NY, USA, 2018. Curran Associates, Inc

  39. [47]

    Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl- Dickstein, and Jeffrey Pennington

    Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl- Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. InAdvances in Neural Information Processing Systems, volume 32, 2019. N...

  40. [48]

    Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit

    Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit. In Kamalika Chaudhuri and Ruslan Salakhut- dinov, editors,Proceedings of Machine Learning Research, volume 99 ofProceedings of Machine...

  41. [49]

    Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao

    Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer.arXiv preprint arXiv:2203.03466, 2022

  42. [50]

    Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou

    Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianine- jad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically, 2017

  43. [51]

    On the slow death of scaling

    Sara Hooker. On the slow death of scaling. SSRN Scholarly Paper ID 5877662. Available at SSRN:https://ssrn.com/abstract=5877662, December 2025

  44. [52]

    Michaud, Berkan Ottlik, and Joseph Turnbull

    Jamie Simon, Daniel Kunin, Alexander Atanasov, Enric Boix-Adser` a, Blake Bordelon, Jeremy Cohen, Nikhil Ghosh, Florentin Guth, Arthur Jacot, Mason Kamb, Dhruva Karkada, Eric J. Michaud, Berkan Ottlik, and Joseph Turnbull. There will be a scientific theory of deep learning. ar...

  45. [53]

    Goodfellow, Jonathon Shlens, and Christian Szegedy

    Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adver- sarial examples. InInternational Conference on Learning Representations (ICLR), 2015

  46. [54]

    Accounting for variance in machine learning benchmarks

    Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Nazanin Mohammadi Sepahvand, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Tal Arbel, Chris Pal, Gael Varoquaux, and Pascal Vincent. Accou...

  47. [55]

    AI researchers allege that machine learning is alchemy

    Matthew Hutson. AI researchers allege that machine learning is alchemy. News article,Sci- ence, May 3 2018. Accessed: 2025-01-05. URL:https://www.science.org/content/article/ ai-researchers-allege-machine-learning-alchemy

  48. [56]

    Brief introduction to deep learning and the “alchemy” controversy

    Sanjeev Arora. Brief introduction to deep learning and the “alchemy” controversy. Video, YouTube, 2019. Presented at Deep Learning: Alchemy or Science?, Institute for Advanced Study. URL:https://www.youtube.com/watch?v=kqhg-o-KEns. 16

  49. [57]

    Zico Kolter

    J. Zico Kolter. Is this really science? A lukewarm defense of alchemy. InWorkshop on Scientific Methods for Understanding Neural Networks, NeurIPS 2024, 2024

  50. [58]

    Vapnik.The Nature of Statistical Learning Theory

    Vladimir N. Vapnik.The Nature of Statistical Learning Theory. Springer, New York, 2 edition, 2000

  51. [59]

    Cambridge University Press, 2014

    Shai Shalev-Shwartz and Shai Ben-David.Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014

  52. [60]

    Approximation theory of the MLP model in neural networks.Acta Numerica, 8:143–195, 1999

    Allan Pinkus. Approximation theory of the MLP model in neural networks.Acta Numerica, 8:143–195, 1999

  53. [61]

    Springer Science & Business Media, New York, second edition, 2006

    Jorge Nocedal and Stephen Wright.Numerical Optimization. Springer Science & Business Media, New York, second edition, 2006

  54. [62]

    Introduction to online convex optimization, 2023

    Elad Hazan. Introduction to online convex optimization, 2023

  55. [63]

    Reconciling modern machine- learning practice and the bias–variance trade-off.Proc

    Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine- learning practice and the bias–variance trade-off.Proc. Natl. Acad. Sci. U.S.A., 116(32):15849– 15854, 2019. PMCID: PMC6689936, PMID: 31341078

  56. [64]

    The implicit bias of gradient descent on separable data, 2017

    Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data, 2017

  57. [65]

    No spurious local minima in nonconvex low rank problems: A unified geometric analysis, 2017

    Rong Ge, Chi Jin, and Yi Zheng. No spurious local minima in nonconvex low rank problems: A unified geometric analysis, 2017

  58. [66]

    Du, Jason D

    Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient descent finds global minima of deep neural networks, 2018

  59. [67]

    Knowledge editing for large language models: A survey.ACM Computing Surveys, 57(3):59:1–59:35, 2024

    Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. Knowledge editing for large language models: A survey.ACM Computing Surveys, 57(3):59:1–59:35, 2024

  60. [68]

    Improving alignment and robustness with circuit breakers

    Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks. Improving alignment and robustness with circuit breakers. InAdvances in Neural Information Processing Systems 37, pages 83345–83373,...

  61. [69]

    Toward faithful retrieval-augmented generation with sparse autoencoders, 2025

    Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, and Aidong Zhang. Toward faithful retrieval-augmented generation with sparse autoencoders, 2025

  62. [70]

    Cor- rectness assessment of code generated by large language models using internal representations

    Tuan-Dung Bui, Thanh Trong Vu, Thu-Trang Nguyen, Son Nguyen, and Hieu Dinh Vo. Cor- rectness assessment of code generated by large language models using internal representations. Journal of Systems and Software, 230:112570, December 2025

  63. [71]

    Zachary C. Lipton. The Mythos of Model Interpretability: In Machine Learning, the Concept of Interpretability Is Both Important and Poorly Defined.Commun. ACM, 61(10):36–43, September 2018

  64. [72]

    Sarthak Jain and Byron C. Wallace. Attention Is Not Explanation. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3543–3556, 2019

  65. [73]

    Attention is not not explanation

    Sarah Wiegreffe and Yuval Pinter. Attention is not not explanation. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural La...

  66. [74]

    Everything, everywhere, all at once: Is mechanistic interpretability identifiable? InThe Thirteenth International Confer- ence on Learning Representations, 2025

    Maxime M´ eloux, Silviu Maniu, Fran¸ cois Portet, and Maxime Peyrard. Everything, everywhere, all at once: Is mechanistic interpretability identifiable? InThe Thirteenth International Confer- ence on Learning Representations, 2025. 17

  67. [75]

    Schiller, Filippos Stamatiou, and Anders Søgaard

    Iwan Williams, Ninell Oldenburg, Ruchira Dhar, Joshua Hatherley, Constanza Fierro, Nina Ra- jcic, Sandrine R. Schiller, Filippos Stamatiou, and Anders Søgaard. Mechanistic interpretability needs philosophy.arXiv:2506.18852 [cs.CL], 2025. Preprint; accessed 2026-01-29

  68. [76]

    Cambridge University Press, 2 edition, 2009

    Judea Pearl.Causality: Models, Reasoning, and Inference. Cambridge University Press, 2 edition, 2009

  69. [77]

    ImageNet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255, 2009

  70. [78]

    Recognition in terra incognita

    Sara Beery, Grant Van Horn, and Pietro Perona. Recognition in terra incognita. InProceedings of the European Conference on Computer Vision (ECCV), pages 456–473, Munich, Germany,

  71. [79]

    Zemel, Wieland Brendel, Matthias Bethge, and Felix A

    Robert Geirhos, J¨ orn-Henrik Jacobsen, Claudio Michaelis, Richard S. Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

  72. [80]

    Benchmarking neural network robustness to common corruptions and perturbations

    Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. InInternational Conference on Learning Representations, 2019

  73. [81]

    GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models

    Seyed Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models. InThe Thirteenth International Conference on Learning Representa- tions, 2025

  74. [82]

    Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

    Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben- jamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei,...

  75. [83]

    Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals

    Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The benchmark lottery, 2021

  76. [84]

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. RealToxi- cityPrompts: Evaluating neural toxic degeneration in language models. In Trevor Cohn, Yulan He, and Yang Liu, editors,Findings of the Association for Computational Linguistics: EMNLP 2020, ...

  77. [85]

    Man ends his life after an AI chatbot ’encour- aged’ him to sacrifice himself to stop climate change, March 2023

    Imane El Atillah. Man ends his life after an AI chatbot ’encour- aged’ him to sacrifice himself to stop climate change, March 2023. Ac- cessed: 2026-03-04. URL:https://www.euronews.com/next/2023/03/31/ man-ends-his-life-after-an-ai-chatbot-encouraged-him-to-sacrifice-himself-t...

  78. [86]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dess` ı, Roberta Raileanu, Maria Lomeli, Eric Ham- bro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. InAdvances in Neural Information Processing Systems 36 (NeurIPS 2...

  79. [87]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022

  80. [88]

    Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534– 46594, 2023

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Shashank Srivastava, Yiming Yang, and Hannaneh Hajishirzi. Self-refine: Iterative refinement with self-feedback.Advances in Neural Information Processing Systems, 36:46534– 46594, 2023

  81. [89]

    Finetuned language models are zero-shot learners

    Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations (ICLR), 2022

  82. [90]

    Deep reinforcement learning from human preferences

    Paul Christiano et al. Deep reinforcement learning from human preferences. InAdvances in Neural Information Processing Systems, 2017

  83. [91]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Misha Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike...

  84. [92]

    Google AI search says to glue pizza and eat rocks, May 2024

    BBC News. Google AI search says to glue pizza and eat rocks, May 2024. Accessed 2026-01-31. URL:https://www.bbc.co.uk/news/articles/cd11gzejgz4o

  85. [93]

    Replit CEO apologizes after AI coding tool wipes company database, July 2025

    Alistair Barr. Replit CEO apologizes after AI coding tool wipes company database, July 2025. Business Insider article on a reported incident where an AI coding agent deleted a production database; accessed 2026-01-31. URL:https://www.businessinsider. com/replit-ceo-apologizes-...

  86. [94]

    A comprehensive study of jailbreak attack versus defense for large language models

    Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. A comprehensive study of jailbreak attack versus defense for large language models. InFindings of the Association for Computational Linguistics: ACL 2024, pages 7432–7449, Bangkok, Thailand, August 2024. Association ...

  87. [95]

    Wenrui Xu and Keshab K. Parhi. A survey of attacks on large language models, 2025

  88. [96]

    A survey of attacks on large vision–language models: Resources, advances, and future trends.IEEE Transactions on Neural Networks and Learning Systems, 36(11):19525–19545, 2025

    Daizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou, Yu Cheng, and Wei Hu. A survey of attacks on large vision–language models: Resources, advances, and future trends.IEEE Transactions on Neural Networks and Learning Systems, 36(11):19525–19545, 2025

  89. [97]

    Ramirez, Song-Kyoo Kim, Hussam M

    Miguel A. Ramirez, Song-Kyoo Kim, Hussam M. N. Al Hamadi, Ernesto Damiani, Young-Ji Byon, Tae-Yeon Kim, Chung-Suk Cho, and Chan Yeob Yeun. Poisoning attacks and defenses on artificial intelligence: A survey.CoRR, abs/2202.10276, 2022. arXiv preprint

  90. [98]

    Data poisoning in deep learning: A survey.CoRR, abs/2503.22759, 2025

    Pinlong Zhao, Weiyao Zhu, Pengfei Jiao, Di Gao, and Ou Wu. Data poisoning in deep learning: A survey.CoRR, abs/2503.22759, 2025. arXiv preprint

  91. [99]

    Introduction to neural network verification.arXiv preprint arXiv:2109.10317, 2021

    Aws Albarghouthi. Introduction to neural network verification.arXiv preprint arXiv:2109.10317, 2021

  92. [100]

    Weiming Xiang, Patrick Musau, Nathan Hamilton, Xiaodong Yang, Jaemin Rosenfeld, and Tay- lor T. Johnson. Verification for machine learning, autonomy, and neural networks survey.arXiv preprint arXiv:1810.01989, 2018

  93. [101]

    Pappas, and Insup Lee

    Radoslav Ivanov, James Weimer, Rajeev Alur, George J. Pappas, and Insup Lee. Verisig: Veri- fying safety properties of hybrid systems with neural network controllers.ACM Transactions on Embedded Computing Systems, 18(5s):1–19, 2019

  94. [102]

    Olympiad-level formal mathematical reasoning with reinforcement learning

    Thomas Hubert, Edward Williams, Jiangjie Chen, Wenxiang Chen, Jiacheng Du, Thomas Han- wen Zhu, et al. Olympiad-level formal mathematical reasoning with reinforcement learning. Nature, Nov 2025. 19

  95. [103]

    Russell and Peter Norvig.Artificial Intelligence: A Modern Approach

    Stuart J. Russell and Peter Norvig.Artificial Intelligence: A Modern Approach. Pearson, Hobo- ken, NJ, 4 edition, 2020

  96. [104]

    Rumelhart and James L

    David E. Rumelhart and James L. McClelland, editors.Parallel Distributed Processing: Explo- rations in the Microstructure of Cognition, Volume 1: Foundations. MIT Press, Cambridge, MA, 1986

  97. [105]

    Economic policy challenges for the age of AI

    Anton Korinek. Economic policy challenges for the age of AI. NBER Working Paper 32980, National Bureau of Economic Research, September 2024

  98. [106]

    Viking, New York, NY, USA, 2024

    Ray Kurzweil.The Singularity Is Nearer. Viking, New York, NY, USA, 2024

  99. [107]

    Hopes and fears for intelligent machines in fiction and reality

    Stephen Cave and Kanta Dihal. Hopes and fears for intelligent machines in fiction and reality. Nature Machine Intelligence, 1(2):74–78, 2019

  100. [108]

    308 — Alison Gopnik on children, AI, and modes of thinking, March 2025

    Sean Carroll and Alison Gopnik. 308 — Alison Gopnik on children, AI, and modes of thinking, March 2025. Podcast episode with transcript. Accessed: 2026-05-19. URL:https://www.preposterousuniverse.com/podcast/2025/03/17/ 308-alison-gopnik-on-children-ai-and-modes-of-thinking/

  101. [109]

    EP20: Yann LeCun, December 15 2025

    The Information Bottleneck. EP20: Yann LeCun, December 15 2025. YouTube video, 01:50:07. URL:https://www.youtube.com/watch?v=7u-DXVADyhc

  102. [110]

    Anthropic CEO Dario Amodei: AI’s potential, OpenAI rivalry, GenAI business, doomerism

    Dario Amodei and Alex Kantrowitz. Anthropic CEO Dario Amodei: AI’s potential, OpenAI rivalry, GenAI business, doomerism. YouTube Video, 2025. Interview published July 31, 2025. Accessed 2026-01-31. URL:https://www.youtube.com/watch?v=mYDSSRS-B5U

  103. [111]

    Artificial General Intelligence Is Already Here,

    Blaise Ag¨ uera y Arcas and Peter Norvig. Artificial General Intelligence Is Already Here,

  104. [112]

    URL:https://www.noemamag.com/ artificial-general-intelligence-is-already-here/

    Published October 10, 2023; Accessed: 2026-01-29. URL:https://www.noemamag.com/ artificial-general-intelligence-is-already-here/

  105. [113]

    A definition of AGI, 2025

    Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, Andy Zou, Lionel Levine, Bo Han, Jie Fu, Ziwei Liu, Jinwoo Shin, Kimin Lee, Mantas Mazeika, Long Phan, George Ingebretsen, Adam Khoja, Cihang Xie, Olawale Salaudeen, Matthias Hei...

  106. [114]

    Galatzer-Levy, Meredith Ringel Morris, Allan Dafoe, Alison M

    Ryan Burnell, Yumeya Yamamori, Orhan Firat, Kate Olszewska, Steph Hughes-Fitt, Oran Kelly, Isaac R. Galatzer-Levy, Meredith Ringel Morris, Allan Dafoe, Alison M. Snyder, Noah D. Good- man, Matthew Botvinick, and Shane Legg. Measuring progress toward AGI: A cognitive frame- wor...

  107. [115]

    Position: Levels of AGI for operationalizing progress on the path to AGI

    Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Alek- sandra Faust, Clement Farabet, and Shane Legg. Position: Levels of AGI for operationalizing progress on the path to AGI. InProceedings of the 41st International Conference on Machine...

  108. [116]

    Three observations

    Sam Altman. Three observations. Blog post on samaltman.com, February 2025. Dis- cusses key trends in the economics and scaling of AI. URL:https://blog.samaltman.com/ three-observations

  109. [117]

    Reward is enough.Artificial Intelligence, 299:103535, 2021

    David Silver, Satinder Singh, Doina Precup, and Richard S Sutton. Reward is enough.Artificial Intelligence, 299:103535, 2021

  110. [118]

    Viking, New York, 2019

    Stuart Russell.Human Compatible: Artificial Intelligence and the Problem of Control. Viking, New York, 2019

  111. [119]

    Artificial intelligence, values, and alignment.Minds and Machines, 30(3):411–437, 2020

    Iason Gabriel. Artificial intelligence, values, and alignment.Minds and Machines, 30(3):411–437, 2020. 20

  112. [120]

    An approach to technical AGI safety and security, 2025

    Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, Jonah Brown-Cohen, Lewis Ho, Neel Nanda, Raluca Ada Popa, Rishub Jain, Rory Greig, Samuel Albanie, Scott Emmons, Sebastian Farquhar, S´ ebastien Krier, Senthooran Rajamanoharan, So- phie Bri...

  113. [121]

    Sharkey, Jacob Pfau, and David Krueger

    Lauro Di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau, and David Krueger. Goal misgener- alization in deep reinforcement learning. InProceedings of the 39th International Conference on Machine Learning, volume 162 ofProceedings of Machine Learning Research, pages 12004–1201...

  114. [122]

    Risks from learned optimization in advanced machine learning systems

    Evan Hubinger et al. Risks from learned optimization in advanced machine learning systems. arXiv preprint arXiv:1906.01820, 2019. 21

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.