Pith. sign in

REVIEW 3 major objections 5 minor 64 references

Open-ended AI needs frame-changing operations, not just better search inside fixed vocabularies and verifiers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 02:08 UTC pith:DFH267DB

load-bearing objection Useful conceptual map of open-ended AI (vocabulary + verifier gaps, L0–L3 ladder) that reorganizes known systems cleanly; the §2.1 generative criteria cannot fully operationalize the verifier gap they diagnose. the 3 major comments →

arxiv 2607.09560 v1 pith:DFH267DB submitted 2026-07-10 cs.AI cs.LG

Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

classification cs.AI cs.LG
keywords open-ended AIvocabulary gapverifier gaprepresentational primitivescognitive discrepancy reductioninnovation autonomyconcept inventionself-extending evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Current AI systems excel at solving problems whose concepts, search space, and success criteria are fixed in advance, yet genuine open-ended innovation requires inventing new representational primitives that change the space being searched. The paper names two structural barriers: the vocabulary gap—the hard problem of creating and stabilizing new concepts rather than recombining inherited ones—and the verifier gap—the hard problem of judging a new primitive when its value may only appear after later reuse. It unifies both under intelligence as cognitive discrepancy reduction: a sequence of transformations that shrink the mismatch between current representation and target, some of which stay inside a fixed frame and some of which rewrite the frame. From that view it defines a ladder of innovation autonomy and sketches what would have to change—objectives that reward useful representational revision, persistent stores for invented primitives, and evaluators that can extend themselves.

Core claim

The distance between present AI and open-ended intelligence is not mainly scale or scaffolding but two gaps: inventing and stabilizing new representational primitives (the vocabulary gap) and judging those primitives when current evaluators cannot yet see their full value (the verifier gap). Closing them requires generative transformations that modify the representational and evaluative frame itself, not only stronger search inside a human-supplied frame.

What carries the argument

Cognitive discrepancy reduction: intelligence as a trajectory of transformations that reduce mismatch between current representational state and a target, distinguishing intra-space operations from generative ones that alter vocabulary and evaluators; operationalized by amortized compression plus feasibility extension for primitives and by a four-level ladder of innovation autonomy.

Load-bearing premise

That a new primitive is worth keeping when it shortens description length across a family of tasks and makes previously unreachable tasks feasible under a fixed search budget—and that those tests plus a short list of cognitive priors are enough to decide value in open-ended domains.

What would settle it

Build or identify a system that invents a reusable primitive outside its training vocabulary, retains it after an ablation-style usefulness test, and revises its own success criteria so that later tasks become solvable; if no such system appears even under the paper’s proposed objectives and memory, or if systems that pass the compression tests still fail to produce lasting open-ended novelty, the gaps and ladder lose force.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that current AI systems excel at search and recombination inside fixed representational frames (vocabulary, admissible solutions, success criteria) but lack open-ended innovation. It characterizes the distance to open-ended intelligence by two gaps: the vocabulary gap (inventing and stabilizing new primitives rather than recombining given ones) and the verifier gap (judging a primitive whose payoff may appear only after future reuse or after the evaluation frame itself has changed). Both are interpreted under a unified schema of intelligence as trajectory-level cognitive discrepancy reduction (Eq. 1), distinguishing intra-space transformations from generative, frame-changing ones. A four-level ladder of innovation autonomy (L0–L3; Table 2, Fig. 1) is proposed, with L3 currently empty. Remedies include objectives that reward representational change, data that expose invention trajectories, persistent primitive stores, surrogate verifiers, and self-extending evaluation.

Significance. If the diagnosis holds, the paper supplies a useful organizing vocabulary for open-ended AI that unifies MDL/AIT, free-energy, structure-mapping, curiosity, and RL as restricted cases of discrepancy reduction (Table 1) and cleanly separates search autonomy from vocabulary and verifier autonomy. The ladder and the two-gap framing give a concrete way to locate systems such as FunSearch, DreamCoder, AI Scientist, and Gödel-machine variants, and the architectural directions in §5 are actionable research targets. The contribution is conceptual rather than empirical; its value lies in clarifying what current benchmarks and agent loops do not test and in making the missing operations (primitive invention, stabilization, and self-extending evaluation) explicit design goals.

major comments (3)
  1. §2.1 (amortized compression and feasibility extension under fixed F and B) is presented as the operational test of a generative primitive, yet §2.2 defines the deep verifier gap precisely as the case in which future tasks and standards of success are not yet expressible in the current frame. The two conditions therefore cannot be evaluated at invention time for the hardest cases the paper cares about. §5’s surrogate verifiers and cognitive priors (Table 1) are offered as provisional substitutes, but their calibration against delayed real outcomes is left unspecified. This undercuts the claim that the ladder and remedies operationally close the verifier gap; either the generative criteria must be relaxed or an explicit provisional retention/revision rule must be stated.
  2. The assertion that L3 is empty (Table 2, Fig. 1, §4) is largely definitional given the current criteria. Systems such as Eurisko, PowerPlay, Darwin/Red-Queen Gödel machines, and DreamCoder already perform limited self-modification of heuristics, libraries, or evaluators. Without a sharper, non-circular demarcation (e.g., an ablation-based test of whether a primitive remains useful after the original task family is removed, or whether the evaluator itself was invented rather than designer-supplied), the placement of all existing systems below L3 remains an assertion rather than a demonstrated empirical fact.
  3. §5 proposes objectives that reward useful representational change and data that expose invention trajectories, but supplies no concrete loss, curriculum construction method, or evaluation protocol that would let a reader test whether a system has narrowed either gap. Without even a toy experimental sketch (e.g., a controlled library-learning or concept-invention setting with an ablation of the proposed memory/consolidation loop), the remedies remain programmatic and the central claim that these changes are necessary and sufficient stays untested.
minor comments (5)
  1. Abstract and §1 contain several missing spaces (“intelligencethroughtwogaps”, “wheretherepresentation”, “intra-spacetransformation”). A careful copy-edit pass is needed.
  2. Eq. (1) and the local objective (2) are schematic; clarifying that they are not claimed to be computable global optimizers (already hinted in the text) would reduce possible misreading.
  3. Figure 1 places many systems along the bottom edge; a short caption note on the precise criteria used for each placement would improve reproducibility of the map.
  4. §6.3–6.4 cite Eurisko’s parasitic heuristics and Lenat & Brown’s later reflection; a one-sentence cross-reference back to the verifier-gap discussion would tighten the related-work connection.
  5. The safety paragraph in §7 is appropriately cautious but brief; a pointer to existing work on objective drift or self-modifying systems would strengthen it without expanding scope.

Circularity Check

0 steps flagged

No significant circularity: conceptual taxonomy and interpretive unification, not a derivation that reduces predictions to fitted inputs or self-citation chains.

full rationale

This is a position/conceptual paper. It defines two gaps, a discrepancy-reduction schema (Eq. 1–2), generative criteria (amortized compression and feasibility extension under budget B in §2.1), a four-level ladder (Table 2 / Fig. 1), and architectural directions. None of these steps fit parameters to data and then call the fit a prediction; there are no uniqueness theorems imported from the same authors; and the reference list does not load-bear on self-citations that themselves encode the target claim. Mapping MDL, free energy, SMT, Bayesian inference, and curiosity onto the schema (Table 1) is an interpretive specialization of D, T, and Ω chosen by the authors—standard unification rhetoric, not definitional equivalence that forces an empirical result. Declaring L3 empty follows from their own high bar plus a survey of existing systems; that is definitional taxonomy plus placement, not a circular derivation. The skeptic’s concern that the §2.1 criteria cannot be evaluated when the verifier gap is deepest is a completeness/operational-force objection, not circularity under the stated patterns. Score 0 is therefore the correct honest finding.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 4 invented entities

As a conceptual position paper the work rests on interpretive axioms rather than fitted constants. The central schema (intelligence as trajectory-level discrepancy reduction) and the two formal conditions for generative primitives are introduced by the authors; background theories (MDL, free energy, structure-mapping) are taken as standard. No numerical free parameters appear.

axioms (4)
  • ad hoc to paper Intelligence can be usefully formalized as trajectory-level minimization of cognitive discrepancy D_t(R_t, G_t) via a sequence of transformations T_t (Eq. 1).
    Introduced in §3.1 as the unifying schema; not derived from a prior theorem.
  • ad hoc to paper A primitive π is generative for task family F and budget B when it satisfies both amortized compression and feasibility extension (Section 2.1).
    Operational definition proposed by the authors; used to ground the vocabulary gap.
  • domain assumption Established frameworks (MDL/AIT, free energy, structure-mapping, curiosity) are restricted special cases of the same discrepancy-reduction dynamics.
    Stated in §3.2–3.3 and Table 1; interpretive claim that organizes the literature.
  • domain assumption Useful cognition is regularized by cognitive priors such as simplicity, systematicity, and learnable progress (Ω term in Eq. 2).
    Drawn from the cited cognitive-science and AIT literature and adopted without new derivation.
invented entities (4)
  • vocabulary gap no independent evidence
    purpose: Name the difficulty of inventing and stabilizing new representational primitives rather than recombining supplied ones.
    Core diagnostic construct of the paper; no independent empirical measure supplied outside the authors’ framing.
  • verifier gap no independent evidence
    purpose: Name the difficulty of judging a new primitive when its value may appear only after future reuse or after the evaluator itself changes.
    Second core diagnostic construct; likewise defined within the paper’s schema.
  • ladder of innovation autonomy (L0–L3) no independent evidence
    purpose: Rank systems by vocabulary autonomy, verifier autonomy, and search pattern; locate current systems and mark L3 empty.
    New taxonomy introduced in §4 and Table 2 / Figure 1.
  • generative (frame-changing) transformations no independent evidence
    purpose: Contrast with intra-space transformations that leave the representational frame fixed.
    Terminological distinction central to the argument in §3.3 and §4.

pith-pipeline@v1.1.0-grok45 · 24931 in / 2842 out tokens · 41531 ms · 2026-07-13T02:08:04.376371+00:00 · methodology

0 comments
read the original abstract

Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which the model operates, including its conceptual vocabulary, the space of admissible solutions it can search, and the criteria by which success is evaluated, is typically fixed and supplied in advance. This paper argues that building stronger intelligent systems capable of open-ended innovation requires additional classes of operations: the creation, stabilization, and reuse of new representational primitives, which alter the space being searched rather than simply searching within it. We characterize the distance between current AI systems and genuinely open-ended intelligence through two gaps. The first is the vocabulary gap, the difficulty of inventing and stabilizing new representational primitives rather than merely recombining existing ones. The second is the verifier gap, the difficulty of judging the value of a new primitive when its full payoff may be visible only after future reuse. We interpret both gaps through a unified framework of intelligence as cognitive discrepancy reduction. By viewing intelligent behaviors as a sequence of cognitive transformations, we distinguish intra-space transformations which operate within a fixed representational frame, from generative transformations which may modify the frame itself. On this basis, we propose a ladder of innovation autonomy and outline several directions for advancing open-ended AI, including objectives that reward useful representational change, persistent memory architectures for invented primitives, and adaptive verification mechanisms capable of evolving alongside the representations they evaluate.

Figures

Figures reproduced from arXiv: 2607.09560 by Haiqian Yang, Yuan Cao.

Figure 1
Figure 1. Figure 1: Axes of open-ended intelligence. Systems can be positioned by Axis X (vocabulary autonomy, from given to invented) and Axis Y (verifier autonomy, from fixed to self-extending). L0 and L1 remain near the bottom-left corner, differing mainly in search autonomy, L2 moves rightward by introducing limited autonomy over primitives, L3 occupies the upper-right region, corresponding to truly open-ended intelligenc… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

64 extracted references · 5 canonical work pages

  1. [1]

    Brian Arthur.The Nature of Technology: What It Is and How It Evolves

    W. Brian Arthur.The Nature of Technology: What It Is and How It Evolves. Free Press, New York, 2009. ISBN 9781439165782

  2. [2]

    Artiles, Martin Weiss, Levin Brinkmann, Iyad Rahwan, Bernhard Schölkopf, Christopher Pal, Hugo Larochelle, Anirudh Goyal, and Nasim Rahaman

    Alejandro H. Artiles, Martin Weiss, Levin Brinkmann, Iyad Rahwan, Bernhard Schölkopf, Christopher Pal, Hugo Larochelle, Anirudh Goyal, and Nasim Rahaman. The alien space of science: Sampling coherent but cognitively unavailable research directions, 2026. URL https://arxiv.org/abs/2603.01092

  3. [3]

    Nova: Fundamental limits of knowledge discovery through ai, 2026

    Salman Avestimehr, Ken Duffy, and Muriel Médard. Nova: Fundamental limits of knowledge discovery through ai, 2026. URLhttps://arxiv.org/abs/2605.15219

  4. [4]

    Honglin Bao, Siyang Wu, Xiao Liu, Sida Li, Shiyun Cao, and James A. Evans. Contemporary ai lacks the imagination to diverge or negate in science, 2026. URLhttps://arxiv.org/abs/ 2606.08251

  5. [5]

    Barsalou and Katja Wiemer-Hastings

    Lawrence W. Barsalou and Katja Wiemer-Hastings. Situating abstract concepts. In Diane Pecher and Rolf Zwaan, editors,Grounding cognition: The role of perception and action in memory, language, and thought, pages 129–163. Cambridge University Press, New York, 2005

  6. [6]

    Creativityandartificialintelligence.Artificial Intelligence, 103(1):347–356,

    MargaretA.Boden. Creativityandartificialintelligence.Artificial Intelligence, 103(1):347–356,

  7. [7]

    doi: https://doi.org/10.1016/S0004-3702(98)00055-1

    ISSN 0004-3702. doi: https://doi.org/10.1016/S0004-3702(98)00055-1. URLhttps:// www.sciencedirect.com/science/article/pii/S0004370298000551. Artificial Intelligence 40 years later. 20

  8. [8]

    Olausson, Catherine Wong, Gabriel Grand, Joshua B

    Matthew Bowers, Theo X. Olausson, Catherine Wong, Gabriel Grand, Joshua B. Tenenbaum, Kevin Ellis, and Armando Solar-Lezama. Top-down synthesis for library learning.Proceedings of the ACM on Programming Languages, 7(POPL), 2023

  9. [9]

    Measuring the gap between human and llm research ideas, 2026

    Ziyu Chen, Yilun Zhao, and Arman Cohan. Measuring the gap between human and llm research ideas, 2026. URLhttps://arxiv.org/abs/2607.01233

  10. [10]

    On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019

    François Chollet. On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019

  11. [11]

    Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and Brain Sciences, 36(3):181–204, 2013

    Andy Clark. Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and Brain Sciences, 36(3):181–204, 2013

  12. [12]

    Ai-gas: Ai-generating algorithms, an alternate paradigm for producing general artificial intelligence.arXiv preprint arXiv:1905.10985, 2019

    Jeff Clune. Ai-gas: Ai-generating algorithms, an alternate paradigm for producing general artificial intelligence.arXiv preprint arXiv:1905.10985, 2019

  13. [13]

    Tenenbaum

    Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sablé-Meyer, Luc Cary, Lucas Morales, Luke Hewitt, Armando Solar-Lezama, and Joshua B. Tenenbaum. Dreamcoder: Bootstrapping inductive program synthesis with wake-sleep library learning. InProceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, 2021

  14. [14]

    Forbus, and Dedre Gentner

    Brian Falkenhainer, Kenneth D. Forbus, and Dedre Gentner. The structure-mapping engine: Algorithm and examples.Artificial Intelligence, 41(1):1–63, 1989

  15. [15]

    The usefulness of useless knowledge.Harper’s Magazine, 1939

    Abraham Flexner. The usefulness of useless knowledge.Harper’s Magazine, 1939

  16. [16]

    The free-energy principle: a unified brain theory?Nature Reviews Neuroscience, 11:127–138, 2010

    Karl Friston. The free-energy principle: a unified brain theory?Nature Reviews Neuroscience, 11:127–138, 2010

  17. [17]

    MIT Press, 2000

    Peter Gärdenfors.Conceptual Spaces: The Geometry of Thought. MIT Press, 2000

  18. [18]

    Leibo, Allan Dafoe, Marcus Hutter, Thore Graepel, and Shane Legg

    Tim Genewein, Matija Franklin, Alexander Lerchner, Laurent Orseau, Samuel Albanie, Adam Bales, Cole Wyeth, Stephanie Chan, Iason Gabriel, Joel Z. Leibo, Allan Dafoe, Marcus Hutter, Thore Graepel, and Shane Legg. From agi to asi, 2026. URLhttps://arxiv.org/abs/2606. 12683

  19. [19]

    Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983

    Dedre Gentner. Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983

  20. [20]

    Dedre Gentner and Arthur B. Markman. Structure mapping in analogy and similarity.Amer- ican Psychologist, 52(1):45–56, 1997

  21. [21]

    MIT press, 2001

    Dedre Gentner, Keith J Holyoak, and Boicho N Kokinov.The analogical mind: Perspectives from cognitive science. MIT press, 2001

  22. [22]

    Gick and Keith J

    Mary L. Gick and Keith J. Holyoak. Schema induction and analogical transfer.Cognitive Psychology, 15(1):1–38, 1983. ISSN 0010-0285. doi: https://doi.org/10.1016/0010-0285(83) 90002-6. URLhttps://www.sciencedirect.com/science/article/pii/0010028583900026

  23. [23]

    Towards an ai co-scientist

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025

  24. [24]

    Grünwald.The Minimum Description Length Principle

    Peter D. Grünwald.The Minimum Description Length Principle. MIT Press, 2007. 21

  25. [25]

    Basic Books, New York, 2013

    Douglas Hofstadter and Emmanuel Sander.Surfaces and Essences: Analogy as the Fuel and Fire of Thinking. Basic Books, New York, 2013

  26. [26]

    Carter, Xin Zhou, Matthew T

    Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Ryan Li, Yusuf Roohani, Lin Qiu, Shiyi Cao, Gavin Li, Junze Zhang, Di Yin, Rick Wierenga, Deniz Kavi, Sherry Liu, Tianwei She, Shruti Marwaha, Jennefer N. Carter, Xin Zhou, Matthew T. Wheeler, Jonathan A. Bernstein, Mengdi Wang, Peng He, Jingtian Zhou, Michael P. Snyder, Le Cong, Aviv Regev...

  27. [27]

    Position: Open-endedness is essential for ar- tificial superhuman intelligence

    Edward Hughes, Michael D Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, and Tim Rocktäschel. Position: Open-endedness is essential for ar- tificial superhuman intelligence. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,Proceed- in...

  28. [28]

    Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, LorenzoSani, NiccolòAlbertoEliaVenanzi, AmbroiseOdonnat, ZeyuCao, BillMarino, Xinchi Qiu, and Nicholas D

    Alex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, LorenzoSani, NiccolòAlbertoEliaVenanzi, AmbroiseOdonnat, ZeyuCao, BillMarino, Xinchi Qiu, and Nicholas D. Lane. The red queen gödel machine: Co-evolving agents and their evaluators, 2026. URLhttps://arxiv.org/abs/2606.26294

  29. [29]

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron- neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reim...

  30. [30]

    The Cambridge Edition of the Works of Immanuel Kant

    Immanuel Kant.Critique of Pure Reason. The Cambridge Edition of the Works of Immanuel Kant. Cambridge University Press, Cambridge, 1998. Originally published 1781 (A) / 1787 (B)

  31. [31]

    Kolmogorov

    Andrei N. Kolmogorov. Three approaches to the quantitative definition of information.Prob- lems of Information Transmission, 1(1):1–7, 1965

  32. [32]

    UniversityofChicagoPress, Chicago, 1980

    GeorgeLakoffandMarkJohnson.Metaphors We Live By. UniversityofChicagoPress, Chicago, 1980

  33. [33]

    Universal intelligence: A definition of machine intelligence

    Shane Legg and Marcus Hutter. Universal intelligence: A definition of machine intelligence. Minds and Machines, 17:391–444, 2007

  34. [34]

    Douglas B. Lenat. AM: An Artificial Intelligence Approach to Discovery in Mathematics as Heuristic Search. 7 1976

  35. [35]

    Douglas B. Lenat. Eurisko: A program that learns new heuristics and domain concepts.Arti- ficial Intelligence, 21(1–2):61–98, 1983. 22

  36. [36]

    Lenat and John Seely Brown

    Douglas B. Lenat and John Seely Brown. Why am and eurisko appear to work.Artificial Intelligence, 23(3):269–294, 1984

  37. [37]

    The abstraction fallacy: Why ai can simulate but not instantiate con- sciousness.PhilPapers, 2026

    Alexander Lerchner. The abstraction fallacy: Why ai can simulate but not instantiate con- sciousness.PhilPapers, 2026

  38. [38]

    Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021. URL https://arxiv.org/abs/2005.11401

  39. [39]

    Springer, 3rd edition, 2008

    Ming Li and Paul Vitányi.An Introduction to Kolmogorov Complexity and Its Applications. Springer, 3rd edition, 2008

  40. [40]

    Pan, Alexander Du, Kurt Keutzer, Alvin Cheung, Alexandros G

    Shu Liu, Shubham Agarwal, Monishwaran Maheswaran, Mert Cemri, Zhifei Li, Qiuyang Mang, Ashwin Naren, Ethan Boneh, Audrey Cheng, Melissa Z. Pan, Alexander Du, Kurt Keutzer, Alvin Cheung, Alexandros G. Dimakis, Koushik Sen, Matei Zaharia, and Ion Stoica. Evox: Meta-evolution for automated discovery, 2026. URLhttps://arxiv.org/abs/2602.23413

  41. [41]

    The ai scientist: Towards fully automated open-ended scientific discovery.arXiv preprint arXiv:2408.06292, 2024

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery.arXiv preprint arXiv:2408.06292, 2024

  42. [42]

    Abstraction and analogy-making in artificial intelligence.Annals of the New York Academy of Sciences, 1505(1):79–101, 2021

    Melanie Mitchell. Abstraction and analogy-making in artificial intelligence.Annals of the New York Academy of Sciences, 1505(1):79–101, 2021. ISSN 1749-6632. doi: 10.1111/nyas.14619. URLhttp://dx.doi.org/10.1111/nyas.14619

  43. [43]

    A compositional framework for open-ended intelli- gence, 2026

    Ida Momennejad and Roberta Raileanu. A compositional framework for open-ended intelli- gence, 2026. URLhttps://arxiv.org/abs/2606.15386

  44. [44]

    Alphaevolve: A coding agent for scientific and algorithmic dis- covery.arXiv preprint arXiv:2506.13131, 2025

    Alexander Novikov, Ngân V˜ u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, et al. Alphaevolve: A coding agent for scientific and algorithmic dis- covery.arXiv preprint arXiv:2506.13131, 2025

  45. [45]

    AutoLabs: Cognitive multi- agentsystemswithself-correctionforautonomouschemicalexperimentation.Scientific Reports, 16(1):19554, June 2026

    Gihan Panapitiya, Emily Saldanha, Heather Job, and Olivia Hess. AutoLabs: Cognitive multi- agentsystemswithself-correctionforautonomouschemicalexperimentation.Scientific Reports, 16(1):19554, June 2026. doi: 10.1038/s41598-026-45593-z. URLhttps://doi.org/10.1038/ s41598-026-45593-z

  46. [46]

    Rajesh P. N. Rao and Dana H. Ballard. Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects.Nature Neuroscience, 2(1):79–87, 1999

  47. [47]

    Modeling by shortest data description.Automatica, 14(5):465–471, 1978

    Jorma Rissanen. Modeling by shortest data description.Automatica, 14(5):465–471, 1978

  48. [48]

    Math- ematical discoveries from program search with large language models.Nature, 625:468–475, 2024

    Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, et al. Math- ematical discoveries from program search with large language models.Nature, 625:468–475, 2024

  49. [49]

    Varshney

    Samuel Schapiro, Sumuk Shashidhar, Alexi Gladstone, Jonah Black, Royce Moon, Dilek Hakkani-Tur, and Lav R. Varshney. Combinatorial creativity: A new frontier in generalization abilities, 2026. URLhttps://arxiv.org/abs/2509.21043. 23

  50. [50]

    Goedel machines: Self-referential universal problem solvers making provably optimal self-improvements, 2006

    Juergen Schmidhuber. Goedel machines: Self-referential universal problem solvers making provably optimal self-improvements, 2006. URLhttps://arxiv.org/abs/cs/0309048

  51. [51]

    Formal theory of creativity, fun, and intrinsic motivation (1990–2010)

    Jürgen Schmidhuber. Formal theory of creativity, fun, and intrinsic motivation (1990–2010). IEEE Transactions on Autonomous Mental Development, 2(3):230–247, 2010

  52. [52]

    Powerplay: Traininganincreasinglygeneralproblemsolverbycontinually searching for the simplest still unsolvable problem.Frontiers in Psychology, 4:313, 2013

    JürgenSchmidhuber. Powerplay: Traininganincreasinglygeneralproblemsolverbycontinually searching for the simplest still unsolvable problem.Frontiers in Psychology, 4:313, 2013

  53. [53]

    A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018

    David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018. doi: 10.1126...

  54. [54]

    Solomonoff

    Ray J. Solomonoff. A formal theory of inductive inference, parts I and II.Information and Control, 7(1–2):1–22, 224–254, 1964

  55. [55]

    Sutton and Andrew G

    Richard S. Sutton and Andrew G. Barto.Reinforcement Learning: An Introduction. MIT Press, 2nd edition, 2018

  56. [56]

    Ai research agents narrow scientific exploration, 2026

    Yixuan Tang and Yi Yang. Ai research agents narrow scientific exploration, 2026. URL https://arxiv.org/abs/2605.27905

  57. [57]

    Tenenbaum, Charles Kemp, Thomas L

    Joshua B. Tenenbaum, Charles Kemp, Thomas L. Griffiths, and Noah D. Goodman. How to grow a mind: Statistics, structure, and abstraction.Science, 331(6022):1279–1285, 2011

  58. [58]

    Yu, and Jianfeng Gao

    Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, Lifang He, Qingsong Wen, Manling Li, Cong Lu, Shuai Li, Pengtao Xie, Yixuan Yuan, Rui Meng, Lei Xing, Lichao Sun, Caiming Xiong, Philip S. Yu, and Jianfeng Gao. Autoresearch ai: Towards ai-powered research automation for scientific d...

  59. [59]

    Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2023

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2023. URLhttps://openreview.net/ forum?id=ehfRiF0R3a

  60. [60]

    Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O. Stanley. Paired open-ended trailblazer: Endlessly generating increasingly complex and diverse learning environments and their solu- tions.arXiv preprint arXiv:1901.01753, 2019

  61. [61]

    Rui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi, Yulun Li, Jeff Clune, and Kenneth O. Stanley. Enhanced poet: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions.arXiv preprint arXiv:2003.08536, 2020

  62. [62]

    Huxley-gödel machine: Human-level coding agent development by an approximation of the optimal self-improving machine, 2025

    Wenyi Wang, Piotr Piękos, Li Nanbo, Firas Laakom, Yimeng Chen, Mateusz Ostaszewski, Mingchen Zhuge, and Jürgen Schmidhuber. Huxley-gödel machine: Human-level coding agent development by an approximation of the optimal self-improving machine, 2025. URLhttps: //arxiv.org/abs/2510.21614

  63. [63]

    Llms can’t jump, January 2026

    Tom Zahavy. Llms can’t jump, January 2026. URLhttps://philsci-archive.pitt.edu/ 28024/. 24

  64. [64]

    Darwin godel machine: Open-ended evolution of self-improving agents, 2026

    Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune. Darwin godel machine: Open-ended evolution of self-improving agents, 2026. URLhttps://arxiv.org/abs/2505. 22954. 25