REVIEW 3 major objections 5 minor 64 references
Open-ended AI needs frame-changing operations, not just better search inside fixed vocabularies and verifiers.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 02:08 UTC pith:DFH267DB
load-bearing objection Useful conceptual map of open-ended AI (vocabulary + verifier gaps, L0–L3 ladder) that reorganizes known systems cleanly; the §2.1 generative criteria cannot fully operationalize the verifier gap they diagnose. the 3 major comments →
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The distance between present AI and open-ended intelligence is not mainly scale or scaffolding but two gaps: inventing and stabilizing new representational primitives (the vocabulary gap) and judging those primitives when current evaluators cannot yet see their full value (the verifier gap). Closing them requires generative transformations that modify the representational and evaluative frame itself, not only stronger search inside a human-supplied frame.
What carries the argument
Cognitive discrepancy reduction: intelligence as a trajectory of transformations that reduce mismatch between current representational state and a target, distinguishing intra-space operations from generative ones that alter vocabulary and evaluators; operationalized by amortized compression plus feasibility extension for primitives and by a four-level ladder of innovation autonomy.
Load-bearing premise
That a new primitive is worth keeping when it shortens description length across a family of tasks and makes previously unreachable tasks feasible under a fixed search budget—and that those tests plus a short list of cognitive priors are enough to decide value in open-ended domains.
What would settle it
Build or identify a system that invents a reusable primitive outside its training vocabulary, retains it after an ablation-style usefulness test, and revises its own success criteria so that later tasks become solvable; if no such system appears even under the paper’s proposed objectives and memory, or if systems that pass the compression tests still fail to produce lasting open-ended novelty, the gaps and ladder lose force.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that current AI systems excel at search and recombination inside fixed representational frames (vocabulary, admissible solutions, success criteria) but lack open-ended innovation. It characterizes the distance to open-ended intelligence by two gaps: the vocabulary gap (inventing and stabilizing new primitives rather than recombining given ones) and the verifier gap (judging a primitive whose payoff may appear only after future reuse or after the evaluation frame itself has changed). Both are interpreted under a unified schema of intelligence as trajectory-level cognitive discrepancy reduction (Eq. 1), distinguishing intra-space transformations from generative, frame-changing ones. A four-level ladder of innovation autonomy (L0–L3; Table 2, Fig. 1) is proposed, with L3 currently empty. Remedies include objectives that reward representational change, data that expose invention trajectories, persistent primitive stores, surrogate verifiers, and self-extending evaluation.
Significance. If the diagnosis holds, the paper supplies a useful organizing vocabulary for open-ended AI that unifies MDL/AIT, free-energy, structure-mapping, curiosity, and RL as restricted cases of discrepancy reduction (Table 1) and cleanly separates search autonomy from vocabulary and verifier autonomy. The ladder and the two-gap framing give a concrete way to locate systems such as FunSearch, DreamCoder, AI Scientist, and Gödel-machine variants, and the architectural directions in §5 are actionable research targets. The contribution is conceptual rather than empirical; its value lies in clarifying what current benchmarks and agent loops do not test and in making the missing operations (primitive invention, stabilization, and self-extending evaluation) explicit design goals.
major comments (3)
- §2.1 (amortized compression and feasibility extension under fixed F and B) is presented as the operational test of a generative primitive, yet §2.2 defines the deep verifier gap precisely as the case in which future tasks and standards of success are not yet expressible in the current frame. The two conditions therefore cannot be evaluated at invention time for the hardest cases the paper cares about. §5’s surrogate verifiers and cognitive priors (Table 1) are offered as provisional substitutes, but their calibration against delayed real outcomes is left unspecified. This undercuts the claim that the ladder and remedies operationally close the verifier gap; either the generative criteria must be relaxed or an explicit provisional retention/revision rule must be stated.
- The assertion that L3 is empty (Table 2, Fig. 1, §4) is largely definitional given the current criteria. Systems such as Eurisko, PowerPlay, Darwin/Red-Queen Gödel machines, and DreamCoder already perform limited self-modification of heuristics, libraries, or evaluators. Without a sharper, non-circular demarcation (e.g., an ablation-based test of whether a primitive remains useful after the original task family is removed, or whether the evaluator itself was invented rather than designer-supplied), the placement of all existing systems below L3 remains an assertion rather than a demonstrated empirical fact.
- §5 proposes objectives that reward useful representational change and data that expose invention trajectories, but supplies no concrete loss, curriculum construction method, or evaluation protocol that would let a reader test whether a system has narrowed either gap. Without even a toy experimental sketch (e.g., a controlled library-learning or concept-invention setting with an ablation of the proposed memory/consolidation loop), the remedies remain programmatic and the central claim that these changes are necessary and sufficient stays untested.
minor comments (5)
- Abstract and §1 contain several missing spaces (“intelligencethroughtwogaps”, “wheretherepresentation”, “intra-spacetransformation”). A careful copy-edit pass is needed.
- Eq. (1) and the local objective (2) are schematic; clarifying that they are not claimed to be computable global optimizers (already hinted in the text) would reduce possible misreading.
- Figure 1 places many systems along the bottom edge; a short caption note on the precise criteria used for each placement would improve reproducibility of the map.
- §6.3–6.4 cite Eurisko’s parasitic heuristics and Lenat & Brown’s later reflection; a one-sentence cross-reference back to the verifier-gap discussion would tighten the related-work connection.
- The safety paragraph in §7 is appropriately cautious but brief; a pointer to existing work on objective drift or self-modifying systems would strengthen it without expanding scope.
Circularity Check
No significant circularity: conceptual taxonomy and interpretive unification, not a derivation that reduces predictions to fitted inputs or self-citation chains.
full rationale
This is a position/conceptual paper. It defines two gaps, a discrepancy-reduction schema (Eq. 1–2), generative criteria (amortized compression and feasibility extension under budget B in §2.1), a four-level ladder (Table 2 / Fig. 1), and architectural directions. None of these steps fit parameters to data and then call the fit a prediction; there are no uniqueness theorems imported from the same authors; and the reference list does not load-bear on self-citations that themselves encode the target claim. Mapping MDL, free energy, SMT, Bayesian inference, and curiosity onto the schema (Table 1) is an interpretive specialization of D, T, and Ω chosen by the authors—standard unification rhetoric, not definitional equivalence that forces an empirical result. Declaring L3 empty follows from their own high bar plus a survey of existing systems; that is definitional taxonomy plus placement, not a circular derivation. The skeptic’s concern that the §2.1 criteria cannot be evaluated when the verifier gap is deepest is a completeness/operational-force objection, not circularity under the stated patterns. Score 0 is therefore the correct honest finding.
Axiom & Free-Parameter Ledger
axioms (4)
- ad hoc to paper Intelligence can be usefully formalized as trajectory-level minimization of cognitive discrepancy D_t(R_t, G_t) via a sequence of transformations T_t (Eq. 1).
- ad hoc to paper A primitive π is generative for task family F and budget B when it satisfies both amortized compression and feasibility extension (Section 2.1).
- domain assumption Established frameworks (MDL/AIT, free energy, structure-mapping, curiosity) are restricted special cases of the same discrepancy-reduction dynamics.
- domain assumption Useful cognition is regularized by cognitive priors such as simplicity, systematicity, and learnable progress (Ω term in Eq. 2).
invented entities (4)
-
vocabulary gap
no independent evidence
-
verifier gap
no independent evidence
-
ladder of innovation autonomy (L0–L3)
no independent evidence
-
generative (frame-changing) transformations
no independent evidence
read the original abstract
Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which the model operates, including its conceptual vocabulary, the space of admissible solutions it can search, and the criteria by which success is evaluated, is typically fixed and supplied in advance. This paper argues that building stronger intelligent systems capable of open-ended innovation requires additional classes of operations: the creation, stabilization, and reuse of new representational primitives, which alter the space being searched rather than simply searching within it. We characterize the distance between current AI systems and genuinely open-ended intelligence through two gaps. The first is the vocabulary gap, the difficulty of inventing and stabilizing new representational primitives rather than merely recombining existing ones. The second is the verifier gap, the difficulty of judging the value of a new primitive when its full payoff may be visible only after future reuse. We interpret both gaps through a unified framework of intelligence as cognitive discrepancy reduction. By viewing intelligent behaviors as a sequence of cognitive transformations, we distinguish intra-space transformations which operate within a fixed representational frame, from generative transformations which may modify the frame itself. On this basis, we propose a ladder of innovation autonomy and outline several directions for advancing open-ended AI, including objectives that reward useful representational change, persistent memory architectures for invented primitives, and adaptive verification mechanisms capable of evolving alongside the representations they evaluate.
Figures
Reference graph
Works this paper leans on
-
[1]
Brian Arthur.The Nature of Technology: What It Is and How It Evolves
W. Brian Arthur.The Nature of Technology: What It Is and How It Evolves. Free Press, New York, 2009. ISBN 9781439165782
2009
-
[2]
Alejandro H. Artiles, Martin Weiss, Levin Brinkmann, Iyad Rahwan, Bernhard Schölkopf, Christopher Pal, Hugo Larochelle, Anirudh Goyal, and Nasim Rahaman. The alien space of science: Sampling coherent but cognitively unavailable research directions, 2026. URL https://arxiv.org/abs/2603.01092
Pith/arXiv arXiv 2026
-
[3]
Nova: Fundamental limits of knowledge discovery through ai, 2026
Salman Avestimehr, Ken Duffy, and Muriel Médard. Nova: Fundamental limits of knowledge discovery through ai, 2026. URLhttps://arxiv.org/abs/2605.15219
Pith/arXiv arXiv 2026
-
[4]
Honglin Bao, Siyang Wu, Xiao Liu, Sida Li, Shiyun Cao, and James A. Evans. Contemporary ai lacks the imagination to diverge or negate in science, 2026. URLhttps://arxiv.org/abs/ 2606.08251
Pith/arXiv arXiv 2026
-
[5]
Barsalou and Katja Wiemer-Hastings
Lawrence W. Barsalou and Katja Wiemer-Hastings. Situating abstract concepts. In Diane Pecher and Rolf Zwaan, editors,Grounding cognition: The role of perception and action in memory, language, and thought, pages 129–163. Cambridge University Press, New York, 2005
2005
-
[6]
Creativityandartificialintelligence.Artificial Intelligence, 103(1):347–356,
MargaretA.Boden. Creativityandartificialintelligence.Artificial Intelligence, 103(1):347–356,
-
[7]
doi: https://doi.org/10.1016/S0004-3702(98)00055-1
ISSN 0004-3702. doi: https://doi.org/10.1016/S0004-3702(98)00055-1. URLhttps:// www.sciencedirect.com/science/article/pii/S0004370298000551. Artificial Intelligence 40 years later. 20
-
[8]
Olausson, Catherine Wong, Gabriel Grand, Joshua B
Matthew Bowers, Theo X. Olausson, Catherine Wong, Gabriel Grand, Joshua B. Tenenbaum, Kevin Ellis, and Armando Solar-Lezama. Top-down synthesis for library learning.Proceedings of the ACM on Programming Languages, 7(POPL), 2023
2023
-
[9]
Measuring the gap between human and llm research ideas, 2026
Ziyu Chen, Yilun Zhao, and Arman Cohan. Measuring the gap between human and llm research ideas, 2026. URLhttps://arxiv.org/abs/2607.01233
Pith/arXiv arXiv 2026
-
[10]
On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019
François Chollet. On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019
Pith/arXiv arXiv 1911
-
[11]
Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and Brain Sciences, 36(3):181–204, 2013
Andy Clark. Whatever next? predictive brains, situated agents, and the future of cognitive science.Behavioral and Brain Sciences, 36(3):181–204, 2013
2013
-
[12]
Jeff Clune. Ai-gas: Ai-generating algorithms, an alternate paradigm for producing general artificial intelligence.arXiv preprint arXiv:1905.10985, 2019
Pith/arXiv arXiv 1905
-
[13]
Tenenbaum
Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sablé-Meyer, Luc Cary, Lucas Morales, Luke Hewitt, Armando Solar-Lezama, and Joshua B. Tenenbaum. Dreamcoder: Bootstrapping inductive program synthesis with wake-sleep library learning. InProceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, 2021
2021
-
[14]
Forbus, and Dedre Gentner
Brian Falkenhainer, Kenneth D. Forbus, and Dedre Gentner. The structure-mapping engine: Algorithm and examples.Artificial Intelligence, 41(1):1–63, 1989
1989
-
[15]
The usefulness of useless knowledge.Harper’s Magazine, 1939
Abraham Flexner. The usefulness of useless knowledge.Harper’s Magazine, 1939
1939
-
[16]
The free-energy principle: a unified brain theory?Nature Reviews Neuroscience, 11:127–138, 2010
Karl Friston. The free-energy principle: a unified brain theory?Nature Reviews Neuroscience, 11:127–138, 2010
2010
-
[17]
MIT Press, 2000
Peter Gärdenfors.Conceptual Spaces: The Geometry of Thought. MIT Press, 2000
2000
-
[18]
Leibo, Allan Dafoe, Marcus Hutter, Thore Graepel, and Shane Legg
Tim Genewein, Matija Franklin, Alexander Lerchner, Laurent Orseau, Samuel Albanie, Adam Bales, Cole Wyeth, Stephanie Chan, Iason Gabriel, Joel Z. Leibo, Allan Dafoe, Marcus Hutter, Thore Graepel, and Shane Legg. From agi to asi, 2026. URLhttps://arxiv.org/abs/2606. 12683
2026
-
[19]
Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983
Dedre Gentner. Structure-mapping: A theoretical framework for analogy.Cognitive Science, 7 (2):155–170, 1983
1983
-
[20]
Dedre Gentner and Arthur B. Markman. Structure mapping in analogy and similarity.Amer- ican Psychologist, 52(1):45–56, 1997
1997
-
[21]
MIT press, 2001
Dedre Gentner, Keith J Holyoak, and Boicho N Kokinov.The analogical mind: Perspectives from cognitive science. MIT press, 2001
2001
-
[22]
Mary L. Gick and Keith J. Holyoak. Schema induction and analogical transfer.Cognitive Psychology, 15(1):1–38, 1983. ISSN 0010-0285. doi: https://doi.org/10.1016/0010-0285(83) 90002-6. URLhttps://www.sciencedirect.com/science/article/pii/0010028583900026
-
[23]
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025
Pith/arXiv arXiv 2025
-
[24]
Grünwald.The Minimum Description Length Principle
Peter D. Grünwald.The Minimum Description Length Principle. MIT Press, 2007. 21
2007
-
[25]
Basic Books, New York, 2013
Douglas Hofstadter and Emmanuel Sander.Surfaces and Essences: Analogy as the Fuel and Fire of Thinking. Basic Books, New York, 2013
2013
-
[26]
Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Ryan Li, Yusuf Roohani, Lin Qiu, Shiyi Cao, Gavin Li, Junze Zhang, Di Yin, Rick Wierenga, Deniz Kavi, Sherry Liu, Tianwei She, Shruti Marwaha, Jennefer N. Carter, Xin Zhou, Matthew T. Wheeler, Jonathan A. Bernstein, Mengdi Wang, Peng He, Jingtian Zhou, Michael P. Snyder, Le Cong, Aviv Regev...
-
[27]
Position: Open-endedness is essential for ar- tificial superhuman intelligence
Edward Hughes, Michael D Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, and Tim Rocktäschel. Position: Open-endedness is essential for ar- tificial superhuman intelligence. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,Proceed- in...
2024
-
[28]
Alex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, LorenzoSani, NiccolòAlbertoEliaVenanzi, AmbroiseOdonnat, ZeyuCao, BillMarino, Xinchi Qiu, and Nicholas D. Lane. The red queen gödel machine: Co-evolving agents and their evaluators, 2026. URLhttps://arxiv.org/abs/2606.26294
Pith/arXiv arXiv 2026
-
[29]
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron- neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reim...
-
[30]
The Cambridge Edition of the Works of Immanuel Kant
Immanuel Kant.Critique of Pure Reason. The Cambridge Edition of the Works of Immanuel Kant. Cambridge University Press, Cambridge, 1998. Originally published 1781 (A) / 1787 (B)
1998
-
[31]
Kolmogorov
Andrei N. Kolmogorov. Three approaches to the quantitative definition of information.Prob- lems of Information Transmission, 1(1):1–7, 1965
1965
-
[32]
UniversityofChicagoPress, Chicago, 1980
GeorgeLakoffandMarkJohnson.Metaphors We Live By. UniversityofChicagoPress, Chicago, 1980
1980
-
[33]
Universal intelligence: A definition of machine intelligence
Shane Legg and Marcus Hutter. Universal intelligence: A definition of machine intelligence. Minds and Machines, 17:391–444, 2007
2007
-
[34]
Douglas B. Lenat. AM: An Artificial Intelligence Approach to Discovery in Mathematics as Heuristic Search. 7 1976
1976
-
[35]
Douglas B. Lenat. Eurisko: A program that learns new heuristics and domain concepts.Arti- ficial Intelligence, 21(1–2):61–98, 1983. 22
1983
-
[36]
Lenat and John Seely Brown
Douglas B. Lenat and John Seely Brown. Why am and eurisko appear to work.Artificial Intelligence, 23(3):269–294, 1984
1984
-
[37]
The abstraction fallacy: Why ai can simulate but not instantiate con- sciousness.PhilPapers, 2026
Alexander Lerchner. The abstraction fallacy: Why ai can simulate but not instantiate con- sciousness.PhilPapers, 2026
2026
-
[38]
Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021. URL https://arxiv.org/abs/2005.11401
Pith/arXiv arXiv 2021
-
[39]
Springer, 3rd edition, 2008
Ming Li and Paul Vitányi.An Introduction to Kolmogorov Complexity and Its Applications. Springer, 3rd edition, 2008
2008
-
[40]
Pan, Alexander Du, Kurt Keutzer, Alvin Cheung, Alexandros G
Shu Liu, Shubham Agarwal, Monishwaran Maheswaran, Mert Cemri, Zhifei Li, Qiuyang Mang, Ashwin Naren, Ethan Boneh, Audrey Cheng, Melissa Z. Pan, Alexander Du, Kurt Keutzer, Alvin Cheung, Alexandros G. Dimakis, Koushik Sen, Matei Zaharia, and Ion Stoica. Evox: Meta-evolution for automated discovery, 2026. URLhttps://arxiv.org/abs/2602.23413
arXiv 2026
-
[41]
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery.arXiv preprint arXiv:2408.06292, 2024
Pith/arXiv arXiv 2024
-
[42]
Melanie Mitchell. Abstraction and analogy-making in artificial intelligence.Annals of the New York Academy of Sciences, 1505(1):79–101, 2021. ISSN 1749-6632. doi: 10.1111/nyas.14619. URLhttp://dx.doi.org/10.1111/nyas.14619
-
[43]
A compositional framework for open-ended intelli- gence, 2026
Ida Momennejad and Roberta Raileanu. A compositional framework for open-ended intelli- gence, 2026. URLhttps://arxiv.org/abs/2606.15386
Pith/arXiv arXiv 2026
-
[44]
Alexander Novikov, Ngân V˜ u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, et al. Alphaevolve: A coding agent for scientific and algorithmic dis- covery.arXiv preprint arXiv:2506.13131, 2025
Pith/arXiv arXiv 2025
-
[45]
Gihan Panapitiya, Emily Saldanha, Heather Job, and Olivia Hess. AutoLabs: Cognitive multi- agentsystemswithself-correctionforautonomouschemicalexperimentation.Scientific Reports, 16(1):19554, June 2026. doi: 10.1038/s41598-026-45593-z. URLhttps://doi.org/10.1038/ s41598-026-45593-z
-
[46]
Rajesh P. N. Rao and Dana H. Ballard. Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects.Nature Neuroscience, 2(1):79–87, 1999
1999
-
[47]
Modeling by shortest data description.Automatica, 14(5):465–471, 1978
Jorma Rissanen. Modeling by shortest data description.Automatica, 14(5):465–471, 1978
1978
-
[48]
Math- ematical discoveries from program search with large language models.Nature, 625:468–475, 2024
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, et al. Math- ematical discoveries from program search with large language models.Nature, 625:468–475, 2024
2024
- [49]
-
[50]
Juergen Schmidhuber. Goedel machines: Self-referential universal problem solvers making provably optimal self-improvements, 2006. URLhttps://arxiv.org/abs/cs/0309048
Pith/arXiv arXiv 2006
-
[51]
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber. Formal theory of creativity, fun, and intrinsic motivation (1990–2010). IEEE Transactions on Autonomous Mental Development, 2(3):230–247, 2010
1990
-
[52]
Powerplay: Traininganincreasinglygeneralproblemsolverbycontinually searching for the simplest still unsolvable problem.Frontiers in Psychology, 4:313, 2013
JürgenSchmidhuber. Powerplay: Traininganincreasinglygeneralproblemsolverbycontinually searching for the simplest still unsolvable problem.Frontiers in Psychology, 4:313, 2013
2013
-
[53]
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018. doi: 10.1126...
2018
-
[54]
Solomonoff
Ray J. Solomonoff. A formal theory of inductive inference, parts I and II.Information and Control, 7(1–2):1–22, 224–254, 1964
1964
-
[55]
Sutton and Andrew G
Richard S. Sutton and Andrew G. Barto.Reinforcement Learning: An Introduction. MIT Press, 2nd edition, 2018
2018
-
[56]
Ai research agents narrow scientific exploration, 2026
Yixuan Tang and Yi Yang. Ai research agents narrow scientific exploration, 2026. URL https://arxiv.org/abs/2605.27905
Pith/arXiv arXiv 2026
-
[57]
Tenenbaum, Charles Kemp, Thomas L
Joshua B. Tenenbaum, Charles Kemp, Thomas L. Griffiths, and Noah D. Goodman. How to grow a mind: Statistics, structure, and abstraction.Science, 331(6022):1279–1285, 2011
2011
-
[58]
Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, Lifang He, Qingsong Wen, Manling Li, Cong Lu, Shuai Li, Pengtao Xie, Yixuan Yuan, Rui Meng, Lei Xing, Lichao Sun, Caiming Xiong, Philip S. Yu, and Jianfeng Gao. Autoresearch ai: Towards ai-powered research automation for scientific d...
Pith/arXiv arXiv 2026
-
[59]
Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2023
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2023. URLhttps://openreview.net/ forum?id=ehfRiF0R3a
2023
-
[60]
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O. Stanley. Paired open-ended trailblazer: Endlessly generating increasingly complex and diverse learning environments and their solu- tions.arXiv preprint arXiv:1901.01753, 2019
Pith/arXiv arXiv 1901
-
[61]
Rui Wang, Joel Lehman, Aditya Rawal, Jiale Zhi, Yulun Li, Jeff Clune, and Kenneth O. Stanley. Enhanced poet: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions.arXiv preprint arXiv:2003.08536, 2020
Pith/arXiv arXiv 2003
-
[62]
Wenyi Wang, Piotr Piękos, Li Nanbo, Firas Laakom, Yimeng Chen, Mateusz Ostaszewski, Mingchen Zhuge, and Jürgen Schmidhuber. Huxley-gödel machine: Human-level coding agent development by an approximation of the optimal self-improving machine, 2025. URLhttps: //arxiv.org/abs/2510.21614
arXiv 2025
-
[63]
Llms can’t jump, January 2026
Tom Zahavy. Llms can’t jump, January 2026. URLhttps://philsci-archive.pitt.edu/ 28024/. 24
2026
-
[64]
Darwin godel machine: Open-ended evolution of self-improving agents, 2026
Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune. Darwin godel machine: Open-ended evolution of self-improving agents, 2026. URLhttps://arxiv.org/abs/2505. 22954. 25
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.