{"id":"c30832a7-ad0e-4346-a3d5-0fa84a0a501a","arxiv_id":"2509.10875","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper argues that the agent paradigm in AI is conceptually ambiguous and anthropocentric, and that next-generation intelligence may be better pursued through system-level, world-model, and material-computing frameworks.","lead":"This paper argues that the 'agent' framing in AI, which treats systems as goal-seeking individuals, may be limiting research and hiding how large language models actually compute. It proposes instead to study system-level dynamics, world models, and material-based intelligence, and it adds a bibliometric analysis to support that shift.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The knowledge-graph diagnosis is not reproducible, and the 'structural crisis' claim exceeds what the 98-concept graph can establish.","rationale":"I agree with the reader that the load-bearing weakness is the unvalidated, unreleased knowledge graph. The paper's central conceptual argument, that the agentic framing of LLM systems is a sophisticated facade that can obscure underlying computational mechanisms, is coherent, well-referenced, and does not depend on the graph. The graph supports the added empirical assertions of a structural crisis, a stable theory-practice gap, and the specific Atlas of Opportunity frontiers. Those assertions go beyond a mere review essay and are the part of the paper that most needs verification. Section V provides only a few paragraphs of description; there is no protocol, no data deposit, no sensitivity analysis, and the graph is built with the authors' own Discovery Engine tool, which invites circularity. The paper's own Section III.A concedes that operationalizing the agentic/agential distinction is difficult, yet the graph treats 'Agential Systems' as a well-defined category, underscoring the risk that the graph encodes the authors' preferred taxonomy rather than the field's actual structure. Because the conceptual contribution is valuable, the appropriate verdict is unchanged: CONDITIONAL. The condition should be that the quantitative diagnosis is either made fully reproducible and validated or explicitly repositioned as an illustrative visualization. I do not see a basis for rejection, since the philosophical argument stands independent of the bibliometric analysis.","tokens_in":23854,"tokens_out":4401,"duration_ms":34523,"concrete_test":"Release the full 98-node, six-category edge list with per-edge source citations, node-category assignment protocol, and period labels. Then rerun Figures 1-4 from that data with two independent coders performing concept extraction and edge assignment on a random 30% subsample, computing Cohen's kappa for node inclusion, category assignment, and edge presence. Rerun the PageRank and Atlas-of-Opportunity computations after (a) removing all self-citations to the Discovery Engine and (b) deleting the most generic Critique/Challenge nodes. If the top frontier scores shift by more than 20%, or the temporal gap flips sign, the quantitative claims are not robust as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest empirical claim, that the field exhibits a 'structural crisis' and a persistent theory-practice gap, rests entirely on Section V's 98-concept knowledge graph. That graph is not released; no inclusion criteria are given for the 98 concepts; no edge-construction protocol is described; and no inter-annotator agreement or external validation is reported. The graph is built with the authors' own Discovery Engine tool (ref [60], per the Acknowledgments), and although the text states concepts were classified into six categories and edges represent explicit links in the source literature, it never specifies who performed the classification, how disagreements were resolved, or whether edges are directed, weighted, or time-stamped. These are not cosmetic omissions: the temporal analysis (Figure 2) requires per-period edges, but the paper gives no period-assignment protocol; the Atlas of Opportunity evidence scores (Section V.D) depend on counts of shared third-party concepts, so node-selection and category-assignment choices directly determine the headline frontier scores of 52 and 45. Absent the underlying data and protocol, the graph cannot independently support strong field-level language like 'structural crisis' or a 'persistent and stable gap.' If the graph merely encodes the authors' taxonomy, the quantitative diagnosis is an illustration of their own framework, not an empirical finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the 'agent' paradigm, especially as applied to LLM-based systems, is a limiting framework for next-generation AI. It introduces a trichotomy of 'agentic' (semi-autonomous AI with an appearance of agency), 'agential' (fully autonomous, self-producing biological systems), and 'non-agentic' (tools without the impression of agency) systems, and it reviews conceptual ambiguities and anthropocentric biases in agent definitions, with Active Inference and LLM-based agents as recurring case studies. The paper proposes an alternative research agenda centered on system-level dynamics, world models, embodied and material intelligence, and agential systems. To support this, Section V presents a knowledge-graph analysis of 98 concepts from the authors' literature review, reporting category-level influence, temporal trends, innovation-strategy quadrants, and an 'Atlas of Opportunity' heatmap. The conclusion states that the field is undergoing a 'structural crisis' characterized by a persistent theory-practice gap, and that critiques of anthropocentrism are now more influential than the foundational agent concept itself.","tokens_in":24183,"tokens_out":3542,"duration_ms":33538,"significance":"The conceptual argument is coherent and well-grounded in a broad literature, including external critiques by Jaeger and Shanahan, and it has practical value as a provocation to reconsider agent-centric assumptions in LLM-based systems. The paper's distinction between agentic and agential systems, while admittedly difficult to operationalize, is a useful framing device for a debate that is often muddled. The 'Atlas of Opportunity' offers concrete, if illustrative, research directions at the method-critique and application-critique interfaces. However, the paper's strongest empirical claim—that the field exhibits a 'structural crisis' and a persistent theory-practice gap—rests entirely on a non-reproducible, unvalidated knowledge graph constructed with the authors' own tool and taxonomy. As it stands, the quantitative diagnosis is best read as an illustration of the authors' framework rather than an independent empirical finding. This significantly limits the current evidentiary weight of the paper, though the conceptual core remains defensible and worth publishing after substantial revision.","major_comments":[{"comment":"The knowledge-graph analysis is not reproducible as reported, and it is load-bearing for the paper's central empirical claims of a 'structural crisis' and a 'persistent and stable gap' between theory and practice. The manuscript does not release the graph, does not specify inclusion criteria for the 98 concepts, does not describe the edge-construction protocol, and does not state who performed the six-category classification or how disagreements were resolved. The temporal analysis in Figure 2 requires per-period edges, but no period-assignment protocol is given; the influence-vs-interdisciplinarity analysis in Figure 3 requires directed or weighted connections whose nature is never specified. Without these details and without external validation, the 'strong empirical evidence' claimed in the Introduction and Section V is not supported.","section":"Section V, Figures 1-4"},{"comment":"The quantitative diagnosis risks circularity because the graph is built from the authors' own Discovery Engine tool and their own six-category taxonomy, and the node set and category assignments directly encode the paper's agentic/agential/non-agentic trichotomy. For example, the conclusion that 'Agential Systems' and 'systemic and emergent intelligence' are at the field's frontier follows in part from the authors choosing to include these as influential concepts in the graph. The finding that 'Critique/Challenge' concepts are increasingly central is similarly shaped by which critique concepts were selected and how they were connected. To make the diagnosis credible, the authors should provide an independent audit, an alternative taxonomy check, or a sensitivity analysis showing that the main conclusions are robust to node selection, category assignment, and edge-construction choices.","section":"Section V.A and Acknowledgments"},{"comment":"The headline evidence scores of 52 and 45 are counts of shared third-party concepts between pairs of categories, so they depend entirely on the manually constructed node set and category assignments. The paper presents these scores as a 'data-driven roadmap' without any null model, permutation test, or confidence interval, so it is unclear whether these frontiers are statistically meaningful or simply reflect the density of the authors' own concept selection. At minimum, the authors should report the raw contingency table and a permutation-based significance test, and they should temper the language that presents the Atlas as an objective empirical result.","section":"Section V.D (Atlas of Opportunity)"},{"comment":"The paper asserts a 'structural crisis' in the conclusion, but the presented analysis only shows correlations within a self-constructed graph; it does not establish that the field is 'under strain' in any causal or structural sense. The conceptual arguments in Sections II-IV are plausible, but the quantitative evidence is not strong enough to move the conclusion from a programmatic position piece to an empirically established diagnosis. The authors should either substantially strengthen the empirical support (data release, validation, sensitivity analysis) or explicitly reframe the conclusion as a hypothesis-illustrating exercise rather than a confirmed finding.","section":"Section III.A and Section VI"}],"minor_comments":[{"comment":"The headings contain typographical spacing errors: 'Chesterton's F ence' and 'Ashby's Law (Requisite V ariety)' should be 'Chesterton's Fence' and 'Ashby's Law (Requisite Variety)'.","section":"Section IV.C"},{"comment":"References [84] and [85] are exact duplicates of the same Chan et al. paper; one should be removed and the in-text citations renumbered accordingly.","section":"References [84] and [85]"},{"comment":"The text refers to colors in Figures 1 and 3 (red, orange, green) and to 'Generative Crossroads,' 'Bridging Niches,' and 'Established Cores' quadrants, but the figures are not included in the manuscript text provided, and the quadrant thresholds are not defined; please add the figures or provide a precise description of the axes and color legend.","section":"Section V.A and V.C"},{"comment":"The phrase 'systematic review' is used without specifying a review protocol (e.g., database search, screening criteria, or number of sources screened); either describe the protocol or replace 'systematic' with 'literature-based'.","section":"Abstract and Section V"},{"comment":"In the sentence beginning 'Rather than being an exclusive attribute of discrete agents, such internal representations” can be seen...', the opening quotation mark is missing or mismatched; please correct the quotation formatting.","section":"Section IV.A"}],"recommendation":"major_revision","confidential_remarks":"The paper's conceptual contribution is real and fits the journal's interest in foundational AI questions, but the quantitative section is currently a self-validating exercise rather than an empirical analysis. The path to acceptance requires either releasing the graph and protocol plus a validation/sensitivity analysis, or explicitly repositioning the knowledge-graph portion as an illustrative tool. I would also flag that the authors' own Discovery Engine is both the construction tool and a proposed solution in the paper, which creates a perceived conflict of interest in the 'data-driven' claims; this should be addressed transparently. The recommendation is major_revision rather than rejection because the conceptual core is defensible and the empirical gap is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper argues that the 'agent' framing, especially for LLM systems, is a limiting, anthropocentric metaphor, and suggests shifting to system-level dynamics, world models, and material intelligence. None of that is new—the authors lean on Jaeger's algorithmic mimicry, Shanahan's simulacra, and the existing anthropocentrism critiques—but it is a fair synthesis and a readable statement of the position. The agentic/agential/non-agentic taxonomy is tidy enough to be pedagogically useful, even if 'agential' is defined so that only biology currently qualifies, which makes the later call to study 'agential systems' a bit confusing.\n\nThe real trouble is Section V. The paper claims 'strong empirical evidence' for a structural crisis and a persistent theory-practice gap, based on a 98-concept knowledge graph. That graph is not released. There are no inclusion criteria for the 98 concepts, no edge-construction protocol, no inter-annotator agreement, no sensitivity analysis. The graph is built with the authors' own Discovery Engine tool. Under those conditions, the quantitative diagnosis is at best an illustration of the authors' taxonomy, not an empirical finding. The 'structural crisis' language is not supported. The Atlas of Opportunity scores (52 and 45) are direct functions of node and category selection, so they cannot stand as 'frontiers' without the data. The citation pattern is otherwise fine; the problem is not self-citation but the opacity of the tool's output.\n\nThe conceptual part of the paper is coherent and honest. They explicitly concede that agentic models retain heuristic value and that abandoning the metaphor requires caution (they even invoke Chesterton's Fence). The literature coverage is appropriate for an opinion piece. My verdict: worth engaging as a standpoint piece, not as a research contribution. If I were refereeing, I would accept it only with major revision: release the graph and fully document its construction, or drop the empirical claims and reframe Section V as a suggestive visual tool. In its current form, the quantitative claims should not be cited as evidence.","headline":"Coherent but largely derivative critique of the agent paradigm; the quantitative 'diagnosis' is too under-specified to support the paper's strong claims.","tokens_in":24598,"tokens_out":2380,"would_cite":false,"duration_ms":21363,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the 'agent' framing of AI, especially for LLM-based systems, is a sophisticated but limiting facade that obscures the underlying tensor computations, and it proposes shifting research toward system-level dynamics…","keywords":["agentic AI","agent paradigm","anthropocentrism","large language models","world models","material intelligence","system-level intelligence","knowledge graph analysis"],"falsifier":"Rebuild the paper's quantitative diagnosis from an independently selected literature corpus with the list of papers and the rules for drawing links fixed in advance; if the resulting knowledge graph no longer shows Critique/Challenge concepts at the center and no longer shows a theory-practice gap, the structural-crisis claim collapses. For the facade claim, run a controlled benchmark comparing an LLM-based agent-framed pipeline against the same model used as a bare input-output tensor function on tasks marketed as 'agentic'; if the agentic framing consistently adds measurable capability, the claim that it only obscures mechanisms is weakened.","tokens_in":23611,"feed_emoji":"🤖","tokens_out":6747,"duration_ms":55354,"temperature":0.7,"pith_summary":"This paper argues that the field's near-universal habit of describing AI systems as 'agents'—autonomous entities with beliefs, goals, and intentions—is a limiting framework rather than a neutral description. The authors distinguish agentic systems (semi-autonomous AI, including LLM-based tools, that give the impression of agency), agential systems (fully autonomous, self-producing systems, currently only biological), and non-agentic systems (tools without any agency-like framing). They claim that the 'agentic' view of LLMs, while heuristically convenient, is a sophisticated facade that obscures the real computational mechanisms, which are high-dimensional tensor operations and pattern completion rather than genuine goal-directedness. A quantitative knowledge-graph analysis of the literature is presented as evidence of a persistent gap between theoretical and critical discourse on one side and practical implementation on the other, with critiques of anthropocentrism now more central than the agent concept itself. If correct, this reorientation would push AI research toward world models, system-level and material intelligence, and away from engineering human-like autonomous entities.","feed_headline":"Calling AI 'agents' may be holding back the field","feed_subtitle":"A new analysis says LLM-based 'agentic' AI is a useful illusion, and system-level design deserves more attention.","key_machinery":"The argument rests on two coupled devices. The first is a tripartite taxonomy: agentic (AI that gives the impression of autonomous, goal-directed behavior without deep autonomy), agential (fully autonomous, self-producing systems, currently only biological), and non-agentic (tools without any agency-like impression). The second is a quantitative knowledge graph built from the paper's literature review: 98 concepts in six categories (Theoretical Concept, Architecture/Model, Entity/System, Method/Technique, Application/Domain, and Critique/Challenge), with edges recording explicit links in the sources. Node influence is measured by a centrality score, interdisciplinarity by the diversity of a concept's connections across categories, and under-explored links by a co-occurrence 'evidence score' heatmap the paper calls the Atlas of Opportunity. The taxonomy supplies the conceptual claim, and the graph supplies the empirical diagnosis of a field whose center of gravity is critical debate rather than foundational theory.","core_discovery":"The central claim is that the agent-centric paradigm, especially in the current wave of LLM-based 'agentic AI', is operationally and conceptually misleading: what these systems compute is not agency but sequences of tensor transformations over high-dimensional embeddings, and describing them as agents with beliefs, plans, or intentions imposes an anthropocentric map onto a mathematical territory. The paper proposes replacing the default agent frame with a focus on agential systems—where intelligence is an emergent, distributed, system-level property—and on non-agentic computing, world models, continuous interaction, and material substrates as legitimate and possibly superior routes to general intelligence. It also claims, on the basis of its knowledge-graph analysis, that the field is in a structural crisis: critique of anthropocentrism is now more central to the discourse than the foundational agent concept, while theory and practice remain persistently disconnected.","pith_inferences":["A testable corollary the paper leaves implicit: if the facade claim is right, then on a fixed benchmark an LLM-based system stripped of agentic scaffolding (no tool loop, no planning language, just direct input-output tensor computation) should match or approach the agent-framed version's performance; where it does not, the framing may be doing real engineering work.","The same knowledge-graph method could be applied to other contested concepts, such as 'intelligence', 'understanding', or 'alignment'; a symmetric finding—critique outweighing foundational theory—would suggest the pattern is general to fields in conceptual transition, not specific to agency.","The taxonomy's claim that agential systems are currently only biological implies that any future non-biological AGI would have to be either agentic (semi-autonomous, facade-like), non-agentic (a tool), or a new kind of agential system; the paper does not say which it expects, but the distinction sets up that question.","If anthropomorphism is mainly a product of interface design and marketing, as the paper suggests, then the agentic framing of consumer AI could be decoupled from the underlying engineering without changing performance—an economic and regulatory lever the paper mentions but does not develop."],"forward_implications":["LLM-based 'agentic' systems should be understood primarily as pattern-completion and tensor-transformation machines; agentic language remains a user-interface convenience, not an explanation of their operation.","Research funding and design effort would shift from building autonomous goal-seeking entities toward world models, continuous sensorimotor interaction, self-organization, and material or unconventional computing substrates.","The paper's Atlas of Opportunity identifies the most promising frontier as work that connects methods to critiques—for example, reinforcement-learning algorithms robust to Goodhart's Law, or formal verification of whether a neural architecture is computationally equivalent to an inferential algorithm.","Governance and accountability for AI would be reframed around verifiable system behavior and emergent properties rather than assumed intentions of an 'agent'.","The agent metaphor would be retained where it has heuristic value, such as human-AI interaction design, but dropped as the default ontology for intelligence research."],"supporting_citations":[{"why":"Supplies the core claim that LLM-based 'agents' are algorithmic mimicry rather than genuine agents.","marker":"[29]"},{"why":"Supplies the 'simulacra of agency' idea used to call the agentic framing a facade.","marker":"[39]"},{"why":"Supplies the literature-processing tool used to build the 98-concept knowledge graph and its quantitative measures.","marker":"[60]"},{"why":"Supplies the system-level, embodied bounded-rationality and world-model alternative to agent-centric design.","marker":"[38]"},{"why":"Supplies the As-If Fallacy critique used against attributing beliefs and inferences to Active Inference systems.","marker":"[23]"},{"why":"Supplies the 'Markov blanket trick' critique that undercuts the formal boundary of the agent concept.","marker":"[70]"},{"why":"Supplies the equivalence between Active Inference and Control as Inference, used to argue agentic framing can be less computationally efficient than direct RL.","marker":"[21]"},{"why":"Supplies the nomenclature-consensus argument that anthropocentric definitions of intelligent systems are contested and need broadening.","marker":"[30]"},{"why":"Supplies the ethical stakes (moral crumple zone, Goodhart's Law) that make the agent metaphor a governance problem.","marker":"[84]"}],"fun_headline_variants":["Agent framing misleads AI, review says","Beyond 'agents': system-level design for smarter AI","Agentic AI is an illusion; let's focus on systems","Retire the agent metaphor to advance AI research","AI's agent obsession may be its biggest limit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantitative evidence for a structural crisis assumes that its self-built map of 98 concepts and the links it drew between them fairly represents the whole field; if that map is idiosyncratic, the claimed theory-practice gap loses its empirical support.","fun_headline_variants_meta":{"raw":{"variants":["Agent framing misleads AI, review says","Beyond 'agents': system-level design for smarter AI","Agentic AI is an illusion; let's focus on systems","Retire the agent metaphor to advance AI research","AI's agent obsession may be its biggest limit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000325,"raw_usage":{"total_tokens":1842,"prompt_tokens":986,"completion_tokens":856,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":781}},"tokens_in":602,"tokens_out":856,"duration_ms":8416,"temperature":1.0,"reasoning_tokens":781,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:51:28.116208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rebuild the paper's quantitative diagnosis from an independently selected literature corpus with the list of papers and the rules for drawing links fixed in advance; if the resulting knowledge graph no longer shows Critique/Challenge concepts at the center and no longer shows a theory-practice gap, the structural-crisis claim collapses. For the facade claim, run a controlled benchmark comparing an LLM-based agent-framed pipeline against the same model used as a bare input-output tensor function on tasks marketed as 'agentic'; if the agentic framing consistently adds measurable capability, the claim that it only obscures mechanisms is weakened.","supporting_citations":[],"review_version":2}