{"id":"aa27b8a5-3240-4a71-822b-0b03a329a9e5","arxiv_id":"2509.02515","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper that maps LLM agents onto classical MAS ideas (BDI, artifacts, ACLs, norms) and warns that many new 'multi-agent' systems fall short of true agency.","lead":"This paper is a reflection essay that compares LLM-driven agents with classic multi-agent systems. It argues that new work should build on decades of MAS research instead of ignoring it.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central warning assumes, without defense, that classical MAS criteria are the correct yardstick; absent that defense, the claim that many LLM systems are 'not true MAS' is unfalsifiable.","rationale":"The reader's weakest_assumption identified exactly the concern I find most load-bearing: the classical MAS yardstick is assumed, not defended. The paper's central warning—that many LLM systems are 'sophisticated distributed LLM applications rather than true MAS'—depends on that yardstick being more than one possible definition. If natural-language-mediated coordination, prompt-embedded norms, and generative planning count as a new form of agency rather than a deficient version of the old one, the warning loses much of its force. The paper even softens its own stance in several places, acknowledging that the mental-state question is 'secondary to this work,' that the field is 'building a new kind of wheel,' and that experimental support for emergent cooperation 'is lacking.' This oscillation between empirical and definitional claims is the soft spot. However, because the paper is explicitly a reflection piece rather than an empirical study, this concern does not change the reader's UNVERDICTED verdict. It reinforces the appropriateness of that verdict: the central claim is not yet empirically testable as stated, and the paper would be stronger if it either operationalized its criteria or explicitly framed its warning as a normative/terminological position rather than a factual finding about the field.","tokens_in":23427,"tokens_out":5895,"duration_ms":54034,"concrete_test":"Operationalize the Section 3.1 criteria and test a representative sample. Have two independent evaluators score three widely used LLM-MAS frameworks (AutoGen, MetaGPT, CrewAI) on five binary criteria derived from the paper's own comparison: (i) agents maintain persistent identities beyond a single session; (ii) agents can select goals without human or central-controller intervention; (iii) agents form intentions/commitments that constrain future actions; (iv) inter-agent messages carry reliably identifiable illocutionary force; (v) agents can detect or enforce norms. If the systems satisfy at least four of five under a neutral reading, the paper's 'distributed LLM application' classification is empirically unsupported and the warning needs re-scoping; if they satisfy two or fewer, the warning is confirmed and gains concrete grounding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 opens by stating that many current systems 'might be more accurately described as sophisticated distributed LLM applications rather than true MAS in the classical sense' and uses this to motivate a call to build on foundational MAS knowledge. The load-bearing assumption is that the classic properties (autonomy, social ability, BDI, formal ACLs, explicit norms) are not only a useful lens but the correct, complete yardstick for judging LLM agents. The paper never defends this against the alternative that natural-language interaction, in-context 'beliefs,' and prompt-embedded 'norms' constitute a genuinely new form of agency whose criteria are not a subset of the classical ones. Indeed, the paper partly concedes this: Section 3.2 says whether LLMs 'genuinely believe' is 'secondary to this work'; Section 3.4 calls the shift 'building a new kind of wheel'; Section 3.5 admits 'experimental support that this will really happen (and will happen every time) is lacking.' If the yardstick is merely one possible definition, then the warning reduces to a terminological preference: LLM systems are 'not true MAS' by definition, which is not an empirical deficiency. The paper therefore oscillates between an empirical claim (systems lack autonomy/sociality) and a definitional claim (they are not classically structured), without giving a test that would separate the two.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper offers a wide-ranging reflection comparing LLM-driven agent systems with classical multi-agent systems. It surveys the architectural building blocks of contemporary LLM agents (profiles, memory, planning, action, communication, perception, learning) and then systematically re-examines classical MAS concepts through the LLM lens: core agent properties, the BDI model, agents-and-artifacts versus tool use, speech act theory versus LLM-mediated dialogue, coordination/cooperation/negotiation as emergent behavior, and norms versus implicit social laws. The central thesis is that many contemporary \"MAS LLMs\" are better described as sophisticated distributed LLM applications than as true MAS in the classical sense, and that the field should build on rather than ignore foundational MAS knowledge. The paper closes with challenges, limitations, and future research directions.","tokens_in":23673,"tokens_out":5117,"duration_ms":46745,"significance":"If its central thesis is accepted, the paper provides a useful integrative perspective that could help align LLM-agent research with decades of MAS theory. Its strengths are a broad and structured comparison, a balanced treatment that acknowledges both continuity and transformation, and an unusually candid limitations section. The paper is not a technical contribution—it contains no experiments, derivations, or formal models—but it can serve as an orientation map for researchers entering the area. It explicitly hedges many empirical claims and cites recent primary sources. Its primary value is as a position or reflection piece, and its success depends on whether the reader accepts the classical MAS yardstick as the appropriate basis for comparison. The paper would be strengthened substantially by separating definitional claims from empirical claims and by specifying what would count as evidence for or against its main warning.","major_comments":[{"comment":"Table 1 lists \"Autonomy, Reactivity, Pro-activeness, Social ability\" as \"Core properties\" shared by both Classic MAS and LLM-based agents, but Section 3.1 argues that LLM agents' autonomy is complicated by prompt dependence and non-determinism, and it cites reference [5] for the claim that many MAS LLM implementations lack genuine autonomy and sophisticated social interaction. This is a direct internal tension on a point that is central to the paper's thesis. The table should distinguish \"properties claimed by system designers\" from \"classical definitional properties,\" or the text of Section 3.1 should be revised to be consistent with the table.","section":"Section 3, opening paragraph; Table 1"},{"comment":"The claim that many systems \"might be more accurately described as sophisticated distributed LLM applications rather than true MAS in the classical sense\" is never operationalized with a falsifiable criterion. The paper oscillates between an empirical deficiency claim (systems lack genuine autonomy, social ability, or strategic reasoning) and a definitional claim (systems are not structured according to classical MAS architectures such as BDI or FIPA-ACL). These two readings require different evidence and have different consequences. I would ask the authors to specify a minimal, testable set of criteria for \"true MAS\"—for example, goal autonomy, explicit commitment or norm representation, and interaction governed by a protocol with defined semantics—and to indicate how the surveyed systems map to that test. Without such a test, the \"reinventing the wheel\" warning risks reducing to a terminological preference rather than a substantive concern.","section":"Section 3, opening paragraph; Section 3.5"},{"comment":"The sentence \"However, experimental support that this will really happen (and will happen every time) is lacking\" is an honest but damaging admission for the paper's earlier assertion that LLM-agent cooperation is an emergent, dialogue-driven behavior. The section then bases its positive claims on a small number of studies—for coordination, reference [38] (LLM-Coordination); for cooperation, reference [46] on Diner's Dilemma; for negotiation, the same reference number [46] on bargaining—each a single experimental setup with limited generality. To keep the reflective claim defensible, the authors should present these as preliminary hypotheses with explicit scope limitations, or they should gather a broader set of corroborating studies before generalizing to \"LLM agents\" as a class.","section":"Section 3.5, Cooperation subsection"}],"minor_comments":[{"comment":"Reference numbers are duplicated: [34] is used both for Bratman's book and for the NatBDI paper by Ichida et al., and [46] is used both for the bargaining paper by Oh et al. and for the Diner's Dilemma paper by Warnakulasuriya et al. This creates ambiguity in the citation trail and should be corrected by renumbering.","section":"References"},{"comment":"The sentence \"Note that similar considerations were made for earlier classic MAS architectures\" would benefit from a citation to one or two classic MAS surveys so that readers can locate the claimed precedent.","section":"Section 2.3, first paragraph"},{"comment":"The phrase \"building a new kind of wheel\" is a helpful nuance, but it partially undercuts the introduction's \"reinventing the wheel\" warning (Section 1). The authors should explicitly reconcile these two framings, for example by stating that the risk is reinventing the wheel for problems already solved (coordination, commitment, norm enforcement) even if the implementation medium is genuinely new.","section":"Section 3.4, last paragraph before the Agora discussion"},{"comment":"Reference [63] is a Medium blog post rather than a peer-reviewed source; consider replacing it with the CrewAI documentation or a peer-reviewed description of the framework if one is available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a position/reflection paper rather than a technical research contribution. Its main value is the structured comparison and the candid discussion of limitations. The major-revision recommendation is driven by the need to sharpen the central thesis: the paper must clearly separate definitional from empirical claims and provide criteria that would make the \"not true MAS\" claim testable. If the authors address that, the paper would be a solid contribution for a venue that welcomes surveys and reflections on emerging areas."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2509.02515. First, it is not a research paper; it is a reflection from a group of senior MAS people, and its value is as a framing document. Second, its central warning—that many LLM-based systems are \"sophisticated distributed LLM applications rather than true MAS\"—is real but under-defended, and the paper never decides whether it is making an empirical claim or a definitional one.\n\nWhat the paper does well: it gives a compact, readable mapping of where LLM agents sit relative to classic MAS concepts: BDI, A&A, ACLs, coordination, norms. Table 1 is genuinely useful as a teaching artifact. The sections on A&A vs. tool use/MCP and on formal ACLs vs. natural-language dialogue are the strongest; they name concrete systems (NatBDI, ChatBDI, JaCaMo, CrewAI) and make specific engineering observations that are easy to miss in the survey literature. The authors also cite the existing position papers (La Malfa et al., Botti) rather than pretending the critique is theirs alone.\n\nWhere it is soft: the load-bearing assumption is that the classical criteria (autonomy, social ability, BDI, formal ACLs, explicit norms) are the correct yardstick. That is asserted, not argued. If the yardstick is just one possible definition, then the claim that LLM systems are \"not true MAS\" is a terminological preference, not a deficiency. The paper occasionally acknowledges this—Section 3.2 says LLM mental states are secondary to its purposes, and Section 3.5 concedes experimental support for emergent cooperation is lacking—but it never steps back to say what empirical observation would count as genuine autonomy or sociality in an LLM system. The stress-test note is right about the oscillation.\n\nAlso, citation hygiene: [34] is used twice (Bratman and Ichida et al.), [46] is used twice, and [64] in the text refers to CrewAI but the reference list has API-Bank there. For a survey, that is a moderate but real problem; it undermines trust in the other references.\n\nWho is this for? Someone new to agentic AI who wants the classical MAS perspective in one place, or a MAS researcher who wants a quick comparison table. It is not a breakthrough, and the authors don't claim one. But it is a competent reflection from people who know the field.\n\nMy recommendation: it deserves peer review, not desk rejection. Send it to a venue that handles position papers, and ask the authors to fix the references and either defend the classical yardstick or reframe the \"not true MAS\" claim as one possible lens rather than a verdict.","headline":"A competent but under-defended reflection on LLM agents vs classic MAS; the comparison table is useful, but the central 'not true MAS' claim never decides whether it is empirical or definitional.","tokens_in":24195,"tokens_out":3049,"would_cite":false,"duration_ms":27110,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that many LLM-powered \"multi-agent\" systems are better described as distributed LLM applications than as true multi-agent systems, and that the field should build on classic MAS theory rather than ignore it.","keywords":["multi-agent systems","large language models","agentic AI","BDI architecture","agent communication","LLM agents","multi-agent LLM systems","norm emergence"],"falsifier":"A concrete result that would weaken the paper's central claim is a benchmark in which LLM-agent teams without BDI or ACL scaffolding reliably satisfy formal coordination and commitment guarantees—for example, completing a shared plan with deadline and resource constraints across repeated trials—matching or exceeding classic MAS. Alternatively, a demonstration that LLM agents trained for inter-agent interaction, rather than user response, exhibit stable social abilities would undercut the claim that they lack genuine social interaction.","tokens_in":1603,"feed_emoji":"🤖","tokens_out":5978,"duration_ms":69346,"temperature":0.7,"pith_summary":"The paper argues that the rapid wave of LLM-powered agent systems is replaying problems that classic multi-agent systems (MAS) research already solved, and that many systems branded \"MAS LLMs\" are better described as distributed LLM applications than as multi-agent systems in the classical sense. The authors map the architectural pillars of modern LLM agents—profile, memory, planning, action, communication, perception, and learning—onto the foundational MAS concepts of autonomy, BDI reasoning, agent communication languages, artifacts, coordination, cooperation, negotiation, and norms. Their aim is to show where LLM agents genuinely extend the classic framework and where they quietly drop its guarantees, so that current innovation can build on rather than ignore decades of MAS theory. A sympathetic reader would take the paper as a call for theoretical grounding and careful terminology, not as a rejection of LLM agents.","feed_headline":"LLM 'agent teams' often aren't true multi-agent systems","feed_subtitle":"A review maps modern LLM agents onto classic MAS theory and finds the social and autonomy core missing.","key_machinery":"The comparative apparatus is the classic MAS characterization itself: the core properties of autonomy, reactivity, pro-activeness and social ability; the BDI model of beliefs, desires and intentions; the agents-and-artifacts view of the environment; formal agent communication based on speech act theory; and the classic treatments of coordination, cooperation, negotiation and norms. Against this yardstick the paper places the LLM-agent scaffolding stack—profile, memory (short-term, long-term, retrieval-augmented), planning, action via tools and the Model Context Protocol, natural-language communication, multimodal perception, and learning via human feedback—and asks, for each pair, what is preserved, what is loosened, and what is lost. The machinery makes the \"true MAS vs distributed LLM app\" distinction concrete: formal verifiability, predictability and explicit intent are traded for flexibility, emergence and implicit intent.","core_discovery":"The paper's central claim is that contemporary LLM-driven systems, for all their power, frequently lack the core properties that define a true multi-agent system in the classic literature: genuine autonomy, robust social interaction, formal communication semantics, and verifiable norms. The authors argue that many current frameworks are sophisticated distributed LLM applications—orchestrated prompts, tool calls, and role assignments—rather than MAS in the classical sense, citing concerns from the MAS community that LLM agents are fine-tuned as single agents responding to users rather than trained for interaction with other agents. At the same time, they hold that classic concepts such as BDI, the artifacts view of the environment, speech act theory, and norm-based social order are not obsolete: they are being re-interpreted through natural language, with LLMs enriching beliefs, generating plans dynamically, and enabling negotiation through persuasive dialogue. The conclusion is an evolution of implementation, not an invalidation of the foundational concepts.","pith_inferences":["If the paper's yardstick is accepted, a testable program follows: benchmark LLM-agent teams specifically on formal coordination guarantees, commitment to intentions, and norm compliance rather than on task completion alone.","The authors' distinction suggests an empirical prediction: LLM-agent systems will show brittle performance precisely in settings that require reliable inter-agent commitments, such as long-horizon cooperative plans with deadlines.","The natural-language turn could be read not as a loss of social ability but as a new kind of social ability whose failure modes—ambiguity, prompt injection, hallucinated commitments—need their own theory; the paper gestures at this but leaves it undeveloped."],"forward_implications":["Systems labeled \"MAS LLMs\" should be expected to demonstrate genuine autonomy and social interaction, not just role-played conversations, before the label is applied.","Tool-access standards like the Model Context Protocol will not by themselves restore MAS-level guarantees; they standardize tool calls but not the semantics of inter-agent communication.","BDI-inspired hybrid architectures point to a path where the BDI reasoning cycle supplies structure and the LLM supplies interpretation.","Formal communication may return in the form of context-adaptive meta-protocols rather than fixed agent communication languages.","Hybrid systems combining rule-based or BDI agents with LLM agents are a likely direction for robustness."],"supporting_citations":[{"why":"Supplies the central critique that many LLM multi-agent implementations miss the mark on genuine autonomy and social interaction.","marker":"[5]"},{"why":"Provides the \"reinventing the wheel\" framing that the paper uses to motivate building on classic MAS knowledge.","marker":"[6]"},{"why":"Defines the classic agent properties and MAS theory that serve as the paper's evaluation yardstick.","marker":"[41]"},{"why":"Represents the classic BDI programming approach used as the reference point for agent architectures.","marker":"[33]"},{"why":"Introduces the agents-and-artifacts meta-model, the classic environment abstraction compared with LLM tool use.","marker":"[61]"},{"why":"Foundational speech act theory source for the classical approach to agent communication.","marker":"[39]"},{"why":"Refines speech act theory and underpins the design of formal agent communication languages.","marker":"[40]"},{"why":"Supplies evidence of emergent social conventions in LLM populations, used in the norms discussion.","marker":"[47]"},{"why":"Describes a meta-protocol for LLM networks, cited as a possible bridge between natural-language flexibility and formal communication structure.","marker":"[44]"}],"fun_headline_variants":["LLM agents: powerful, but often not truly multi-agent","LLM 'teams' miss the core of classic agent systems","Modern agent hype vs classic MAS: missing social core","LLM agents lack what makes MAS genuine","Are LLM agent teams real MAS? Usually not"],"cache_read_input_tokens":26368,"weakest_assumption_plain":"The argument assumes that the classic MAS criteria—autonomy, social ability, formal communication, BDI-style commitment, and verifiable norms—are the right and complete standard for judging LLM agents, so that deviations from them count as deficiencies rather than as a genuinely new form of agency.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents: powerful, but often not truly multi-agent","LLM 'teams' miss the core of classic agent systems","Modern agent hype vs classic MAS: missing social core","LLM agents lack what makes MAS genuine","Are LLM agent teams real MAS? Usually not"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":1120,"prompt_tokens":801,"completion_tokens":319,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":417,"completion_tokens_details":{"reasoning_tokens":240}},"tokens_in":417,"tokens_out":319,"duration_ms":3205,"temperature":1.0,"reasoning_tokens":240,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:36:38.250687+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete result that would weaken the paper's central claim is a benchmark in which LLM-agent teams without BDI or ACL scaffolding reliably satisfy formal coordination and commitment guarantees—for example, completing a shared plan with deadline and resource constraints across repeated trials—matching or exceeding classic MAS. Alternatively, a demonstration that LLM agents trained for inter-agent interaction, rather than user response, exhibit stable social abilities would undercut the claim that they lack genuine social interaction.","supporting_citations":[{"cited_title":"Zhang, Elizabeth Black, Michael Luck, Philip Torr, and Michael Wooldridge","cited_arxiv_id":null,"evidence_quote":"Supplies the central critique that many LLM multi-agent implementations miss the mark on genuine autonomy and social interaction."},{"cited_title":"An Introduction to MultiAgent Systems (2nd ed.), Wiley, 2009","cited_arxiv_id":null,"evidence_quote":"Defines the classic agent properties and MAS theory that serve as the paper's evaluation yardstick."},{"cited_title":"Bordini, Jomi F","cited_arxiv_id":null,"evidence_quote":"Represents the classic BDI programming approach used as the reference point for agent architectures."}],"review_version":2}