{"id":"0afe927a-663c-4c7b-a86b-c93a4bb6123a","arxiv_id":"2412.07975","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper offers a philosophical definition of meaning for AI systems and argues current language models are already 'machines of meaning' but are limited by fixed vocabularies and full-distribution outputs.","lead":"This paper argues that meaning for artificial agents should be defined as the learned connection between symbols and their contexts, grounded in an agent's goals. It proposes a conceptual framework, 'machines of meaning,' to clarify how we talk about what language models do and don't understand.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that LLMs are already machines of meaning is either trivial on the paper's own Definition 3 or contradicted by Definition 2 and Section 6's admission that LLMs lack proper grounding.","rationale":"Section 5 asserts as a conclusion what Definition 3 has effectively stipulated. Because definitional stipulation alone cannot carry the claimed substantive discovery, the weakest load-bearing point is the equivocation between a weak statistical reading of 'meaning' (which any n-gram model satisfies) and the strong goal-directed grounding reading of Definition 2 (which the paper's own conclusion denies to LLMs). This is more fundamental than the Chalmers functionalism premise flagged by the reader: even granting Chalmers, the paper must decide what its definitions require. I therefore partially agree with the reader's weakest_assumption check, but recommend a different emphasis. I would mark the paper CONDITIONAL rather than simply UNVERDICTED, because the central assertion can be made coherent only by revising either Definition 2, Definition 3, or the Section 5 claim; as written it is internally unstable. This is not an objection to the paper's philosophical synthesis, which is often careful, but to the one sentence that carries the headline claim.","tokens_in":20431,"tokens_out":6624,"duration_ms":73561,"concrete_test":"Construct two formal interpretations of Definition 3: (W) any learned conditional distribution P(symbol|context) optimized by a prediction objective counts as meaning; (S) Definition 3 requires Definition 2's goal-guided experience and world-model update. Then test whether a bigram model and a transformer LLM satisfy W and S respectively. If the bigram model satisfies W, the paper's 'already machines of meaning' assertion is true of all statistical language models and is not informative about LLMs; if the LLM fails S because its 'goals' do not guide context selection in the required sense, then Section 5's assertion is incompatible with Section 6's admission. A second decisive check: replace 'machine of meaning' in Section 5, first paragraph, with the formal predicate from Definition 3 and verify whether the sentence remains true under the paper's own definitions.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim appears in Section 5, first paragraph: 'modern implementations of language models are already machines of meaning.' For this to be substantive, Definition 3 must be non-vacuous and Definition 2's grounding requirement must be met. The paper never shows that a current LLM has 'experience' of contexts in which symbol use is relevant to its goals in the sense of Definition 2; a transformer is trained on fixed text with a next-token objective and does not update a world model from goal-relevant experience at inference. Worse, the conclusion explicitly says current approaches are 'lacking proper grounding mechanisms for language semantics,' and Section 5.3 lists fixed lexicons and full output distributions as 'major obstacles to their full potential as MoMs.' So the assertion is either read weakly, in which case any n-gram model with a prediction objective already qualifies as a machine of meaning and the claim loses its force, or read strongly, in which case the authors' own limitations section contradicts it. This is an internal equivocation in the central claim, independent of whether Chalmers' functionalist principle from Section 3 is accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a conceptual and philosophical essay that proposes definitions of symbol, grounding, and meaning for artificial agents. It argues against Searle's Chinese Room Argument by appealing to Chalmers' functionalist principle that a simulation of X can be an X when the property of being an X depends only on functional organization. The paper reviews the trajectory from structural linguistics and distributional semantics to modern neural language models, and then claims in Section 5 that modern language model implementations are already 'machines of meaning.' It identifies two major obstacles to their full potential: the reliance on fixed lexicons and the training objective of outputting full probability distributions. The proposed way forward combines compressive encoding (e.g., random projections) with likelihood-free estimation methods to enable incremental, open-vocabulary learning.","tokens_in":20662,"tokens_out":4191,"duration_ms":45121,"significance":"If the paper's central claim were established, it would reframe debates about whether LLMs have meaning and would focus attention on specific architectural obstacles. The paper's strengths are its clear separation of symbols, grounding, and meaning; its broad and relevant synthesis from philosophy of language through computational semantics to LLMs; and its explicit naming of the 'prediction frame problem' with a concrete research agenda. The paper does not provide empirical evaluations, machine-checked formalizations, or falsifiable predictions; its contribution is definitional and programmatic. This kind of conceptual work can be valuable for a field that often conflates model behavior with understanding, but the central assertion that current LLMs are already machines of meaning must be internally consistent and non-vacuous for the paper to be convincing.","major_comments":[{"comment":"The paper asserts 'modern implementations of language models are already machines of meaning', but Section 6 states that current approaches are 'lacking proper grounding mechanisms for language semantics', and Section 5.3 lists fixed lexicons and full output distributions as 'major obstacles to their full potential as MoMs'. Under Definition 2, grounding is the process of updating a world model by experiencing contexts where symbol use is relevant, guided by the agent's goals. A transformer trained on static text with a next-token objective does not update a world model during inference and has no goal-directed experience in the sense of Definition 2. If the assertion is read weakly, any n-gram model with a prediction objective already qualifies as a machine of meaning, which trivializes the claim; if read strongly, it contradicts the paper's own limitations. This equivocation is load-bearing and must be resolved.","section":"Section 5, first paragraph"},{"comment":"Definition 3 defines meaning as a learned connection between a symbol and its referent in a context, with relevance determined by a goal. The paper never gives an operational criterion for when a model's parameters instantiate such a learned connection, nor does it show that a current LLM's training objective (e.g., next-token prediction or RLHF reward) amounts to the goal-directed relevance required by Definition 2. Without such a criterion, the claim that current LLMs are machines of meaning is either vacuously true under a permissive reading of 'goal' or unsupported under a strict reading. The authors should state which reading they intend and justify it, especially because the conclusion explicitly denies that current approaches have proper grounding.","section":"Definitions 2 and 3"},{"comment":"The refutation of the Chinese Room Argument depends on Chalmers' principle that 'A simulation of X can be an X when the property of being an X depends only on the functional organization of the underlying system, and not on any other details.' This premise is asserted rather than defended; the artificial-heart and artificial-limb analogies are suggestive but do not establish the premise. Because this functionalist move is what licenses treating LLM computational states as potential carriers of meaning, the authors should either provide an argument for the premise or explicitly mark the claim as conditional on a contested philosophical position. As it stands, the central claim rests on an unargued assumption that many readers, including Searle and subsequent critics, will reject.","section":"Section 3"}],"minor_comments":[{"comment":"There is a typo: 'Or arse they functional duplicates of hearts?' should read 'Or are they functional duplicates of hearts?'.","section":"Section 3"},{"comment":"The name 'Dennett' is misspelled as 'Dennet' in the discussion of the Twin Earth thought experiment.","section":"Section 4"},{"comment":"Several references are incomplete or inconsistent in formatting; for example, [8], [12], [22], and [34] lack complete journal, volume, or publisher details. A journal submission should have full bibliographic entries.","section":"References"},{"comment":"The 'prediction frame problem' is a useful concept, but its statement would be clearer with a formal or semi-formal definition, as currently it is described through examples rather than a precise problem formulation.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"X, here is the take. This is a philosophy paper, not a technical one, and judged as a philosophical essay it is mostly good. The authors survey meaning-as-use, symbol grounding, and the frame problem, and connect them cleanly to the fixed-vocabulary and full-output-distribution limitations of LLMs. The discussion of Putnam and Kripke is competent, and the critique of hype around LLM 'understanding' is on target. The paper would make a decent reading-group piece.\n\nThe central claim in Section 5, that modern LLMs are already machines of meaning, is equivocal. Under Definition 3, meaning is a learned connection between a symbol and its referent in a context that matters for a goal. A next-token predictor has no referents, only distributional contexts, so either the definition is stretched until any statistical language model qualifies, or the claim is trivial. Under Definition 2, grounding requires an agent to update a world model by experiencing contexts relevant to its goals. A transformer trained on fixed text with a prediction objective does not do this at inference. The conclusion then admits exactly this: current approaches are 'lacking proper grounding mechanisms for language semantics.' The paper cannot have it both ways. This is not a minor slip; the headline claim loses its force once you pin down which definition is doing the work.\n\nThe other soft spots are minor by comparison. The 'prediction frame problem' is a relabeling of known limitations, not a new problem. The proposed direction involving random projections and likelihood-free estimation is speculative and leans on the authors' own prior work, which is fine but is not evidence. The functionalist premise borrowed from Chalmers is asserted rather than defended; that is more a gap than a fatal flaw, since the paper could proceed on a weaker claim.\n\nWho should read this: people working on AI semantics or safety who want a clear map of the philosophical vocabulary and a reminder that 'meaning' is often undefined. It deserves a serious referee at a venue that publishes conceptual work, but the authors should be pushed to either weaken the 'already machines of meaning' claim to 'partial candidates' or revise Definition 3 so the claim is substantive.","headline":"A thoughtful conceptual essay on meaning and grounding whose central claim about LLMs being already machines of meaning does not survive contact with its own definitions.","tokens_in":21139,"tokens_out":2278,"would_cite":false,"duration_ms":24523,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that meaning for an AI is a learned, goal-directed link between a symbol and its context, and that current large language models already satisfy this definition.","keywords":["meaning","grounding","symbol grounding problem","large language models","frame problem","computational semantics","artificial intelligence","language games"],"falsifier":"A concrete experiment: train two identical language models, one with a goal-dependent training signal such as human preference and one with pure next-token prediction, then measure for the same symbols whether each model's internal updates track goal relevance as Definition 2 requires. If the pure predictor already shows the same goal-relevant grounding, the goal clause in Definition 3 is superfluous; if the goal-trained model alone shows it, the paper's account is confirmed.","tokens_in":20234,"feed_emoji":"🤖","tokens_out":5572,"duration_ms":55366,"temperature":0.7,"pith_summary":"The paper tries to establish a precise, non-anthropocentric definition of meaning for artificial agents, then argues that today's large language models already qualify as machines of meaning. It defines meaning as a learned connection between a symbol and its referent in a context, where the context matters because it serves a goal. Grounding, on this view, is the process by which an agent updates its world model by experiencing the contexts in which symbols are used, guided by its goals. From this perspective, many apparent failures of language models to understand come from misconceptions about what meaning requires, not from a missing human-like capacity. The paper then identifies two concrete architectural obstacles that prevent current models from adapting to open, evolving linguistic worlds: fixed lexicons and the need to output full probability distributions.","feed_headline":"LLMs already make meaning, paper argues","feed_subtitle":"A goal-directed definition of meaning puts today's LLMs on the right side; the real gap is open vocabularies.","key_machinery":"The load-bearing machinery is the paper's three-part conceptual apparatus: the definition of a symbol as a behavioural pattern with no intrinsic meaning; the definition of grounding as goal-guided updating of a world model through contexts of use; and Definition 3, which makes meaning a learned, goal-directed symbol-to-context connection rather than propositional content. This apparatus is used to reject the claim that programs cannot understand, via the principle that a simulation of something can be that thing when the property in question depends only on functional organisation. It also introduces the 'prediction frame problem', the impossibility of fixing the lexicon and output space in advance, which the paper presents as the real obstacle separating current language models from machines of meaning that adapt to open linguistic worlds.","core_discovery":"The paper's central claim is that meaning, for an artificial agent, is the learned connection between a symbol and its referent in a context, where the features of that context are meaningful because they are important for a goal such as survival or coordination (Definition 3). Grounding is the process whereby an agent updates its world model by experiencing the contexts where a symbol is used, with relevance guided by the agent's goals (Definition 2). On these definitions, the paper argues, modern large language models already qualify as machines of meaning: they learn symbol-to-context associations from data and use them to pursue prediction and human-alignment goals. Their apparent failures to 'understand' stem not from a missing capacity but from a misconception about what meaning requires, namely that it must be human-like. The paper then identifies two architectural obstacles that prevent these models from being machines of meaning in open, evolving linguistic domains: fixed lexicons and the output of full probability distributions, and it proposes likelihood-free estimation and compressive encoding as candidate solutions.","pith_inferences":["Editorial inference: If Definition 3 is taken literally, almost any goal-directed reinforcement learner qualifies as a machine of meaning for its input symbols, so the term becomes a general property of goal-directed representation learning rather than a specifically linguistic achievement.","Editorial inference: A direct test of the 'already machines of meaning' claim is to compare a pure next-token predictor with a goal-conditioned variant on the same corpus; only if the latter shows measurably different symbol-to-context structure does the goal clause carry explanatory weight.","Editorial inference: The paper's proposed fixes point to a concrete extension: build a small-scale model with a continually growing lexicon and test its ability to acquire new symbols online, which would operationalize 'machines of meaning' in a way current frozen-vocabulary benchmarks cannot."],"forward_implications":["If the definitions are right, current LLMs can be said to possess machine-relative meaning without needing human-like embodiment or subjective experience.","The debate about LLM understanding shifts from 'can they mean at all' to 'what goals and contexts are their symbols grounded in'.","Scaling alone will not produce machines of meaning that keep pace with evolving language; architectures must allow incremental learning with open vocabularies.","Likelihood-free training and compressive symbol encodings become first-class research directions rather than merely practical tricks.","AI risk and capability assessment can be made more precise by evaluating goal-relevant grounding rather than anthropomorphic notions of understanding."],"supporting_citations":[{"why":"Supplies the principle that a simulation of X can be an X when the property depends only on functional organization, the step that rejects the Chinese Room conclusion.","marker":"[18]"},{"why":"Defines the symbol grounding problem and provides the Chinese Room analysis that the paper builds on.","marker":"[8]"},{"why":"The Chinese Room argument that the paper targets and reframes in terms of system-level grounding.","marker":"[13]"},{"why":"Argues the symbol grounding problem has been solved and distinguishes linguistic symbols from programming-language symbols.","marker":"[4]"},{"why":"Provides the 'meaning is use' view that the paper refines into its goal-directed Definition 3.","marker":"[32]"},{"why":"The neural probabilistic language model that anchors the claim that current approaches already learn goal-directed symbol-to-context connections.","marker":"[65]"},{"why":"The critique of LLM hype that the paper considers a misconception about what meaning requires.","marker":"[2]"},{"why":"Shows human feedback is used as a training goal, supporting the goal-directed component of meaning.","marker":"[68]"}],"fun_headline_variants":["Meaning is goal-driven; LLMs already have it","LLMs qualify as meaning machines, says paper","Goal-directed meaning: LLMs already pass the test","LLMs make meaning; open vocab is the real gap","Define meaning by goals, and LLMs already win"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that being a thing that understands depends only on how it is functionally organised, not on what it is made of; if that premise fails, the claim that language models can genuinely have meaning loses its footing.","fun_headline_variants_meta":{"raw":{"variants":["Meaning is goal-driven; LLMs already have it","LLMs qualify as meaning machines, says paper","Goal-directed meaning: LLMs already pass the test","LLMs make meaning; open vocab is the real gap","Define meaning by goals, and LLMs already win"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1247,"prompt_tokens":964,"completion_tokens":283,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":207}},"tokens_in":580,"tokens_out":283,"duration_ms":3592,"temperature":1.0,"reasoning_tokens":207,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:19:55.559904+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete experiment: train two identical language models, one with a goal-dependent training signal such as human preference and one with pure next-token prediction, then measure for the same symbols whether each model's internal updates track goal relevance as Definition 2 requires. If the pure predictor already shows the same goal-relevant grounding, the goal clause in Definition 3 is superfluous; if the goal-trained model alone shows it, the paper's account is confirmed.","supporting_citations":[{"cited_title":"John Wiley & Sons (2010)","cited_arxiv_id":null,"evidence_quote":"Provides the 'meaning is use' view that the paper refines into its goal-directed Definition 3."},{"cited_title":"In: Leen, T., Dietterich, T., Tresp, V","cited_arxiv_id":null,"evidence_quote":"The neural probabilistic language model that anchors the claim that current approaches already learn goal-directed symbol-to-context connections."},{"cited_title":"Advances in Neural Information Processing Systems 35, 27730–27744 (2022)","cited_arxiv_id":null,"evidence_quote":"Shows human feedback is used as a training goal, supporting the goal-directed component of meaning."}],"review_version":1}