{"id":"6db53974-bd2d-4581-8790-53e9989d152b","arxiv_id":"2505.07079","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A NARS-based agent reaches 100% accuracy on derived same/opposite matching-to-sample tests, but the generalization is largely produced by explicitly pretrained entailment patterns and a hand-authored relational schema.","lead":"The paper reports a simulation in which the NARS reasoning system answers same/opposite matching tasks after training phases, including tests on stimulus pairs it never saw during training. It claims this shows 'arbitrarily applicable relational responding' similar to human symbolic reasoning, but the target inference patterns were explicitly trained and partly hand-coded into the system.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The BC result is pre-structured: Section 3.2's hand-authored variable schema and Phase 1's explicit symmetric/transitive training already contain the target generalization, so 100% BC accuracy does not by itself demonstrate emergent AARR.","rationale":"I read the paper as claiming that NARS, from sensorimotor training in a matching-to-sample procedure, acquires generalized same/opposite responding and derives untrained BC relations. The strongest support would be that the 100% BC result arises from learning restricted to reinforced AB/AC trials plus NARS's normal inference. That is exactly what is not shown. Sections 3.2-3.3 describe a generalized variable schema and explicit relational naming as part of the implementation, but give no induction or abstraction algorithm and no trace demonstrating these structures were derived from experience. Phase 1 explicitly trains the two abilities named in the abstract, symmetry and transitivity, across multiple exemplars, and Section 5.4 presents a rule that is a direct statement of combinatorial entailment. The accuracy result therefore cannot discriminate between a genuinely learned relational frame and execution of a hand-authored rule base. This is a correctness and interpretability risk, not a disagreement with external consensus. Independent support is limited: there are no machine-checked proofs, no released code or data, and the Discussion itself concedes reliance on empirical validation without formal analysis. An ablation would settle the matter. Since the Reader already identifies the same weakest assumption and rejects on that basis, my stress-test does not change the verdict.","tokens_in":6380,"tokens_out":4736,"duration_ms":51630,"concrete_test":"Release the model with a strict ablation: disable the hand-authored variable schema from Section 3.2 and the explicit symmetric/transitive pretraining of Phase 1; train only the Phase 2 AB/AC SAME/OPPOSITE mappings with feedback, then test the untrained BC pairs. If BC accuracy remains 100%, the emergent-AARR claim is supported; if accuracy drops toward chance, the headline result is carried by the pre-specified rules rather than by learned relational frames.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that same/opposite AARR emerges from minimal experience, with untrained BC trials derived from previously learned relations. The implementation as described makes that conclusion unsupported. In Section 3.2, the authors hand-author abstraction to variables ($1, $2, $3, $4), producing an implication that maps any related stimulus pair and any related location pair to a match action; Section 3.3 then adds 'explicit relational naming' without providing an acquisition mechanism. Phase 1 explicitly trains mutual entailment (Y->X) and combinatorial entailment (X->Z) across multiple exemplars. The illustrative rule in Section 5.4 is already a hand-formulated statement of SAME/OPPOSITE combinatorial entailment. If these components are part of the initial implementation rather than learned from the Phase 1-2 trials, the measured 100% BC accuracy simply executes the provided generalization. The paper's own Discussion concedes reliance on empirical validation rather than formal analysis of the acquired-relations mechanism, and no code or data are released to show that the schema was actually induced from experience. The risk is not that NARS cannot make the inference; it is that the paper has not shown the inference is emergent from minimal training.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the Non-Axiomatic Reasoning System (NARS) with an 'acquired relations' mechanism and applies it to a same/opposite matching-to-sample (MTS) task inspired by Relational Frame Theory (RFT). The experiment has three phases: Phase 1 explicitly pre-trains mutual entailment (symmetry) and combinatorial entailment (transitivity) for SAME and OPPOSITE relations; Phase 2 trains relational networks on novel stimulus sets AB and AC; Phase 3 tests untrained BC stimulus pairs. The reported results are perfect accuracy (100%) in all phases, including the critical BC phase, along with internal confidence metrics. The authors interpret this as demonstrating the emergence of arbitrarily applicable relational responding within NARS.","tokens_in":6636,"tokens_out":3625,"duration_ms":38048,"significance":"If the central claim were supported, the paper would offer a useful computational bridge between RFT and NARS and a concrete implementation of same/opposite relational responding. The formalization in Narsese and the explicit MTS design are clear, and the target phenomenon is genuinely interesting. However, the paper does not currently support the emergence claim: the generalized variable schema and the pretrained entailment rules appear to encode the tested generalization in advance, and no code, data, or learning trace is provided. As a result, the reported 100% BC accuracy is better described as execution of hand-authored rules than as evidence of learned, arbitrary applicable relational responding.","major_comments":[{"comment":"The generalized variable schema with $1, $2, $3, $4 is introduced as an implementation construct, not as something NARS learns from experience. The schema already states that if two stimuli are related and two locations are related, then the corresponding match action leads to the goal. Because this exactly covers the logic needed for the BC test in Section 5.3, the perfect BC accuracy is evidence that the hand-authored rule executes, not that arbitrarily applicable relational responding emerged.","section":"§3.2"},{"comment":"Phase 1 explicitly trains both symmetric (mutual entailment) and transitive (combinatorial entailment) relational frames before the BC test. Since the paper itself identifies mutual and combinatorial entailment as the components required for BC responding, the design cannot support the claim that BC responses emerge from untrained combination; the required component inferences were installed by explicit pretraining.","section":"§4, Phase 1"},{"comment":"The example combinatorial entailment hypothesis, written as <(<($1 * #1) --> SAME> && <(#1 * $2) --> OPPOSITE>) ==> <($1 * $2) --> OPPOSITE>>, is a hand-formulated Narsese rule that encodes the target SAME/OPPOSITE composition. The paper provides no learning trace or induction procedure showing that this rule was derived from the Phase 1–2 trials, so the internal-representation example does not establish the acquisition claim.","section":"§5.4"},{"comment":"The 'explicit relational naming' step is described as something NARS 'explicitly abstracted and internally represented' after learning explicit contingencies, but no mechanism is given for how the named relational statements are acquired from sensorimotor experience. This is a load-bearing gap because the named relational representations are the basis for the novel derivations reported in the BC phase.","section":"§3.3"},{"comment":"The reported 100% accuracy in the BC phase is presented without code, data, number of independent runs, or error bars. With four blocks of 16 trials, the result cannot be statistically assessed, and the 'internal confidence' metrics are not precisely defined, so the claim of 'strong internalization' is not independently verifiable from the supplied artifacts.","section":"§5.3 and Figure 3"}],"minor_comments":[{"comment":"The text says 'All phases included four blocks of 16 trials each,' but Section 5 reports accuracy per block without giving total trial counts or numbers of runs; please state these explicitly.","section":"§4"},{"comment":"The phrase 'minimal explicit training' should be qualified, because Phase 1 explicitly trains both mutual and combinatorial entailment, which are the very capabilities tested in the BC phase.","section":"Abstract and §4"},{"comment":"The figure caption says 'Average confidence' but does not define how the confidence values are computed from NARS; please add a precise definition and clarify the axis labels.","section":"Figure 3"},{"comment":"The term 'combinational' appears where 'combinatorial' is intended; please correct this for consistency with the rest of the paper.","section":"§5.4"}],"recommendation":"reject","confidential_remarks":"This manuscript presents a demonstration that is substantially pre-structured by the implementation: the variable schema and Phase 1 pretraining already contain the generalization steps that the BC test is supposed to reveal. To make a publishable claim of emergent AARR, the authors would need to provide a mechanism by which the generalized schema and relational names are actually learned from experience, along with code and data. As it stands, the central claim is unsupported, and the issue cannot be resolved by minor edits."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this paper is a NARS implementation of same/opposite relational responding, but the emergence claim doesn't hold up. The authors pre-train mutual and combinatorial entailment and hand-code a variable-abstraction schema that already contains the target generalization; the 100% BC accuracy is the schema running, not something the system derived from minimal experience.\n\nWhat's genuinely new: it extends the authors' prior stimulus-equivalence work in NARS to handle two relational frames (SAME/OPPOSITE), with an explicit Narsese encoding of MTS trials, acquired relations, and named relational statements. The experimental structure is clear: Phase 1 trains symmetry/transitivity, Phase 2 trains AB/AC, Phase 3 tests BC. The confidence metrics are a nice touch, though under-specified.\n\nThe soft spots are central, not peripheral. Section 3.2 presents a hand-authored generalization schema using variables ($1,$2,$3,$4) that maps any related stimulus pair and any related location pair to a match action. That schema is exactly the same/opposite combinatorial entailment the paper claims to demonstrate. Phase 1 explicitly trains mutual and combinatorial entailment. Section 3.3 adds explicit relational naming with no acquisition mechanism. The example in Section 5.4 is a hand-formulated rule for combining SAME and OPPOSITE. So the BC test is constructed to pass. The paper's own limitation section admits reliance on empirical rather than formal validation, and no code or data are released. The 100% accuracy is therefore not evidence of AARR emerging; it is evidence that a hand-built rule executes.\n\nI'd also flag the absence of error bars or trial-level data; with one system and deterministic behavior, 100% is almost tautological given the schema. The citation pattern is fine, mostly RFT foundations and the authors' own prior NARS work, which is appropriate here.\n\nWho gets value: people working on computational models of RFT or machine psychology may want this as an existence proof that NARS can implement relational frames when the rules are given. It is not a demonstration of derived relational responding in the RFT sense.\n\nMy recommendation: I would not desk-reject this outright, there is a real implementation and a clear write-up, but I'd send it to review only with the expectation of major revision: either release code and show the variable schema is induced from experience, or reframe the paper as 'NARS can execute a hand-coded same/opposite schema,' which is a weaker but honest result. If the venue is a top journal, the overclaim is enough to reject as is.","headline":"The 100% BC accuracy is built into the hand-coded schema, so the emergence claim fails; the paper is a useful NARS implementation but not a demonstration of AARR.","tokens_in":7137,"tokens_out":2774,"would_cite":false,"duration_ms":28393,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that NARS, with its acquired-relations extension, combines mutual and combinatorial entailment to answer never-trained same/opposite matching-to-sample trials at perfect accuracy, thereby exhibiting arbitrarily applicable…","keywords":["same/opposite relational responding","arbitrarily applicable relational responding","Non-Axiomatic Reasoning System","relational frame theory","mutual entailment","combinatorial entailment","matching-to-sample","acquired relations"],"falsifier":"Run the BC test with Phase 1 pretraining omitted while keeping the acquired-relations machinery; if accuracy drops to chance, the symmetric/transitive pretraining carried the generalization. Alternatively, train only A opposite B and A opposite C and test B same C without any prior symmetry or transitivity training; a chance result would show the generalization is not learned from minimal experience alone.","tokens_in":6142,"feed_emoji":"🧠","tokens_out":7828,"duration_ms":75199,"temperature":0.7,"pith_summary":"The paper aims to show that a reasoning system called NARS can learn same/opposite relations from a few reinforced examples and then respond correctly to stimulus pairs it has never seen. To do this, the authors add an 'acquired relations' mechanism that abstracts trained pairings into higher-order rules linking stimulus identities to screen locations. In a context-controlled matching-to-sample design, after explicit pretraining of symmetric and transitive relations, the system scores 100% on the critical test phase, including deriving a same-relation from two trained opposite-relations. If the claim holds, it provides a computational demonstration of relational generalization of the kind psychologists attribute to human symbolic cognition, and a path toward more flexible reasoning in artificial intelligence.","feed_headline":"NARS hits 100% on never-trained relational pairs","feed_subtitle":"A reasoning system derives same/opposite rules for unseen stimuli, a step toward human-like symbolic cognition.","key_machinery":"The load-bearing component is the acquired relation, a higher-order NARS implication that abstracts a trained relation between two stimulus identities together with a spatial relation between two locations into a rule for choosing a match action. In its generalized form it uses variables ($v_1,v_2,v_3,v_4$) and states: if $v_1$ and $v_2$ stand in a learned relation and $v_3$ and $v_4$ stand in a location-pair relation, then matching $v_3$'s stimulus to $v_4$'s stimulus leads to the goal. Paired with explicit relational naming (e.g., `<(X1 * Y1) --> SAME>`), this lets one trained rule apply to arbitrary novel stimuli and locations. Phase 1 pretraining supplies the symmetrical and transitive inference patterns, and Phase 2 maps them onto the AB/AC stimulus network, so the Phase 3 BC test can be passed without any direct BC feedback.","core_discovery":"On the paper's own terms, the central discovery is that NARS can generalize explicitly trained same/opposite relations to entirely new stimulus pairs by combining mutual entailment (symmetry) and combinatorial entailment (transitivity). In the critical BC phase, pairs such as B and C had never been reinforced together, yet NARS selected the correct SAME or OPPOSITE match on 100% of trials, using internal relational hypotheses that connect acquired relations between stimulus identities to matching actions between locations. The system's internal confidence for both mutual and combinatorial entailment rises with training and stays high during novel testing, which the authors read as evidence that the relational principles have been internalized rather than applied as isolated memorized pairings.","pith_inferences":["Editorial inference: The experiment does not isolate emergence from execution, because Phase 1 explicitly trains symmetry and transitivity and Section 3.2 provides a hand-written generalized schema; a stronger test would run BC with both removed to see whether the system can learn the rule from the AB/AC reinforcement alone.","Editorial inference: If the hand-authored schema is considered part of the model rather than an experimental prompt, the contribution is better described as showing that NARS can execute relational-logic rules in a matching task, not that relational frames emerge from experience.","Editorial inference: A natural extension would be to train only OPPOSITE pairs (A opposite B, A opposite C) and test B same C with no preceding symmetric/transitive pretraining; the paper's design already includes this logic, but an ablation would make the derivation path visible.","Editorial inference: The perfect deterministic accuracy suggests that a more demanding test, such as noisy or conflicting context cues, or novel relational frames like 'larger/smaller,' would better discriminate learned relational abstraction from rule-following."],"forward_implications":["NARS achieves 100% accuracy on BC trials without feedback, which is above the 50% chance baseline and therefore demonstrates generalized rather than random responding.","Because the same acquired-relations mechanism is domain-general over stimulus identities and locations, it can be applied to new stimulus sets without retraining the relational rules.","The internal confidence measures for mutual and combinatorial entailment rise with training and stay high in the derived test phase, supporting the authors' claim that the relational principles are internalized.","The approach extends prior NARS stimulus-equivalence work by handling two relational frames (SAME and OPPOSITE) and their combination, not just equivalence."],"supporting_citations":[{"why":"Defines arbitrarily applicable relational responding, mutual entailment, combinatorial entailment, and the matching-to-sample paradigm that the study operationalizes.","marker":"[2]"},{"why":"Supplies the Non-Axiomatic Reasoning System and its logic, the architecture in which the acquired-relations extension is implemented.","marker":"[7]"},{"why":"Prior demonstration of stimulus equivalence in NARS that this work extends to same/opposite relational frames.","marker":"[5]"},{"why":"Describes the Narsese temporal encoding of operant conditioning and matching behavior used to represent MTS trials.","marker":"[4]"}],"fun_headline_variants":["NARS generalizes same/opposite to unseen stimuli","NARS aces novel relational pairs after minimal training","Arbitrary same/opposite reasoning emerges in NARS","NARS derives novel relational rules with 100% accuracy","NARS: novel same/opposite pairs, 100% correct"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on the assumption that the hand-written generalized schema and the Phase 1 symmetric/transitive pretraining do not already encode the same/opposite generalization; if they do, the perfect BC score shows only that the programmed rule executes, not that the relational response emerged from experience.","fun_headline_variants_meta":{"raw":{"variants":["NARS generalizes same/opposite to unseen stimuli","NARS aces novel relational pairs after minimal training","Arbitrary same/opposite reasoning emerges in NARS","NARS derives novel relational rules with 100% accuracy","NARS: novel same/opposite pairs, 100% correct"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001607,"raw_usage":{"total_tokens":6382,"prompt_tokens":909,"completion_tokens":5473,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":5389}},"tokens_in":525,"tokens_out":5473,"duration_ms":37977,"temperature":1.0,"reasoning_tokens":5389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:24:53.988249+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the BC test with Phase 1 pretraining omitted while keeping the acquired-relations machinery; if accuracy drops to chance, the symmetric/transitive pretraining carried the generalization. Alternatively, train only A opposite B and A opposite C and test B same C without any prior symmetry or transitivity training; a chance result would show the generalization is not learned from minimal experience alone.","supporting_citations":[{"cited_title":"Kluwer Academic/Plenum Publishers (2001)","cited_arxiv_id":null,"evidence_quote":"Defines arbitrarily applicable relational responding, mutual entailment, combinatorial entailment, and the matching-to-sample paradigm that the study operationalizes."},{"cited_title":"World Scientific (2013)","cited_arxiv_id":null,"evidence_quote":"Supplies the Non-Axiomatic Reasoning System and its logic, the architecture in which the acquired-relations extension is implemented."},{"cited_title":"In: International Conference on Artificial General Intelligence","cited_arxiv_id":null,"evidence_quote":"Prior demonstration of stimulus equivalence in NARS that this work extends to same/opposite relational frames."},{"cited_title":"Frontiers in Robotics and AI 11, 1440631 (2024)","cited_arxiv_id":null,"evidence_quote":"Describes the Narsese temporal encoding of operant conditioning and matching behavior used to represent MTS trials."}],"review_version":1}