{"id":"7e0cd881-8242-4e21-aa8c-d3a69e81d5bd","arxiv_id":"2506.15468","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Human-AI pairs using a Metropolis-Hastings acceptance rule in a joint-attention naming game achieved better categorization and shared sign convergence than pairs with always-accept or always-reject AI agents.","lead":"A study tested whether humans and AI can jointly learn a shared category system by playing a naming game where each sees only part of the same images. Pairs using a Bayesian-inspired acceptance rule improved joint categorization and converged on shared labels, suggesting a path toward AI that learns with humans rather than from them.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The human acceptance alignment is computed from Inter-GM parameters fit to the same interaction stream the accept/reject decisions shaped; without an out-of-sample test, the r_MH correlation does not independently verify MHNG-based co-creative learning.","rationale":"The reader's conditional verdict is appropriate. The paper has real strengths: the MHNG framework is theoretically grounded, the JA-NG is a clean partial-observability task, and the MH-vs-AA/AR differences in computer ARI and sign agreement are reported with inferential statistics. However, the central theoretical claim is not just that stochastic acceptance helps; it is that the human-AI dyad as a whole is running MHNG-based decentralized Bayesian inference. That requires evidence that human decisions are governed by MH probabilities. The current evidence is generated from a model fit to the same data that those decisions shaped. This is a classic in-sample circularity: the fitted Inter-GM will tend to make the human's own labeling behavior look MH-consistent regardless of whether the human is performing MH sampling. The out-of-sample test would settle this. The post-hoc exclusion and lack of a random-accept control are additional concerns, but they are secondary; the in-sample acceptance analysis is the load-bearing point because it is the only direct evidence for the mechanism. Therefore I do not change the reader's verdict: conditional acceptance pending the test.","tokens_in":15792,"tokens_out":7885,"duration_ms":83594,"concrete_test":"Run a held-out prediction check on the acceptance model. For each participant, fit Inter-GM parameters (θ, φ) using only the first 10 rounds (100 interactions) of that participant's log. For each listener decision in the remaining 10 rounds, freeze those parameters, take the participant's logged current category and current sign at that moment, compute r_MH from Eq. (2), and score the constrained linear model's predicted acceptance probability against the actual accept/reject choice (report AUC and calibration, separately per participant and pooled). As a control, compare against a non-Bayesian baseline that predicts acceptance from whether the proposed sign matches the participant's current category label. If held-out AUC is near chance or no better than the label-match baseline, the Fig.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's mechanism claim—that the dyad is performing decentralized Bayesian inference, not merely exchanging labels—rests on the reported correlation between human accept/reject decisions and the MH acceptance probability r_MH (Eq. 2; §5, Fig. 4). This correlation is not independent evidence. r_MH is computed from Inter-GM parameters and category assignments inferred by Gibbs sampling from the same interaction stream that the human's accept/reject decisions helped generate (§3.2). The accepted sign sequence, the participant's updated category labels, and the fitted model are therefore jointly determined; a sufficiently flexible Inter-GM will assign high r_MH to proposals consistent with the participant's learned label-category mapping (which is exactly the set the participant tends to accept) and low r_MH to inconsistent proposals. The fitted positive slope \\hat a=0.645 is thus an in-sample description of how well the model reconstructs the participant's own labeling behavior, not an independent measurement that the human is sampling from the MH target. Since the computer's MH rule is imposed by design, this alignment is the only direct support for the claim that the human half of the dyad is performing Bayesian inference. Without out-of-sample evaluation, the headline claim overreaches.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'co-creative learning' as a paradigm in which a human and an AI integrate partial observations through decentralized Bayesian inference, instantiated by the Metropolis-Hastings naming game (MHNG) with an Interpersonal Gaussian Mixture (Inter-GM) model. The authors formalize this as minimization of collective free energy and report an online joint-attention naming game experiment (N=69 after exclusions) with three computer-agent conditions (MH-based, always-accept, always-reject). They report that the MH condition yields significantly higher computer categorization accuracy (ARI) and sign-posterior agreement than controls, and that human accept/reject decisions are positively correlated with the model-derived MH acceptance probability r_MH. The paper interprets this as first empirical evidence for co-creative learning in human-AI dyads.","tokens_in":16006,"tokens_out":5511,"duration_ms":54667,"significance":"If the empirical claims held, the paper would be a meaningful step: it extends experimental semiotics and MHNG from human-human dyads to human-AI dyads, connects the mechanism to collective predictive coding and bidirectional alignment, and releases its data. The theoretical formalization of co-creative learning via collective free energy is useful, and the three-condition design is a sensible way to contrast co-creative, supervised-like, and unsupervised-like interaction. However, the central mechanism evidence is in-sample, the sign-agreement target is model-derived from the same data, and the human-side advantages over the always-accept control are not statistically significant; these issues currently limit the strength of the headline claim.","major_comments":[{"comment":"The human-acceptance alignment is computed in-sample: r_MH uses Inter-GM parameters and category assignments inferred by Gibbs sampling from the same interaction stream on which the human accept/reject decisions were recorded. A flexible model fit to the data will assign high r_MH to proposals consistent with the participant's learned label-category mapping, exactly the proposals the participant tends to accept, so the positive fitted slope a=0.645 does not independently verify that the human is sampling according to the MH target. Please add an out-of-sample evaluation, for example fitting the Inter-GM on the initial categorization or the first half of each participant's interactions and predicting accept/reject decisions in the held-out half, and compare prediction accuracy against simple baselines (e.g., always accept, accept-when-label-matches-current-category). Absent such a test, the claim that the human half of the dyad performs Bayesian inference is not supported.","section":"§3.2, §4, Eq. (2), Fig. 4"},{"comment":"The sign-posterior agreement target is obtained by Gibbs sampling the full Inter-GM model on the same dyad's observations and sign sequences, and the empirical sign distributions being compared are generated by the same model family. High agreement can therefore partly reflect the fitted model's ability to reconstruct the data rather than convergence to the true integrated posterior. Please validate the target using held-out data or an independent target (for instance, posterior over ground-truth categories or a model fit on one half of the interactions), and report agreement against human-only and AI-only targets as a reference.","section":"§4, Evaluation Metrics, Table 1"},{"comment":"The analysis excludes 21 of 90 participants based on 'inactivity or failure to follow instructions,' but the exclusion criteria are not pre-registered or quantified (e.g., no threshold for number of identical responses), and no information is given about excluded participants by condition. Because the MH-vs-AA computer ARI difference is marginal (p=0.046), the results may be sensitive to these exclusions. Please specify the exact exclusion rules, report counts per condition, and provide a sensitivity analysis including all participants.","section":"§4, Participants"},{"comment":"The human-side outcomes do not show a significant MH advantage: human ARI in MH (0.490) is not significantly different from AA (0.473; p=0.799), and human sign agreement in MH (0.729) is not significantly different from AA (0.722; p=0.762). The evidence for 'mutual' integration is therefore confined to the computer agent's metrics. Please either provide analyses demonstrating a human-side benefit (e.g., individual learning curves, within-dyad alignment over time) or temper the claim that the dyad co-creatively learns beyond what the always-accept baseline achieves.","section":"§5, Table 1"}],"minor_comments":[{"comment":"The figure contains Japanese text in its axis labels and title; these should be translated to English for the intended readership.","section":"Fig. 4"},{"comment":"The column layout of Table 1 is difficult to parse, especially the placement of Initial, Final, and the agreement columns; please reformat so that each phase and metric is clearly labeled.","section":"Table 1"},{"comment":"For the Welch t-tests, please report effect sizes and confidence intervals in addition to p-values, particularly for the marginally significant MH-vs-AA comparison.","section":"§5"},{"comment":"The statement that the collective free energy decrease follows from detailed balance via the Data Processing Inequality is loose; please cite the standard Metropolis-Hastings convergence argument or prove the KL decrease directly.","section":"§2.2"},{"comment":"The data link is a Google Drive folder; please provide a stable archival identifier (e.g., a DOI) and include the analysis code to make the in-sample/out-of-sample distinction reproducible.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The abstract and conclusion are considerably stronger than the reported evidence warrants. The central empirical test of the MH mechanism is in-sample, and the human-side metrics do not differ from the always-accept control. I would encourage the editor to require the out-of-sample and sensitivity analyses outlined in the major comments before considering publication; with those changes, the paper could be a solid contribution. The manuscript draws heavily on the authors' own prior work (Okumura et al., 2023; Taniguchi et al.), but the human-AI extension and the co-creative learning framing are sufficiently novel for this venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuinely new empirical result, but the headline mechanism claim is not yet supported by the evidence. The behavioral contrast between conditions is the solid part; the human-alignment analysis is an in-sample fit and should be presented as descriptive, not confirmatory.\n\nWhat's new: first test of MHNG with human-AI pairs in a joint attention naming game. 69 online participants, three computer agent strategies. The MH-based agent achieves higher final ARI and sign-posterior agreement than always-accept or always-reject controls, especially on the computer side, and those differences are significant. That is a useful data point for designing interactive learning systems. The theory section gives a clean formal definition of co-creative learning as collective free-energy minimization; the framing is standard FEP/CPC but clearly stated.\n\nSoft spots, in increasing order of concern. First, participant exclusion: 90 recruited, 69 analyzed, with no per-condition counts or robustness check on the exclusion criteria. Minor, but easy to fix. Second, no no-interaction control, so some of the ARI improvement could be practice effects. Given initial ARI differences across groups, an ANCOVA on initial ARI would help. Third, the target sign-posterior distribution is derived from the same Inter-GM model family fitted to the same dyad data, so the agreement metric is partly model-consistent rather than an independent external criterion. Fourth and most important, the human acceptance alignment: r_MH is computed from Inter-GM parameters inferred from the same interaction stream the accept/reject decisions helped generate. The fitted positive slope is an in-sample description of how well the model reconstructs the participant's labeling behavior, not independent evidence that the human is sampling from the MH target. The stress-test note has it right. You'd need out-of-sample prediction or at least a bootstrap/held-out evaluation, and ideally a model comparison against a simpler heuristic (e.g., accept if proposal matches current label) to show r_MH adds predictive value.\n\nThe citation pattern is fine: prior human-human and computational work by the same group is clearly cited, and the stated novelty is appropriately limited to the human-AI dyad. Data are available at a public link, which is good.\n\nWho this is for: people working on emergent communication, human-in-the-loop learning, and experimental semiotics. The behavioral data are worth engaging with, but the paper's central theoretical claim currently overreaches its evidence.\n\nRecommendation: send it to peer review as a major revision. The behavioral experiment deserves referee time, and the authors should be asked to address the in-sample circularity and add a no-interaction control or ANCOVA.","headline":"Useful empirical extension of MHNG to human-AI dyads, but the mechanism claim rests on an in-sample fit; the behavioral contrast between conditions is the real contribution.","tokens_in":16555,"tokens_out":3097,"would_cite":true,"duration_ms":31854,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Accepting and rejecting names like a Bayesian sampler lets a human and an AI build a shared category system.","keywords":["co-creative learning","human-AI interaction","symbol emergence","Metropolis-Hastings naming game","joint attention naming game","decentralized Bayesian inference","AI alignment","human-in-the-loop machine learning"],"falsifier":"Run the same partial-observation naming game but ask participants to categorize held-out novel stimuli or rate pairwise similarities independently before, during, and after the interaction; if their accept decisions track the MH acceptance probability computed from those independently elicited beliefs as well as they track the probability computed from the game-fitted model, the finding is genuine, while a mismatch would show the reported alignment is specific to the fitted model rather than to human psychology.","tokens_in":15573,"feed_emoji":"🧩","tokens_out":10117,"duration_ms":97042,"temperature":0.7,"pith_summary":"This paper argues that human-AI collaboration can work as co-creative learning: two agents, each seeing only part of the same objects, integrate what they have by exchanging names and accepting or rejecting proposals according to the Metropolis-Hastings rule. The claim is that this simple game makes the pair behave like one decentralized Bayesian system, converging on categories that reflect both partners' partial views. An online experiment with 69 participants paired people with an MH-rule computer agent, an always-accept agent, or an always-reject agent; the MH agent ended with the highest categorization accuracy and closest convergence to a shared sign system, and participants' accept decisions tracked the theoretical MH acceptance probability. If correct, this gives a concrete, testable mechanism for AI that learns with people instead of merely being taught by them, relevant to building shared representations and aligning AI to human perception.","feed_headline":"Naming game lets humans and AI build shared categories","feed_subtitle":"In a 69-person study, a Bayesian accept/reject rule outperformed always-agree and always-reject partners.","key_machinery":"The load-bearing mechanism is the Metropolis-Hastings naming game (MHNG), a protocol in which a listener accepts a speaker's proposed name with probability $r_{\\mathrm{MH}} = \\min(1, P(c_{\\mathrm{Li}}|\\theta_{\\mathrm{Li}}, s^*)/P(c_{\\mathrm{Li}}|\\theta_{\\mathrm{Li}}, s_{\\mathrm{Li}}))$, the likelihood ratio of the listener's own inferred category under the proposed sign versus its current sign. The game converts a conversation into a distributed Metropolis-Hastings sampler with the joint posterior as its target, so local accept/reject choices implement collective Bayesian inference without any agent revealing private observations or gradients. To measure this in humans, the paper analyzes interaction logs with the Inter-GM generative model, in which each agent's observations are Gaussian mixtures over latent categories linked by shared signs, infers each participant's parameters and category assignments by Gibbs sampling, and uses the fitted model both to compute the acceptance probability and to construct the target sign posterior. The three experimental conditions instantiate three learning paradigms: always-accept as supervised learning, always-reject as unsupervised learning, and the MH rule as co-creative learning.","core_discovery":"The central discovery the paper reports is that a human and a computer agent playing a joint attention naming game under partial observability achieve co-creative learning: the dyad's accepted names follow a Metropolis-Hastings update rule, so the interaction is a distributed Bayesian sampler whose stationary distribution is the posterior over shared signs conditioned on the union of both agents' observations. In the experiment, the computer agent using the MH rule reached final categorization ARI of 0.609, significantly above 0.469 for the always-accept condition and 0.404 for the always-reject condition, and its final sign distribution agreed with the combined-observation posterior at 0.765, against 0.717 and 0.469 respectively. Human accept/reject behavior was also well described by the theoretical acceptance probability, with a fitted sensitivity parameter of 0.645. The paper interprets these results as the first empirical evidence that human-AI interaction of this kind performs decentralized Bayesian inference through the MHNG mechanism.","pith_inferences":["The argument should extend from dyads to larger teams: because the MHNG update uses only local messages, a network of humans and AIs with different sensors could in principle converge to one shared naming scheme without any central coordinator. The paper does not test this.","A sharper, independently grounded test would elicit participants' category beliefs outside the naming game, for example by asking them to sort held-out novel stimuli before and after the session, and then check whether the MH acceptance curve predicts those independent judgments; the current experiment derives the target posterior from the same model family it uses to score participants.","One can derive a quantitative prediction for the acceptance curve: shifting category overlap or prior uncertainty in the stimulus generation should shift the fitted slope and baseline in a specific direction, which would distinguish genuine Bayesian updating from a generic tendency to accept plausible names."],"forward_implications":["Dyads using the MH rule converge to a shared sign system that encodes information from both sensors, so the computer's categories can improve beyond what the human's labels alone or its own observations alone would support.","The acceptance rule is not a neutral interface: an AI that accepts or rejects proposals probabilistically shapes human categorization as well, making the interaction a two-way influence rather than one-way teaching.","Because the mechanism never requires sharing raw observations or model gradients, co-creative learning of this kind could operate in privacy-sensitive settings where centralized training is impossible.","Human acceptance behavior being predictable from the MH probability suggests that the same algorithm could be used to design AI agents that communicate in ways people find natural to accept or reject."],"supporting_citations":[{"why":"Supplies the JA-NG experimental paradigm, the Inter-GM model, and the earlier human-human finding that acceptance behavior aligns with MHNG probabilities, which this study extends to human-AI dyads.","marker":"Okumura et al. [2023]"},{"why":"Introduced the Metropolis-Hastings naming game as a decentralized Bayesian inference mechanism, providing the algorithmic basis for the experimental condition and the theoretical prediction.","marker":"Hagiwara et al. [2019]"},{"why":"Extends MHNG to deep generative models and formalizes decentralized inference through communication, supporting the claim that interactions converge to the joint posterior.","marker":"Taniguchi et al. [2023b]"},{"why":"States the Collective Predictive Coding hypothesis that shared symbol systems emerge through decentralized Bayesian inference, the theoretical framework the experiment tests.","marker":"Taniguchi [2024]"},{"why":"Formalizes collective free energy and its decomposition, which the paper uses to define co-creative learning and its convergence criterion.","marker":"Taniguchi et al. [2025]"},{"why":"Supplies the information-theoretic basis, detailed balance and the data-processing inequality, for the claim that the collective free energy decreases under MHNG interaction.","marker":"Cover and Thomas [2006]"},{"why":"Provides the Adjusted Rand Index used to measure categorization accuracy against ground truth in all three experimental conditions.","marker":"Hubert and Arabie [1985]"},{"why":"Provides the Hungarian-algorithm matching procedure used to compute sign-posterior agreement between empirical and target sign distributions.","marker":"Inukai et al. [2023]"}],"fun_headline_variants":["Co-creative learning emerges from human-AI naming game","Bayesian accept/reject rule helps human-AI pairs share categories","Human-AI dyads co-create signs via Metropolis-Hastings interchange","Naming game with Bayesian partners yields shared sign systems","Metropolis-Hastings turns human-AI naming into co-creation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the statistical category model fitted to each participant's interaction log faithfully reproduces how that person actually categorized the stimuli, since both the acceptance probability and the target sign distribution come from that same fitted model.","fun_headline_variants_meta":{"raw":{"variants":["Co-creative learning emerges from human-AI naming game","Bayesian accept/reject rule helps human-AI pairs share categories","Human-AI dyads co-create signs via Metropolis-Hastings interchange","Naming game with Bayesian partners yields shared sign systems","Metropolis-Hastings turns human-AI naming into co-creation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000614,"raw_usage":{"total_tokens":2861,"prompt_tokens":961,"completion_tokens":1900,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":1811}},"tokens_in":577,"tokens_out":1900,"duration_ms":14293,"temperature":1.0,"reasoning_tokens":1811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:33:40.451657+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same partial-observation naming game but ask participants to categorize held-out novel stimuli or rate pairwise similarities independently before, during, and after the interaction; if their accept decisions track the MH acceptance probability computed from those independently elicited beliefs as well as they track the probability computed from the game-fitted model, the finding is genuine, while a mismatch would show the reported alignment is specific to the fitted model rather than to human psychology.","supporting_citations":[],"review_version":2}