{"id":"1a692b0f-579c-4938-bcc3-6bc99da4003e","arxiv_id":"2509.03736","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM agents produce survey responses that sometimes match human patterns but show major inconsistencies when their behavior is checked across isolated questions and group conversations.","lead":"This paper tests if LLM agents keep consistent 'latent profiles' from survey questions when they interact in conversations with other agents. Smart generalists should read it because it directly challenges whether AI can reliably stand in for real humans in social science experiments.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Unvalidated mapping from latent profile questions to expected conversational behaviors","rationale":"The reader's weakest_assumption already isolates the profile-to-behavior mapping as the critical untested link; this remains the single most load-bearing point because the headline finding of 'failure to be empirically consistent' is only as strong as that mapping. No other internal inconsistency or missing control appears more central from the abstract and described design.","tokens_in":1700,"tokens_out":261,"duration_ms":23071,"concrete_test":"Run the full protocol (profile-revealing questions followed by the conversational task) with human participants as a control; if humans exhibit comparable mismatches between profile and observed behavior, the mapping is unreliable and the agent incoherence conclusion does not follow.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim of empirical inconsistency requires that any deviation between an agent's revealed profile (from the question set) and its conversational actions is attributable to agent incoherence rather than error in the profile-to-behavior mapping. The paper invokes 'standard behavioral hypotheses' to generate these predictions, but without independent validation that the chosen questions produce stable, predictive profiles for human conversational behavior in the exact multi-agent setup, mismatches could reflect flawed or underspecified hypotheses instead of LLM limitations.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper develops a two-stage design to test behavioral coherence in LLM agents for social simulation: (a) a question set to elicit an agent's latent profile, and (b) observation of the agent's conversational behavior in multi-agent interactions, evaluated against predictions generated from standard behavioral hypotheses. It reports significant inconsistencies across model families and sizes, concluding that LLMs fail to maintain empirical consistency and thus cannot reliably substitute for human participants despite matching isolated survey responses.","tokens_in":1791,"tokens_out":395,"duration_ms":26804,"significance":"If the reported inconsistencies prove robust, the work identifies a substantive limitation in current LLM agents for human-subject research and social simulation, shifting focus from response alignment to cross-context coherence. The empirical, hypothesis-driven approach is a positive feature that could inform more reliable agent architectures.","major_comments":[{"comment":"The central claim of empirical inconsistency rests on the assumption that deviations between revealed latent profiles and observed conversational actions are attributable to agent incoherence rather than error in the profile-to-behavior mapping. The manuscript invokes 'standard behavioral hypotheses' to generate predictions but provides no independent validation that the chosen questions produce stable, predictive profiles for human conversational behavior in the exact multi-agent setup used here. This is load-bearing for the inconsistency conclusion.","section":"Study Design and Behavioral Hypotheses"}],"minor_comments":[{"comment":"The abstract and methods summary omit key details on sample sizes, statistical controls, exact question wording, and how conversational outcomes were coded or scored; these should be added for reproducibility.","section":"Abstract and Methods"},{"comment":"Clarify whether the same agents participate in both the profile-elicitation and conversational phases or whether fresh instances are used, as this affects interpretation of within-agent consistency.","section":"Experimental Procedure"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive review. The comment raises an important point about the foundational assumptions in our study design, and we address it directly below.","responses":[{"response":"We agree this is a substantive concern. The question battery draws from established instruments in social psychology and behavioral economics (e.g., scales for trust, cooperation, and personality traits commonly used in prior human-subject studies), and the behavioral hypotheses are drawn from standard predictions in the literature on social dilemmas and group interaction. Nevertheless, we did not run a parallel human experiment using the identical multi-agent conversational protocol to re-validate predictive accuracy in this specific setup. In the revised manuscript we will (1) add an explicit subsection in the Methods and Discussion that states the sources of each question and hypothesis, (2) clarify that the mapping is treated as a benchmark drawn from the existing literature rather than newly validated here, and (3) acknowledge this as a limitation while outlining how future human validation studies could be conducted. These changes will make the assumptions more transparent without requiring new data collection for the current paper.","revision_made":"partial","referee_comment":"The central claim of empirical inconsistency rests on the assumption that deviations between revealed latent profiles and observed conversational actions are attributable to agent incoherence rather than error in the profile-to-behavior mapping. The manuscript invokes 'standard behavioral hypotheses' to generate predictions but provides no independent validation that the chosen questions produce stable, predictive profiles for human conversational behavior in the exact multi-agent setup used here. This is load-bearing for the inconsistency conclusion."}],"tokens_in":1261,"tokens_out":340,"duration_ms":37459,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper tests LLM agents for behavioral coherence by first pulling a latent profile from a set of questions and then checking whether their actions in multi-agent conversations line up with what standard behavioral hypotheses would predict from that profile. They find significant mismatches across model families and sizes, even in cases where survey-style responses match human patterns. This suggests agents may not be reliable substitutes for real participants in social simulations or human-subject research setups.","headline":"LLM agents show inconsistencies between elicited profiles and conversational behavior, but the profile-to-behavior mapping lacks independent validation.","tokens_in":2277,"tokens_out":156,"would_cite":false,"duration_ms":33219,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"LLM agent consistency testing in social simulation has no overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper's central machinery (latent-profile elicitation via preference/openness questions, agent pairing for dialogue, bootstrap sampling of agreement scores, and six statistical tests for behavioral coherence) operates entirely within empirical AI evaluation. It invokes no recognition cost J, golden-ratio ladder, 8-tick periodicity, or any parameter-free derivation of constants. RS theorems such as reality_from_one_distinction, absolute_floor_iff_bare_distinguishability, and alexander_duality_circle_linking are irrelevant; the paper neither echoes nor contradicts them.","tokens_in":57139,"confidence":"high","tokens_out":153,"duration_ms":5790,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLM agents can match human survey answers yet fail to act consistently with their own revealed profiles in conversations.","keywords":["LLM agents","behavioral consistency","latent profiles","social simulation","empirical consistency","human substitution","conversational behavior","multi-agent interaction"],"falsifier":"Agents whose conversational actions reliably follow the predictions derived from their profile-revealing answers across repeated trials and settings would contradict the reported inconsistency.","tokens_in":2617,"feed_emoji":"🤖","tokens_out":580,"duration_ms":36763,"temperature":0.7,"pith_summary":"The paper tests whether large language model agents can stand in for real people in human-subject research. It first asks agents a set of questions to uncover their latent profile, then places them in conversations with other agents to see if their behavior follows what that profile should predict. Standard behavioral hypotheses supply the expected links between profile and action. The results show clear mismatches across model families and sizes. Matching isolated answers does not guarantee that the agent will behave as its profile requires when the setting changes.","feed_headline":"LLM agents match answers but break profile consistency in talk","feed_subtitle":"Revealed latent profiles fail to guide their conversational actions as expected, limiting reliable social simulation.","key_machinery":"A two-stage design that elicits a latent profile through targeted questions and then measures whether subsequent conversational turns conform to the behavioral implications of that profile.","core_discovery":"LLM agents produce responses that align with human counterparts on profile-revealing questions, yet their conversational behavior in multi-agent settings deviates from the predictions that standard behavioral hypotheses would draw from those same profiles, indicating a failure of empirical consistency.","pith_inferences":["Future work could test whether fine-tuning on paired profile-and-behavior data reduces the observed gaps.","The same design might be applied to other domains such as economic games or policy deliberation to check generalizability.","If profile-to-behavior links prove unstable for current models, simulation studies may need hybrid human-agent setups rather than full replacement."],"forward_implications":["Individual response matching is insufficient to establish agents as reliable stand-ins for human participants.","Inconsistencies appear across different model families and sizes, limiting substitution claims.","Social simulation experiments that rely on behavioral coherence will require additional validation steps.","Prompting alone does not produce agents whose actions remain predictable from their stated profiles."],"fun_headline_variants":["LLM agents match on profiles yet chats contradict","Revealed profiles don't shape LLM conversations","Behavioral tests expose LLM agent inconsistencies","LLMs align in queries but falter in interactions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The selected questions reveal a stable latent profile whose implications for conversational behavior are accurately predicted by ordinary behavioral hypotheses, so that any mismatch must be blamed on the agent rather than on the profile-to-behavior mapping.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents match on profiles yet chats contradict","Revealed profiles don't shape LLM conversations","Behavioral tests expose LLM agent inconsistencies","LLMs align in queries but falter in interactions"]},"model":"grok-4.3","cost_usd":0.00873,"raw_usage":{"total_tokens":3825,"prompt_tokens":612,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":87303000,"prompt_tokens_details":{"text_tokens":612,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3159,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":612,"tokens_out":54,"duration_ms":15314,"temperature":1.0,"reasoning_tokens":3159,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-18T18:40:57.719166+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Agents whose conversational actions reliably follow the predictions derived from their profile-revealing answers across repeated trials and settings would contradict the reported inconsistency.","supporting_citations":[],"review_version":1}