{"id":"edc79aee-9824-43cd-a952-6ea50b9e0488","arxiv_id":"2507.09657","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLM-based generative agents simulate personality-driven temperature negotiations in a shared residential building, with positive personality traits associated with higher happiness and stronger friendships.","lead":"Researchers used LLM-powered agents to simulate how personality traits shape heating temperature decisions in a shared apartment building. The simulations suggest that positive traits are linked to higher happiness and stronger friendships, showing how generative agents could model human social behavior in building energy settings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline correlations rest on single stochastic trajectories and unclustered p-values; rerunning with seeds and clustered errors could overturn them.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the inference treats within-run daily observations as independent and relies on single simulation trajectories. My reading of the manuscript confirms this is the most serious threat to the central claim. The paper's contribution is partly methodological, and the simulation design is described in enough detail to be replicable in principle, but the headline findings are framed as empirical correlations with significance levels. Those significance levels are not credible under the current design. I also considered the 'prompt steering' concern, since agents are explicitly told to let traits influence decisions; that is a real limitation on external validity, but it is less decisive than the statistical issue because the paper's stated goal is to demonstrate a simulation methodology, not to discover real-world personality laws. If multi-seed replication and properly clustered inference preserve the qualitative differences, the conditional acceptance would be justified; if not, the empirical claims would need to be weakened. The lack of released code and data compounds the problem but is secondary to the statistical validity of the p-values. Therefore I do not move the reader's conditional verdict; the concern strengthens the need for the stated conditions rather than changing the verdict direction.","tokens_in":15325,"tokens_out":3560,"duration_ms":48348,"concrete_test":"Run each of the three trait-distribution settings at least 20 times with different random seeds and LLM sampling seeds, keeping prompts and architecture fixed. For the network-level claim, fit a mixed model with run-level random effects or cluster by run, and report between-run standard deviations and 95% confidence intervals for the differences in average happiness and friendship weight between all-positive and all-negative settings at day 15 and day 30. For the node-level claim, recompute Tables 4 and 5 with two-way cluster-robust standard errors clustered by agent and by day, pooling across runs and including run as a random or fixed effect. If the point estimates shrink materially or p-values for assertiveness, selflessness, and temperature preferences cross 0.05, the central correlations are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that positive personality traits correlate with higher happiness and stronger friendships, and that assertiveness, selflessness, and temperature preferences significantly predict happiness and decisions, depends on statistical analyses whose independence assumptions are not met. Section 4.1 Table 2 reports OLS slopes over 15 days per setting, apparently from one simulation run per trait distribution. Section 4.2 builds a CRE model on 3480 node-day observations from one 30-day run in the 50% positive setting. The simulation is stochastic: LLM generation, agent ordering, and initial random family sizes all vary, yet no multiple seeds are used. The observations are also serially correlated through shared weather, the building-wide temperature, friends' last three degree choices, and happiness updates. OLS standard errors and the CRE p-values treat these as independent, so significance levels are overstated. With only 15 network-level observations, even a strong autocorrelated trend can produce p<0.001 without any true trait effect. At the node level, errors are correlated within agents, within days, and between friends, so the reported standard errors in Tables 4 and 5 are likely too small. The authors themselves acknowledge in Section 5 that the setups are limited, but that admission does not repair the inference. If these effects vanish once clustering and multiple seeds are introduced, the abstract's correlational claims are not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents an LLM-based generative-agent simulation of a shared residential building. Thirty-four households are derived from Zachary's Karate Club network; each day family members propose a heating temperature, family representatives aggregate these proposals, and representatives then vote in a building-wide poll. Agents are assigned personality adjectives from Big Five facets, temperature preferences, and evolving friendship closeness weights, and their decisions are produced by a quantized Mistral 7B model. The paper runs simulations under three trait distributions (all positive, 50% positive, all negative) and analyzes the outputs with network-level OLS regressions and node-level Correlated Random Effects models. The headline claims are that positive trait distributions correlate with higher average happiness and stronger friendships, and that assertiveness, selflessness, and temperature preferences significantly predict individual happiness and degree choices.","tokens_in":15558,"tokens_out":5344,"duration_ms":63406,"significance":"If the statistical claims were properly supported, the paper would make a useful methodological contribution: it demonstrates an end-to-end integration of generative agents with a social-network simulation framework (Crowd), including prompt design, JSON parsing, and multi-day data collection. The example prompts and outputs in Figures 6-9 are clear and the observed building-level temperature convergence to the neutral range is an interesting emergent pattern. However, the paper's central empirical findings rest on inference from single stochastic runs and on standard errors that ignore the data's dependence structure. The methodological demonstration is credible, but the statistical evidence for the personality-correlation claims is not yet established. The authors do not provide code or data artifacts, so reproducibility cannot be independently checked from the manuscript alone.","major_comments":[{"comment":"The network-level trends in Table 2 are OLS slopes over 15 daily observations from a single simulation run per trait setting, yet the simulation is stochastic in at least the random assignment of family sizes and the LLM decoding process. With no multiple seeds, the reported p-values only summarize how well a line fits one autocorrelated trajectory; they cannot support the claim that the all-positive setting produces increasing friendship weight while the all-negative setting declines, because those differences could be run-specific. A concrete fix is to rerun each setting with several seeds, report the distribution of slopes, and use run-level cluster-robust standard errors or a mixed model with run random effects.","section":"Section 4.1, Table 2"},{"comment":"The CRE models use 3480 node-day observations from one 30-day run in the 50% positive setting. The standard errors treat each node-day as independent, but residuals are correlated within agents over time (happiness carries over from previous days), within days (shared weather and a single building temperature), and between friends whose choices are fed into each other's prompts. With only one run there is no way to separate the trait effect from the particular stochastic trajectory. I would ask for cluster-robust standard errors on agent and day, or a two-way cluster, plus multiple seeds, before the significant coefficients in Tables 4 and 5 can be interpreted as evidence about personality effects.","section":"Section 4.2, Tables 4 and 5"},{"comment":"The prompts explicitly ask the LLM to 'mention how your traits influence your decision' and list the agent's personality adjectives. The regression estimates for assertiveness and selflessness in Tables 4 and 5 therefore partly reflect the experimental manipulation, because the trait labels are provided as inputs, rather than necessarily an emergent social process. This is a validity concern rather than a fatal flaw, but the paper should either be reframed as a test of prompt-driven trait effects or supplemented with an ablation in which trait labels are withheld, or with a rule-based baseline, to show that the observed associations are not simply the LLM following the instruction to use the listed traits.","section":"Section 3.5, Figures 6 and 8; Section 4.2"},{"comment":"The pooled regression in Table 3 stacks 30 days from the 50% positive run with 15 days from each of the other two runs, giving 60 observations from three settings. These observations are not independent: each setting contributes a single autocorrelated time series, and the regressor 'average friendship weight' is an outcome of the same simulation dynamics as happiness. The reported p-values therefore do not provide a valid basis for the claim that friendship weight causes happiness. A multi-level model with run-level random effects, or at least Newey-West standard errors treating each run as a cluster, would be needed.","section":"Section 4.1, Table 3"}],"minor_comments":[{"comment":"The row label 'Alturism' should be 'Altruism'.","section":"Table 1"},{"comment":"The footnote contains the typo 'p =< 0.001'; it should read 'p < 0.001'.","section":"Table 5"},{"comment":"The sentence 'The mean happiness over the 15 iterations' should say '15 days' or '15 daily observations' to match the rest of the section.","section":"Section 4.1"},{"comment":"The paper does not explain why the all-positive and all-negative settings were run for only 15 days while the 50% positive setting ran for 30 days; this asymmetry should be stated, and its implications for the comparability of the slopes in Table 2 should be discussed.","section":"Section 4"},{"comment":"Because C is set to 1, the 'cost' variable is not in meaningful physical or monetary units; a brief statement that the cost is an ordinal proxy for energy use would improve interpretability.","section":"Section 3.4, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper is better framed as a demonstration of an LLM-agent simulation methodology than as a statistical finding about personality effects. The authors' own Section 5 admission that the setups are limited does not repair the inference problems in Section 4. I do not see a novelty disclosure issue, but the novelty claim should focus on the Crowd integration and the end-to-end simulation, not on the regression results as currently presented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper is a methodology demonstration, not an empirical finding. The genuinely new thing is the wiring: generative LLM agents placed in an explicit weighted social network (Karate Club as friendship graph, families as cliques) and asked each day to vote on a central heating setpoint. The two-stage family/building poll, the prompt design with adjective-based Big Five facets, and the JSON parsing for closeness updates are concrete and reusable. I learned something from how they calibrated the prompt to avoid absurd degree outputs. The Crowd framework example is a nice bonus.\n\nWhat the paper does not establish is the headline: 'positive traits correlate with higher happiness and stronger friendships.' Table 2 is OLS slopes over 15 days from one run per trait distribution. The CRE in Table 4 uses 3480 node-day observations from one 30-day run, with no clustering, no multiple seeds, and errors that are certainly correlated within agents, across days, and within friendship dyads. The p-values are overconfident; the stress-test note is right that these effects could vanish with proper replication. There is also a circularity problem: the prompts explicitly tell the agent to mention how traits influence decisions, so the regression of happiness on assertiveness is partly reading back the prompt. The authors admit the setups are limited in Section 5, but the admission does not repair the inference.\n\nThat said, this is not a dishonest paper. The qualitative contrast between the all-positive and all-negative settings is large and visible in the charts, and the mechanism they describe is plausible: all-positive settings build stronger friendships, the building temperature stays in the neutral range, and agents with warm/hot preferences end up less happy. The problem is that plausible is not the same as supported at p<0.001. The absence of released code and data makes independent re-running harder.\n\nWho is this for? Anyone building LLM-based agent simulations for building energy or social dynamics. It is a useful case study in both prompt engineering and in how not to report statistics on a single trajectory. With multi-seed runs, clustered standard errors, and a public artifact, it could become a solid contribution. I would send it to a serious referee, but the referee should treat the statistical claims as unproven and require revision. I would not cite the correlations; I might cite the methodology.","headline":"A genuinely useful LLM-agent methodology demo whose headline statistical claims rest on single runs and prompt-steered traits; worth refereeing for the methods, not for the correlations as established.","tokens_in":16103,"tokens_out":2063,"would_cite":true,"duration_ms":25269,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM-powered generative agents can simulate personality-driven social negotiation in a shared residential building, with positive personalities producing happier, more connected communities.","keywords":["generative agents","large language models","social network simulation","building energy modeling","personality traits","agent-based modeling","happiness","central heating"],"falsifier":"Repeat each of the three settings with at least ten random seeds and compare the distributions of daily average happiness and friendship weights; if the all-positive and all-negative confidence intervals overlap once seed-to-seed variation is included, the paper's central contrast is not supported.","tokens_in":15109,"feed_emoji":"🏢","tokens_out":7291,"duration_ms":77257,"temperature":0.7,"pith_summary":"Generative agents are LLM chatbots asked to behave as specific people, and this paper tries to prove they can carry a social simulation. On each of 30 days, 116 virtual residents in a 34-apartment building first negotiate a temperature with their family, then their representatives negotiate with friends; every agent is prompted with personality adjectives, a heater preference, a happiness score, and the weather. The paper compares three runs in which 100%, 50%, or 0% of assigned trait adjectives are positive, and reports that the all-positive run has higher average happiness and strengthening friendships while the all-negative run shows declining friendships. A sympathetic reader would care because, if true, LLM agents could prototype social and energy decisions that would be costly or intrusive to test with real people.","feed_headline":"Positive personalities make simulated neighbors happier and closer","feed_subtitle":"In a 30-day LLM building simulation, assertive and selfless agents and temperature preferences shape happiness and votes.","key_machinery":"The load-bearing mechanism is the two-stage prompt-engineered decision loop. Each day, every agent receives a text identity card containing its nine personality adjective pairs, heater preference, happiness, and outside temperature; the LLM returns a degree suggestion and a new happiness. Family suggestions are averaged, and each Family Representative then receives friends' closeness levels and recent votes before returning a final building vote plus friendship weight updates. The statistical machinery is the Correlated Random Effects model, a panel-data estimator chosen after a Hausman test rejected random effects, which lets the authors estimate the contribution of time-invariant traits such as assertiveness and temperature preference while pooling within- and between-agent variation.","core_discovery":"On its own terms, this paper establishes that LLM-powered generative agents can simulate personality-driven negotiation in a shared residential building and that the resulting trends are statistically measurable. In a 30-day, 116-agent simulation built on a real-world social network, the all-positive personality run averaged 91.01 happiness against 85.99 for the all-negative run, with friendship weight rising at 0.0298 per day in the positive run and falling at -0.0295 in the negative run. Node-level regressions on 3,480 agent-days find that each extra degree an agent suggests adds 0.2468 happiness points, that assertiveness raises happiness by 3.2413 points, and that selflessness lowers it by 2.6886 points, while warm and hot temperature preferences cost 7.7128 and 13.584 happiness points respectively because the building settles near 21-22°C. The paper's claim is that these effects make personality and preference meaningful drivers of the emergent social outcomes.","pith_inferences":["Editorial inference: the universal 21-22°C convergence may reflect anchoring to the 'Neutral' range in the prompt's reference table as much as genuine emergent consensus; a direct test would shift or remove that reference anchor.","Editorial inference: because the node-level analysis uses one 30-day run and treats agent-days as independent observations, the reported p-values are likely optimistic; multiple seeds with clustered standard errors could change the significance profile.","Editorial inference: the same family-then-building polling design could be applied to other shared-resource negotiations, such as water, electricity, or common-space decisions, to see whether personality effects generalize across domains."],"forward_implications":["All-positive trait settings end with higher average happiness, around 91.0 versus 86.0 for all-negative, and rising friendship weights, while all-negative settings show falling friendship weights.","The building temperature converges to 21-22°C regardless of the personality mix, so agents with warm or hot preferences systematically lose happiness in the collective decision.","Individual-level factors are measurable: each extra degree an agent suggests adds roughly 0.25 happiness points, assertiveness adds about 3.24, and selflessness removes about 2.69.","LLM-driven generative agents can generate realistic daily occupant decisions for building energy modeling where human data are scarce or privacy-sensitive."],"supporting_citations":[{"why":"Introduces the generative-agent architecture of memory, reflection, planning, and dialogue that this simulation adapts for daily two-stage voting.","marker":"Park et al. (2023)"},{"why":"Provides the Big Five facet taxonomy and the adjective pairs used to assign each agent a personality.","marker":"Costa Jr and McCrae (1995)"},{"why":"Supplies the real-world social network whose 34 nodes and 78 edges become the building's households and friendships.","marker":"Zachary (1977)"},{"why":"Supplies the Mistral 7B LLM whose 8-bit quantized version generates every agent decision.","marker":"Jiang et al. (2024a)"},{"why":"Supplies the correlated-random-effects estimator used for the node-level panel regressions.","marker":"Mundlak (1978)"},{"why":"Supplies the simulation framework used to execute the daily methods and collect the network data.","marker":"Rende et al. (2025)"},{"why":"Supplies the Typical Meteorological Year weather data through pvlib that drives daily outside temperatures.","marker":"Jensen et al. (2023)"}],"fun_headline_variants":["LLM agents reveal personality drives happiness in shared housing","Positive agents outhappy negative ones in 30-day building sim","Simulated neighbors: personality matters more than thermostat","AI building sim links assertiveness and selflessness to joy","In LLM social sim, nice personalities yield 91 vs 86 happiness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a single 30-day run per personality setting gives a representative picture, with each agent-day treated as an independent observation; if the runs are not representative or the observations are correlated, the reported differences between positive and negative personalities could be noise.","fun_headline_variants_meta":{"raw":{"variants":["LLM agents reveal personality drives happiness in shared housing","Positive agents outhappy negative ones in 30-day building sim","Simulated neighbors: personality matters more than thermostat","AI building sim links assertiveness and selflessness to joy","In LLM social sim, nice personalities yield 91 vs 86 happiness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1305,"prompt_tokens":862,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":478,"tokens_out":443,"duration_ms":5342,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:50:50.220845+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat each of the three settings with at least ten random seeds and compare the distributions of daily average happiness and friendship weights; if the all-positive and all-negative confidence intervals overlap once seed-to-seed variation is included, the paper's central contrast is not supported.","supporting_citations":[{"cited_title":", author O'Brien, J","cited_arxiv_id":null,"evidence_quote":"Introduces the generative-agent architecture of memory, reflection, planning, and dialogue that this simulation adapts for daily two-stage voting."},{"cited_title":", author McCrae, R.R","cited_arxiv_id":null,"evidence_quote":"Provides the Big Five facet taxonomy and the adjective pairs used to assign each agent a personality."},{"cited_title":", year 1977","cited_arxiv_id":null,"evidence_quote":"Supplies the real-world social network whose 34 nodes and 78 edges become the building's households and friendships."},{"cited_title":", year 1978","cited_arxiv_id":null,"evidence_quote":"Supplies the correlated-random-effects estimator used for the node-level panel regressions."},{"cited_title":", author Yilmaz, T","cited_arxiv_id":null,"evidence_quote":"Supplies the simulation framework used to execute the daily methods and collect the network data."},{"cited_title":", author Anderson, K.S","cited_arxiv_id":null,"evidence_quote":"Supplies the Typical Meteorological Year weather data through pvlib that drives daily outside temperatures."}],"review_version":1}