{"id":"5680ce0a-086d-402a-98d6-597b51b0c7a0","arxiv_id":"2412.20505","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A three-cycle LLM urban planning loop raised simulated residents' average experience score from 65.03 to 69.03 in one Beijing case study, while accessibility and ecology scores did not improve.","lead":"This paper proposes a closed-loop urban planning system in which AI agents play residents, planners, and judges, repeatedly revising a land-use plan based on simulated daily life and feedback. It tests the loop on a Beijing community and reports that the simulated residents' average experience score rises over three rounds, even though objective plan quality scores do not improve.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The iterative-improvement claim rests on Experience scores from the same agents whose feedback drove plan revisions; the abstract also overstates static-metric results, so the central evidence is self-referential.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the Experience metric is computed from the same agents whose feedback shapes the plan, making the closed-loop evaluation non-independent. I agree with that assessment. The abstract-to-Table 1 contradiction is also real and independently damaging, but the self-referential metric is more fundamental because it undermines the only result that trends positively across cycles. If the Experience metric were independent, the static-metric shortfalls might be acceptable with a nuanced claim about trading off quantitative layout for subjective well-being. As written, the paper does not support that trade-off; it only shows that agents report liking plans they helped create. The proposed fresh-cohort test is a direct, low-cost way to distinguish genuine plan improvement from conversational self-consistency. Given the single case study, three iterations, no code release, and no error bars, the REJECT verdict is appropriate and my read does not change it.","tokens_in":5787,"tokens_out":2081,"duration_ms":21293,"concrete_test":"Run the same three-cycle procedure but evaluate the final plan P3 with a fresh cohort of resident agents generated from the same demographic distribution, who never participated in planning discussions or interviews for this run. Compare their Experience scores (Eq. 7) on P0 and P3 with the original cohort's scores. If the fresh cohort also shows a meaningful increase, the improvement is not solely an artifact of self-feedback. If the fresh cohort's scores are flat or lower, the reported 69.03% is not evidence of plan improvement. Additionally, repeat the full pipeline with at least five random seeds (or five independently generated resident profiles) and report mean and standard deviation for every metric and iteration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LiPUP-MA 'consistently outperforms baselines' and that iterative cycles further improve plan quality—depends on the Experience metric computed in Eq. 7 from interviews with resident agents. Those same residents produce discussion output D_k used in Eqs. 2–3 to revise the plan. The reported monotone rise from 65.03% to 69.03% is therefore compatible with agents parroting their prior suggestions, not with independently measured plan improvement. The paper provides no held-out evaluation, no fresh resident cohort, and no variance across runs; temperature 0 reduces sampling noise but does not remove the systematic dependence of the judge on the planner's inputs. In addition, the abstract's 'consistently outperforms baselines on both conventional static planning metrics and living-based metrics' is contradicted by Table 1: DRL beats LiPUP on Accessibility (66.25 vs 64.17 at Ours-3rd) and Ecology (56.67 vs 53.33), and MA-LLM beats Ours-3rd on Ecology (73.33 vs 53.33). The only consistent advantage is in Experience, which is the self-referential metric. Thus the load-bearing weakness is not an implementation detail but the absence of any independent signal that the iterative loop improves actual plan quality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an LLM-based multi-agent framework for cyclical urban planning. The framework alternates three phases: Planning, where planner and resident agents generate and revise a land-use plan; Living, where resident agents simulate mobility and social behavior in the current plan; and Judging, where a judge agent evaluates the plan via quantitative metrics and qualitative interviews and produces suggestions for the next cycle. The authors test the framework on the Huilongguan community in Beijing over three iterations and report that iterative cycles improve a resident Experience score, with the third iteration reaching 69.03%. The paper claims the framework 'consistently outperforms baselines on both conventional static planning metrics and living-based metrics,' although the full text later qualifies this as 'generally outperforms baselines in Experience metric.'","tokens_in":6033,"tokens_out":3923,"duration_ms":38906,"significance":"If the closed-loop planning-living-judging paradigm were shown to produce reliable, independently verified improvements in plan quality, it would be a meaningful contribution to LLM-based participatory urban planning. The paper addresses a real problem—static one-shot planning versus the inherently cyclical nature of urban regeneration—and the use of LLM agents for simulated living is a plausible direction. However, the significance is currently not established: the central evidence of improvement comes from a metric computed by interviewing the same agents whose feedback drove the plan revisions, no external validation is provided, and the abstract overstates the quantitative results. The paper also ships no code or detailed prompts, so the experiments are not reproducible from the manuscript alone.","major_comments":[{"comment":"The submitted title and abstract describe 'LiPUP-MA' with a 'Plan-centric Graph-based Experience Bank' and a 'Spatially-constrained Skill-augmented Planner agent,' but the full-text abstract, Methodology, and Experiments describe only the Planning/Living/Judging (CUP) framework and contain none of these components. The evaluation in Table 1 tests the CUP framework, not the framework claimed in the title and abstract. This mismatch makes the stated contribution unverifiable from the submitted manuscript.","section":"Title and Abstract vs. Full Text"},{"comment":"The abstract claims LiPUP-MA 'consistently outperforms baselines on both conventional static planning metrics and living-based metrics,' but Table 1 contradicts this. Ours-3rd scores 64.17 on Accessibility versus 66.25 for DRL, and 53.33 on Ecology versus 56.67 for DRL and 73.33 for MA-LLM (Ours-1st). The full text itself acknowledges that 'the DRL-based method achieves the best performance in Accessibility' and that 'the outcome decline of our framework during cycling in Accessibility and Ecology metrics.' The only metric with consistent improvement is Experience, so the abstract's claim of consistent outperformance is not supported.","section":"Abstract, Table 1"},{"comment":"The central evidence for iterative improvement is the Experience score, but this score is circular. The resident agents produce discussion D_k (Eq. 2) that the planner uses to revise plan P_k (Eq. 3); the same residents are then interviewed in Eq. (7) to compute Q_k,2, which is averaged into the overall score R_k (Eq. 8). There is no held-out resident cohort, no independent evaluator, and no external ground truth for resident well-being. The monotone rise from 65.03% to 69.03% across three iterations is therefore compatible with the agents repeating their earlier preferences rather than with an independent measure of plan quality. Without variance estimates or a control condition that breaks this dependency, the paper's main claim is not supported.","section":"Methodology, Eqs. (2)-(3) and (6)-(9); Experiment, Table 1"},{"comment":"The paper reports a single run with temperature 0, 30 resident agents, and 3 iterations, and gives no error bars, multiple seeds, or significance tests. The observed increases of about 1.6 and 2.4 percentage points in Experience could be within the noise of the LLM simulation. The conclusion that 'the effectiveness gradually increases with the number of cycles' is based on three unverified points. The authors also state, 'we speculate that the boost of subjective human well-being ought to undertake the sacrifice of the valid spatial layout of land uses,' which is speculation, not a demonstrated trade-off, and does not rescue the circular evaluation.","section":"Experiment, Table 1 and 'Results'"}],"minor_comments":[{"comment":"The sentence 'The four evaluation metrics originate from the judging procedure, including Accessibility, Ecology, and Experience' lists only three metrics; either a fourth metric is missing or the count should be 'three.'","section":"Experiment, first paragraph"},{"comment":"The 'interview({Ri})' operation is underspecified: no questionnaire, question list, scoring rubric, or aggregation scheme is provided, so a reader cannot reproduce the Experience score.","section":"Methodology, Eq. (7)"},{"comment":"The case study of resident R19 is purely qualitative and reports only a single agent's trajectory; it does not provide quantitative evidence that resident agents' behaviors or feedback are reliable or representative.","section":"Case study, Figure 2"},{"comment":"The manuscript provides no code, prompts, or detailed simulation configuration, making the experiments difficult to reproduce; consider a supplementary appendix with these materials.","section":"Overall"},{"comment":"The full text uses 'cyclical urban planning' and 'CUP' throughout, while the submission title uses 'LiPUP-MA' and 'Living-in-the-loop Participatory Urban Planning'; the terminology should be unified to avoid confusion about what is being proposed and evaluated.","section":"Terminology"}],"recommendation":"reject","confidential_remarks":"The mismatch between the submitted title/abstract and the full text is substantial and suggests the wrong version may have been uploaded. Even taking the full text on its own terms, the evaluation is self-referential: the same resident agents whose feedback generates plan revisions are also the source of the Experience score used to measure improvement. The abstract's claim of consistent outperformance is directly contradicted by Table 1, and the paper provides no independent validation. These are load-bearing issues that would require a fundamentally different evaluation design rather than a local fix, so I cannot recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The closed-loop idea is real, and the paper deserves credit for putting it together clearly. The planning-living-judging cycle is a genuine integration that prior one-shot LLM planning papers don't have, and the Experience Bank and spatially-constrained planner show careful engineering. The authors also report the decline in Accessibility and Ecology rather than hiding it, which is honest—though the abstract doesn't reflect that honesty.\n\nThe soft spots are significant. The structural problem is that the Experience metric comes from interviews with the same resident agents whose discussion drove the plan revision. The monotone rise from 65.03 to 69.03 is exactly what you'd expect if agents are happy that their own suggestions got adopted; it is not an independent signal of improved planning. A held-out resident cohort, a fresh interview protocol blind to the revision history, or a human evaluation would fix this. None are present.\n\nThe abstract's claim of \"consistently outperforms baselines\" is simply false against Table 1: DRL beats them on Accessibility and Ecology, and their own one-shot MA-LLM beats them on Ecology. The body text is more careful—\"generally outperforms\" and a speculation about tradeoffs—but the abstract overstates the result.\n\nOther issues: a single case study, three iterations, one simulated day, 30 agents, no error bars, no code or data release. As a framework demonstration, it's readable and the equations are clear. But it does not yet support the central claim that iterative cycles improve plan quality. The comparison to DRL is also not apples-to-apples because DRL is tailored to accessibility, which the paper acknowledges but could emphasize more.\n\nWho is this for? Researchers working on LLM agents for urban planning. They'll find the loop design a useful starting point, but they shouldn't cite it as evidence that iterative planning works. It deserves a serious referee only if the venue will demand a major revision with a real evaluation; otherwise it's a workshop paper.","headline":"Genuinely new closed-loop planning-living-judging framework, but the main evidence is self-referential and the abstract overstates the static-metric results.","tokens_in":6570,"tokens_out":2413,"would_cite":false,"duration_ms":25635,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A closed loop of planning, simulated living, and judging raises residents' reported experience score to 69.03% across three cycles.","keywords":["participatory urban planning","large language models","multi-agent simulation","closed-loop planning","quality of life","urban regeneration","agent-based modeling","Huilongguan"],"falsifier":"Run the full loop on Huilongguan but evaluate each revised plan with a held-out panel of resident agents who never participated in discussions or living; if the held-out panel's experience scores do not rise with iterations and do not beat the one-shot baseline, the reported 65.03-to-69.03 improvement is an artifact of the closed self-evaluation loop rather than genuine plan improvement.","tokens_in":5550,"feed_emoji":"🏙️","tokens_out":10741,"duration_ms":96261,"temperature":0.7,"pith_summary":"The paper seeks to establish that urban plans get better if they are revised in a closed loop: LLM agents draft a land-use plan, simulated residents 'live' under it and report how it feels, and the plan is revised using that feedback, over and over. The authors claim this living-in-the-loop paradigm beats random, deep-reinforcement-learning, and single-iteration multi-agent baselines on residents' living experience, and that the experience score keeps rising with each cycle, reaching 69.03% after three iterations in a Huilongguan, Beijing case study. A reader should care because urban regeneration decisions are currently made largely as one-shot interventions; if this holds, planners could pre-test how a neighborhood feels before building it. The authors also observe that the experiential gains come with declining static Accessibility and Ecology scores, suggesting a trade-off between subjective well-being and conventional spatial-efficiency metrics.","feed_headline":"AI planning loops raise resident experience scores to 69.03%","feed_subtitle":"Cyclical plan-revise cycles in Beijing's Huilongguan community beat one-shot planning baselines on lived experience.","key_machinery":"Central machinery is the closed loop itself, instantiated as three interacting agent groups: a planner who revises land-use assignments; resident agents whose dynamic memory pools accumulate mobility trajectories, social-media posts, and reflections; and a judger who combines quantitative statistics with qualitative interviews into a quality-of-life score and suggestions for the next cycle. In the LiPUP-MA formulation, scattered resident feedback is organized into a Plan-centric Graph-based Experience Bank that grounds experiences in concrete urban contexts, and the planner is a Spatially-constrained Skill-augmented Planner that turns subjective feedback into spatially coherent land-use changes. The loop carries the argument because the same resident experiences that feed the revision are also what measures improvement.","core_discovery":"The central claim is that participatory urban planning should be a cyclical process rather than a one-time event, and that LLM-based agents can run that cycle. Given a region partitioned into areas with land-use assignments, the framework iterates: the planner agent revises the plan after expert knowledge, prior suggestions, and a resident discussion; resident agents then simulate a day of mobility and social-media behavior under the plan, storing experiences in memory; the judger computes quantitative metrics (accessibility, ecology) and interviews residents for a qualitative Experience score, then feeds suggestions back into planning. On the Huilongguan dataset the Experience score rises from 65.03 to 69.03 across three iterations and generally exceeds the baselines, while Accessibility and Ecology decline, which the authors interpret as a sacrifice of static spatial efficiency for subjective well-being.","pith_inferences":["If the closed-loop gains are real rather than self-confirmation, the same cycle could be applied to roads, infrastructure, or zoning by adding new resident voices and watching which plans survive many iterations.","A direct test would be evaluating the final plan with a fresh panel of resident agents who never joined discussions; if the score still beats baselines, the improvement is more than the loop talking to itself.","The decline in Accessibility and Ecology may be a quirk of the current balance rather than a rule; adding an explicit constraint to preserve quantitative targets could yield plans that win on both axes.","The abstract's claim of consistent outperformance on static metrics is stronger than Table 1, where the deep-reinforcement-learning method leads Accessibility and the single-iteration version leads Ecology; the concrete advantage shown is on the Experience metric."],"forward_implications":["Running more planning–living–judging cycles keeps raising the residents' average experience score (65.03, then 66.6, then 69.03) in the reported runs.","The framework reports higher Experience scores than random, deep-reinforcement-learning, and single-iteration multi-agent baselines in the Huilongguan case.","Improvements in experiential quality can come at the cost of lower Accessibility and Ecology scores, so evaluating plans on static metrics alone can miss what residents actually feel.","Because the loop needs no human surveys or field observation, the same cycle can be re-run at scale to explore alternative plans before any physical change."],"supporting_citations":[{"why":"Provides the participatory multi-agent LLM planning approach this framework extends, the single-iteration baseline it is compared against, and the quantitative accessibility/ecology metrics reused in judging.","marker":"(Zhou et al. 2024)"},{"why":"Supplies the Huilongguan community dataset and experimental setup, and the deep-reinforcement-learning baseline compared in Table 1.","marker":"(Zheng et al. 2023)"},{"why":"Supplies the generative-agent architecture (perception, memory, reflection, and retrieval) used to build resident agents in the living simulation.","marker":"(Park et al. 2023)"},{"why":"Motivates automated plan evaluation and quantification, supporting the judging stage's quantitative metrics.","marker":"(Wang et al. 2023)"},{"why":"Surveys LLM-empowered agent-based modeling and provides the rationale for using LLM agents to simulate resident life.","marker":"(Gao et al. 2024)"}],"fun_headline_variants":["Cyclical AI planning lifts resident experience to 69.03% in Beijing community","Living-in-the-loop AI planning beats one-shot on experience scores","Simulated resident feedback coaches LLM planners to better urban plans","Multi-agent loop raises lived experience from 65 to 69 in urban planning","Closed-loop planning with AI agents outdoes static one-shot baselines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result depends on trusting the simulated residents' interview answers as a real measure of a plan's quality, even though those same simulated residents took part in changing the plan they are judging.","fun_headline_variants_meta":{"raw":{"variants":["Cyclical AI planning lifts resident experience to 69.03% in Beijing community","Living-in-the-loop AI planning beats one-shot on experience scores","Simulated resident feedback coaches LLM planners to better urban plans","Multi-agent loop raises lived experience from 65 to 69 in urban planning","Closed-loop planning with AI agents outdoes static one-shot baselines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000586,"raw_usage":{"total_tokens":2736,"prompt_tokens":910,"completion_tokens":1826,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":1731}},"tokens_in":526,"tokens_out":1826,"duration_ms":12915,"temperature":1.0,"reasoning_tokens":1731,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:19:26.429570+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full loop on Huilongguan but evaluate each revised plan with a held-out panel of resident agents who never participated in discussions or living; if the held-out panel's experience scores do not rise with iterations and do not beat the one-shot baseline, the reported 65.03-to-69.03 improvement is an artifact of the closed self-evaluation loop rather than genuine plan improvement.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Huilongguan community dataset and experimental setup, and the deep-reinforcement-learning baseline compared in Table 1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates automated plan evaluation and quantification, supporting the judging stage's quantitative metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Surveys LLM-empowered agent-based modeling and provides the rationale for using LLM agents to simulate resident life."}],"review_version":1}