REVIEW 2 major objections 5 minor 39 references
LLM community simulations cannot justify governance claims until they report necessary versus sufficient causation under explicit assumptions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 12:00 UTC
load-bearing objection Clean workshop position paper that imports PN/PS correctly and maps them to governance stakeholders; useful framing, no new empirics, monotonicity already flagged by the authors. the 2 major comments →
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For simulations to inform online-community governance, researchers must stop treating believability as enough and instead report probabilities of necessary and sufficient causation, mapped to distinct stakeholder questions and interpreted strictly as simulator-conditional estimates whose real-world value depends on fidelity.
What carries the argument
Probabilities of necessary causation (PN) and sufficient causation (PS). PN asks whether an observed harm would have been avoided without the cause; PS asks whether introducing the cause would have produced the outcome when it was absent. Under exogeneity and monotonicity they collapse to simple functions of the two outcome rates from factual and counterfactual simulation runs.
Load-bearing premise
The simple formulas work only if interventions usually push outcomes in one consistent direction; when social reactions reverse the effect, those formulas become mere lower bounds.
What would settle it
Instantiate a real, previously studied intervention (for example a quarantine or removal-explanation policy) inside a simulator, compute its p0, p1, PN and PS profiles, and check whether they preserve the rank-order and qualitative asymmetries observed in the corresponding real-platform study; systematic reversals would falsify policy transfer.
If this is right
- Moderators can treat high PN as evidence that a particular comment or action was needed for an observed escalation.
- Platform designers can treat high PS as evidence that an intervention is worth scaling.
- Simulator adequacy becomes concrete: p0 and p1 must approximate real-world rates for the interventions of interest.
- Reporting PN and PS across outcome thresholds yields a causal profile richer than any single average effect size.
- Detected monotonicity violations themselves become useful findings about heterogeneous or backfiring effects.
Where Pith is reading between the lines
- Standard PN/PS tooling layered on existing multi-agent platforms could turn causal profiling into a shared evaluation benchmark across simulators.
- The same necessity-versus-sufficiency split could clarify claims in other multi-agent LLM settings that currently report only behavioral plausibility.
- If monotonicity fails often, the practical output of policy simulations may need to shift from aggregate rates to distributions of thread-level effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This workshop position paper argues that LLM-based social simulations of online communities currently support only plausibility judgments, not the causal claims needed for governance decisions. It proposes adopting the causal counterfactual framework of Hannart et al. and Pearl, distinguishing the probability of necessary causation PN = P(Y0=0 | Y=1, X=1) from the probability of sufficient causation PS = P(Y1=1 | Y=0, X=0). Under exogeneity and monotonicity these reduce to simple functions of the outcome rates p0 and p1 (Equation 1), which paired simulation runs can estimate. The paper maps PN to moderators and liability questions and PS to platform designers and policy selection (Table 1), works a numerical counter-speech example, and insists that all such quantities be interpreted as simulator-conditional estimates whose policy relevance depends on fidelity of p0 and p1. A four-part research agenda covers tooling, real-world calibration, assumption testing, and stakeholder-centered presentation.
Significance. If the field adopts this framing, simulation-based governance work would gain a shared, explicit semantics for intervention claims instead of informal statements that an intervention 'seemed to help.' The paper correctly imports standard PN/PS definitions, states the identifying assumptions, notes that monotonicity violations turn Equation 1 into lower bounds, and treats estimates as simulator-conditional rather than automatically transferable. That combination is a useful conceptual contribution for a PoliSim@CHI workshop: it clarifies what 'adequate fidelity' would mean and gives stakeholders distinct, decision-relevant quantities. Strengths include the clean stakeholder mapping in Table 1, the worked numerical example, and an agenda that defers empirical validation rather than overclaiming it.
major comments (2)
- Section 4.1 and Equation 1: the operational payoff of the paper rests on PN and PS reducing to functions of p0 and p1 under monotonicity. The manuscript itself notes that social reflexivity (patronizing counter-speech, rule defiance) can reverse effects, in which case the expressions become only lower bounds. For a position paper this is correctly flagged, but the central claim that simulations can support 'precise, probabilistic, and policy-relevant claims' still needs a short, concrete protocol (even schematic) for when to report the simple estimators versus bounds or heterogeneous-effect diagnostics; otherwise the stakeholder mapping in Table 1 loses its clean operational form in the common reflexive case.
- Section 5: the adequacy criterion ('a simulator is adequate if its p0 and p1 approximate real-world values') is the load-bearing bridge from simulator-conditional estimates to policy relevance, yet it remains purely qualitative. Without even a provisional notion of approximation (e.g., rank-order preservation of interventions, or tolerance on severe-outcome rates), the claim that the framework 'helps define what adequate fidelity means' is asserted rather than specified. A brief, falsifiable calibration target would strengthen the argument without requiring new experiments in this manuscript.
minor comments (5)
- Section 1: the phrase 'for this gap for this gap' is a duplicated typo and should be corrected.
- Section 4.1: the sentence beginning 'However, PN and PS as we defined above...' repeats 'However' awkwardly; a single transition would improve readability.
- Section 4.2 worked example: the coding of X=1 as absence of counter-speech is correct but easy to misread; a one-line reminder that the 'cause' is the risk factor (absence) would reduce confusion for readers used to X=1 as treatment.
- References: several arXiv preprints and author self-citations are appropriate for related work, but a short note distinguishing prior empirical governance results from the present conceptual contribution would help readers place novelty.
- Table 1: 'Policy/Legal Team' row uses PN for liability; a brief caveat that legal standards of causation may not coincide with probabilistic PN would avoid overclaiming transferability.
Circularity Check
No circularity: position paper imports external PN/PS definitions and maps them conceptually without self-definitional reduction or load-bearing self-citation.
full rationale
The paper is a workshop position piece whose central claim is conceptual: adopt the necessary/sufficient causation framework of Hannart et al. (2016) and Pearl (2009) for LLM social simulations, map PN to moderators and PS to platform designers (Table 1), note that under exogeneity+monotonicity the quantities reduce to functions of estimable rates p0/p1 (Eq. 1), and interpret all results as simulator-conditional (Section 5). PN and PS are not defined in terms of the paper’s own conclusions; they are quoted from the external literature and then applied. There are no fitted parameters re-labeled as predictions, no uniqueness theorems imported from the authors’ prior work, no ansatz smuggled via self-citation, and no renaming of a known empirical pattern. Self-citations appear only in Related Works and the research agenda for the authors’ prior simulation and governance studies; none of them underwrite the derivation of the causal semantics themselves. Because the strongest claim is an invitation to adopt external semantics rather than an empirical result forced by construction, the derivation chain is self-contained and non-circular.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Exogeneity: intervention X is assigned independently of system dynamics that also affect Y (satisfied by researcher injection in simulation).
- domain assumption Monotonicity: X does not increase Y for some units and decrease it for others.
- standard math Standard counterfactual definitions PN = P(Y0=0|Y=1,X=1) and PS = P(Y1=1|Y=0,X=0) and their reduction under the above assumptions.
- ad hoc to paper A simulator is adequate for governance if its p0 and p1 approximate real-world values for the interventions and outcomes of interest.
invented entities (1)
-
simulator-conditional causal estimates
no independent evidence
read the original abstract
LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tested before deployment. But believability is not causality. Claims like ``intervention $A$ reduces escalation'' require causal semantics that current simulation work typically does not specify. We propose adopting the causal counterfactual framework, distinguishing \textit{necessary causation} (would the outcome have occurred without the intervention?) from \textit{sufficient causation} (does the intervention reliably produce the outcome?). This distinction maps onto different stakeholder needs: moderators diagnosing incidents require evidence about necessity, while platform designers choosing policies require evidence about sufficiency. We formalize this mapping, show how simulation design can support estimation under explicit assumptions, and argue that the resulting quantities should be interpreted as simulator-conditional causal estimates whose policy relevance depends on simulator fidelity. Establishing this framework now is essential: it helps define what adequate fidelity means and moves the field from simulations that look realistic toward simulations that can support policy changes.
Reference graph
Works this paper leans on
-
[1]
Jacy Reese Anthis, Ryan Liu, Sean M Richardson, Austin C Kozlowski, Bernard Koch, James Evans, Erik Brynjolfsson, and Michael Bernstein. 2025. Llm social simulations are a promising research method.arXiv preprint arXiv:2504.02234 (2025)
Pith/arXiv arXiv 2025
-
[2]
Alex Atcheson, Vinay Koshy, and Karrie Karahalios. 2024. Not What it Used to Be: Characterizing Content and User-base Changes in Newly Created Online Communities. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 738, 12 pages. doi:10....
doi:10.1145/3613904 2024
-
[3]
Jinyu Cai, Yusei Ishimizu, Mingyue Zhang, Munan Li, Jialong Li, and Kenji Tei
-
[4]
Simulation of language evolution under regulated social media platforms: A synergistic approach of large language models and genetic algorithms.arXiv preprint arXiv:2502.19193(2025)
Pith/arXiv arXiv 2025
-
[5]
Eshwar Chandrasekharan, Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2022. Quarantined! Examining the effects of a community-wide moderation interven- tion on Reddit.ACM Transactions on Computer-Human Interaction (TOCHI)29, 4 (2022), 1–26
2022
-
[6]
Eshwar Chandrasekharan, Umashanthi Pavalanathan, Anirudh Srinivasan, Adam Glynn, Jacob Eisenstein, and Eric Gilbert. 2017. You Can’t Stay Here: The Efficacy of Reddit’s 2015 Ban Examined Through Hate Speech.Proc. ACM Hum.-Comput. Interact.1, CSCW, Article 31 (Dec. 2017), 22 pages. doi:10.1145/3134666
doi:10.1145/3134666 2017
-
[7]
Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert. 2018. The Internet’s Hidden Rules: An Empirical Study of Reddit Norm Violations at Micro, Meso, and Macro Scales.Proc. ACM Hum.-Comput. Interact.2, CSCW, Article 32 (Nov. 2018), 25 pages. doi:10.1145/3274301
doi:10.1145/3274301 2018
-
[8]
Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy Rogers. 2024. Simu- lating opinion dynamics with networks of llm-based agents. InFindings of the association for computational linguistics: NAACL 2024. 3326–3346
2024
-
[9]
Yun-Shiuan Chuang, Krirk Nirunwiroj, Zach Studdiford, Agam Goyal, Vincent V Frigo, Sijia Yang, Dhavan V Shah, Junjie Hu, and Timothy T Rogers. 2024. Beyond demographics: Aligning role-playing LLM-based agents using human belief networks. InFindings of the Association for Computational Linguistics: EMNLP
2024
-
[10]
Yun-Shiuan Chuang, Siddharth Suresh, Nikunj Harlalka, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. 2023. The wisdom of partisan crowds: Comparing collective intelligence in humans and llm-based agents.arXiv preprint arXiv:2311.09665(2023)
Pith/arXiv arXiv 2023
-
[11]
Yi Feng, Chen Huang, Zhibo Man, Ryner Tan, Long P Hoang, Shaoyang Xu, and Wenxuan Zhang. 2026. MoltNet: Understanding Social Behavior of AI Agents in the Agent-Native MoltBook.arXiv preprint arXiv:2602.13458(2026). PoliSim@CHI 2026, April 16, 2026, Barcelona, Spain Goyal et al
Pith/arXiv arXiv 2026
-
[12]
Agam Goyal, Charlotte Lambert, and Eshwar Chandrasekharan. 2025. The language of approval: Identifying the drivers of positive feedback online.arXiv preprint arXiv:2509.10370(2025)
arXiv 2025
-
[13]
Agam Goyal, Charlotte Lambert, Yoshee Jain, and Eshwar Chandrasekharan. 2024. Uncovering the Internet’s Hidden Values: An Empirical Study of Desirable Be- havior Using Highly-Upvoted Content on Reddit.arXiv preprint arXiv:2410.13036 (2024)
Pith/arXiv arXiv 2024
-
[14]
Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, and Koustuv Saha. 2026. Social Simulacra in the Wild: AI Agent Communities on Moltbook. arXiv preprint arXiv:2603.16128(2026)
arXiv 2026
-
[15]
Agam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha, and Eshwar Chan- drasekharan. 2025. Momoe: Mixture of moderation experts framework for ai- assisted online governance. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 12656–12671
2025
-
[16]
Agam Goyal, Xianyang Zhan, Charlotte Lambert, Koustuv Saha, and Eshwar Chandrasekharan. 2026. VASTU: Value-Aligned Social Toolkit for Online Content Curation.arXiv preprint arXiv:2601.12491(2026)
arXiv 2026
-
[17]
Alexis Hannart, Judea Pearl, Friederike EL Otto, Philippe Naveau, and Michael Ghil. 2016. Causal counterfactual theory for the attribution of weather and climate-related events.Bulletin of the American Meteorological Society97, 1 (2016), 99–110
2016
-
[18]
Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2019. Does transparency in moderation really matter? User behavior after content removal explanations on reddit.Proceedings of the ACM on Human-Computer Interaction3, CSCW (2019), 1–27
2019
-
[19]
Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2019. Does Transparency in Moderation Really Matter?: User Behavior After Content Removal Explanations on Reddit.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 150 (Nov. 2019), 27 pages. doi:10.1145/3359252
doi:10.1145/3359252 2019
-
[20]
Vinay Koshy, Frederick Choi, Yi-Shyuan Chiang, Hari Sundaram, Eshwar Chan- drasekharan, and Karrie Karahalios. 2025. Venire: A Machine Learning-Guided Panel Review System for Community Content Moderation.Proceedings of the ACM on Human-Computer Interaction9, 7 (2025), 1–35
2025
-
[21]
2012.Building successful online communities: Evidence-based social design
Robert E Kraut and Paul Resnick. 2012.Building successful online communities: Evidence-based social design. Mit Press
2012
-
[22]
Deepak Kumar, Yousef Anees AbuHashem, and Zakir Durumeric. 2024. Watch your language: Investigating content moderation with large language models. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 18. 865–878
2024
-
[23]
Posi- tive reinforcement helps breed positive behavior
Charlotte Lambert, Frederick Choi, and Eshwar Chandrasekharan. 2024. " Posi- tive reinforcement helps breed positive behavior": Moderator Perspectives on Encouraging Desirable Behavior.Proceedings of the ACM on Human-Computer Interaction8, CSCW2 (2024), 1–33
2024
-
[24]
Charlotte Lambert, Agam Goyal, Eunice Mok, and Eshwar Chandrasekharan
-
[25]
Mind Your Ps and Qs: Supporting Positive Reinforcement in Moderation Through a Positive Queue.arXiv preprint arXiv:2509.18437(2025)
arXiv 2025
-
[26]
Charlotte Lambert, Koustuv Saha, and Eshwar Chandrasekharan. 2025. Does Positive Reinforcement Work?: A Quasi-Experimental Study of the Effects of Positive Feedback on Reddit. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–16
2025
-
[27]
Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, and Saadia Gabriel
Genglin Liu, Vivian T. Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, and Saadia Gabriel. 2025. MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Viol...
doi:10.18653/v1/2025 2025
-
[28]
J Nathan Matias. 2019. Preventing harassment and increasing group participation through social norms in 2,190 online science discussions.Proceedings of the National Academy of Sciences116, 20 (2019), 9785–9789
2019
-
[29]
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th annual acm symposium on user interface software and technology. 1–22
2023
-
[30]
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. InProceedings of the 35th annual ACM symposium on user interface software and technology. 1–18
2022
-
[31]
Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109 (2024)
Pith/arXiv arXiv 2024
-
[32]
2009.Causality
Judea Pearl. 2009.Causality. Cambridge university press
2009
-
[33]
Giuseppe Russo, Maciej Styczen, Manoel Horta Ribeiro, and Robert West. 2025. Does content moderation lead users away from fringe movements? evidence from a recovery community. InProceedings of the International AAAI Conference on Web and Social Media, Vol. 19. 1719–1734
2025
-
[34]
Morgan Klaus Scheuerman, Jialun Aaron Jiang, Casey Fiesler, and Jed R Brubaker
-
[35]
A framework of severity for harmful content online.Proceedings of the ACM on Human-Computer Interaction5, CSCW2 (2021), 1–33
2021
-
[36]
2017.The reflective practitioner: How professionals think in action
Donald A Schön. 2017.The reflective practitioner: How professionals think in action. Routledge
2017
-
[37]
Galen Cassebeer Weld, Leon Leibmann, Amy X Zhang, and Tim Althoff. 2025. Perceptions of Moderators as a Large-Scale Measure of Online Community Gov- ernance.Proceedings of the ACM on Human-Computer Interaction9, 7 (2025), 1–29
2025
-
[38]
Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, et al . 2024. Oasis: Open agent social interaction simulations with one million agents.arXiv preprint arXiv:2411.11581(2024)
Pith/arXiv arXiv 2024
-
[39]
Xianyang Zhan, Agam Goyal, Yilun Chen, Eshwar Chandrasekharan, and Kous- tuv Saha. 2025. SLM-mod: Small language models surpass LLMs at content moderation. InProceedings of the 2025 Conference of the Nations of the Ameri- cas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 8774–8790
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.