Pith. sign in

REVIEW 2 major objections 5 minor 39 references

LLM community simulations cannot justify governance claims until they report necessary versus sufficient causation under explicit assumptions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 12:00 UTC

load-bearing objection Clean workshop position paper that imports PN/PS correctly and maps them to governance stakeholders; useful framing, no new empirics, monotonicity already flagged by the authors. the 2 major comments →

arxiv 2604.03920 v2 submitted 2026-04-05 cs.CL

From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities

classification cs.CL
keywords LLM Social SimulationsPolicy InterventionsOnline Community GovernanceCounterfactual CausationNecessary CausationSufficient CausationSimulator FidelityContent Moderation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

LLM-based social simulations can generate believable community interactions and serve as policy wind tunnels for testing moderation and design changes before deployment. Believability alone, however, cannot answer causal questions such as whether an intervention reduces escalation. This paper argues that the field should adopt the causal counterfactual framework, which distinguishes necessary causation—would the outcome still have occurred without the intervention?—from sufficient causation—does the intervention reliably produce the desired outcome? Those two quantities map onto different stakeholder needs: moderators diagnosing specific incidents need necessity evidence, while platform designers choosing policies need sufficiency evidence. Under the assumptions of exogeneity and monotonicity, both can be estimated from paired simulation runs that toggle the intervention, but the resulting numbers are only as policy-relevant as the simulator’s fidelity to real community dynamics.

Core claim

For simulations to inform online-community governance, researchers must stop treating believability as enough and instead report probabilities of necessary and sufficient causation, mapped to distinct stakeholder questions and interpreted strictly as simulator-conditional estimates whose real-world value depends on fidelity.

What carries the argument

Probabilities of necessary causation (PN) and sufficient causation (PS). PN asks whether an observed harm would have been avoided without the cause; PS asks whether introducing the cause would have produced the outcome when it was absent. Under exogeneity and monotonicity they collapse to simple functions of the two outcome rates from factual and counterfactual simulation runs.

Load-bearing premise

The simple formulas work only if interventions usually push outcomes in one consistent direction; when social reactions reverse the effect, those formulas become mere lower bounds.

What would settle it

Instantiate a real, previously studied intervention (for example a quarantine or removal-explanation policy) inside a simulator, compute its p0, p1, PN and PS profiles, and check whether they preserve the rank-order and qualitative asymmetries observed in the corresponding real-platform study; systematic reversals would falsify policy transfer.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Moderators can treat high PN as evidence that a particular comment or action was needed for an observed escalation.
  • Platform designers can treat high PS as evidence that an intervention is worth scaling.
  • Simulator adequacy becomes concrete: p0 and p1 must approximate real-world rates for the interventions of interest.
  • Reporting PN and PS across outcome thresholds yields a causal profile richer than any single average effect size.
  • Detected monotonicity violations themselves become useful findings about heterogeneous or backfiring effects.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Standard PN/PS tooling layered on existing multi-agent platforms could turn causal profiling into a shared evaluation benchmark across simulators.
  • The same necessity-versus-sufficiency split could clarify claims in other multi-agent LLM settings that currently report only behavioral plausibility.
  • If monotonicity fails often, the practical output of policy simulations may need to shift from aggregate rates to distributions of thread-level effects.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This workshop position paper argues that LLM-based social simulations of online communities currently support only plausibility judgments, not the causal claims needed for governance decisions. It proposes adopting the causal counterfactual framework of Hannart et al. and Pearl, distinguishing the probability of necessary causation PN = P(Y0=0 | Y=1, X=1) from the probability of sufficient causation PS = P(Y1=1 | Y=0, X=0). Under exogeneity and monotonicity these reduce to simple functions of the outcome rates p0 and p1 (Equation 1), which paired simulation runs can estimate. The paper maps PN to moderators and liability questions and PS to platform designers and policy selection (Table 1), works a numerical counter-speech example, and insists that all such quantities be interpreted as simulator-conditional estimates whose policy relevance depends on fidelity of p0 and p1. A four-part research agenda covers tooling, real-world calibration, assumption testing, and stakeholder-centered presentation.

Significance. If the field adopts this framing, simulation-based governance work would gain a shared, explicit semantics for intervention claims instead of informal statements that an intervention 'seemed to help.' The paper correctly imports standard PN/PS definitions, states the identifying assumptions, notes that monotonicity violations turn Equation 1 into lower bounds, and treats estimates as simulator-conditional rather than automatically transferable. That combination is a useful conceptual contribution for a PoliSim@CHI workshop: it clarifies what 'adequate fidelity' would mean and gives stakeholders distinct, decision-relevant quantities. Strengths include the clean stakeholder mapping in Table 1, the worked numerical example, and an agenda that defers empirical validation rather than overclaiming it.

major comments (2)
  1. Section 4.1 and Equation 1: the operational payoff of the paper rests on PN and PS reducing to functions of p0 and p1 under monotonicity. The manuscript itself notes that social reflexivity (patronizing counter-speech, rule defiance) can reverse effects, in which case the expressions become only lower bounds. For a position paper this is correctly flagged, but the central claim that simulations can support 'precise, probabilistic, and policy-relevant claims' still needs a short, concrete protocol (even schematic) for when to report the simple estimators versus bounds or heterogeneous-effect diagnostics; otherwise the stakeholder mapping in Table 1 loses its clean operational form in the common reflexive case.
  2. Section 5: the adequacy criterion ('a simulator is adequate if its p0 and p1 approximate real-world values') is the load-bearing bridge from simulator-conditional estimates to policy relevance, yet it remains purely qualitative. Without even a provisional notion of approximation (e.g., rank-order preservation of interventions, or tolerance on severe-outcome rates), the claim that the framework 'helps define what adequate fidelity means' is asserted rather than specified. A brief, falsifiable calibration target would strengthen the argument without requiring new experiments in this manuscript.
minor comments (5)
  1. Section 1: the phrase 'for this gap for this gap' is a duplicated typo and should be corrected.
  2. Section 4.1: the sentence beginning 'However, PN and PS as we defined above...' repeats 'However' awkwardly; a single transition would improve readability.
  3. Section 4.2 worked example: the coding of X=1 as absence of counter-speech is correct but easy to misread; a one-line reminder that the 'cause' is the risk factor (absence) would reduce confusion for readers used to X=1 as treatment.
  4. References: several arXiv preprints and author self-citations are appropriate for related work, but a short note distinguishing prior empirical governance results from the present conceptual contribution would help readers place novelty.
  5. Table 1: 'Policy/Legal Team' row uses PN for liability; a brief caveat that legal standards of causation may not coincide with probabilistic PN would avoid overclaiming transferability.

Circularity Check

0 steps flagged

No circularity: position paper imports external PN/PS definitions and maps them conceptually without self-definitional reduction or load-bearing self-citation.

full rationale

The paper is a workshop position piece whose central claim is conceptual: adopt the necessary/sufficient causation framework of Hannart et al. (2016) and Pearl (2009) for LLM social simulations, map PN to moderators and PS to platform designers (Table 1), note that under exogeneity+monotonicity the quantities reduce to functions of estimable rates p0/p1 (Eq. 1), and interpret all results as simulator-conditional (Section 5). PN and PS are not defined in terms of the paper’s own conclusions; they are quoted from the external literature and then applied. There are no fitted parameters re-labeled as predictions, no uniqueness theorems imported from the authors’ prior work, no ansatz smuggled via self-citation, and no renaming of a known empirical pattern. Self-citations appear only in Related Works and the research agenda for the authors’ prior simulation and governance studies; none of them underwrite the derivation of the causal semantics themselves. Because the strongest claim is an invitation to adopt external semantics rather than an empirical result forced by construction, the derivation chain is self-contained and non-circular.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

The central claim rests on importing standard counterfactual identities and two classical attribution assumptions, then adding the domain claim that simulation injection satisfies exogeneity and that fidelity is defined by matching real p0/p1. No free parameters are fitted; outcome thresholds are analyst choices. The main invented framing is the simulator-conditional status of the estimates.

axioms (4)
  • domain assumption Exogeneity: intervention X is assigned independently of system dynamics that also affect Y (satisfied by researcher injection in simulation).
    Invoked in Section 4.1 so that PN/PS reduce to functions of p0 and p1; standard for experiments, asserted by construction for sims.
  • domain assumption Monotonicity: X does not increase Y for some units and decrease it for others.
    Section 4.1; required for Equation 1 equality rather than lower bounds; paper notes social reflexivity can violate it.
  • standard math Standard counterfactual definitions PN = P(Y0=0|Y=1,X=1) and PS = P(Y1=1|Y=0,X=0) and their reduction under the above assumptions.
    Taken from Hannart et al. and Pearl (Sections 3–4); not re-derived.
  • ad hoc to paper A simulator is adequate for governance if its p0 and p1 approximate real-world values for the interventions and outcomes of interest.
    Section 5 defines fidelity operationally via these rates; this is the paper’s proposed adequacy criterion, not an external theorem.
invented entities (1)
  • simulator-conditional causal estimates no independent evidence
    purpose: Qualify PN/PS computed inside LLM social simulations so policy relevance is explicitly tied to fidelity rather than believability.
    Framing introduced in Abstract and Section 5; no independent measurement protocol beyond the proposed research agenda.

pith-pipeline@v1.1.0-grok45 · 14926 in / 2765 out tokens · 33424 ms · 2026-07-13T12:00:06.789238+00:00 · methodology

0 comments
read the original abstract

LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tested before deployment. But believability is not causality. Claims like ``intervention $A$ reduces escalation'' require causal semantics that current simulation work typically does not specify. We propose adopting the causal counterfactual framework, distinguishing \textit{necessary causation} (would the outcome have occurred without the intervention?) from \textit{sufficient causation} (does the intervention reliably produce the outcome?). This distinction maps onto different stakeholder needs: moderators diagnosing incidents require evidence about necessity, while platform designers choosing policies require evidence about sufficiency. We formalize this mapping, show how simulation design can support estimation under explicit assumptions, and argue that the resulting quantities should be interpreted as simulator-conditional causal estimates whose policy relevance depends on simulator fidelity. Establishing this framework now is essential: it helps define what adequate fidelity means and moves the field from simulations that look realistic toward simulations that can support policy changes.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 7 linked inside Pith

  1. [1]

    Jacy Reese Anthis, Ryan Liu, Sean M Richardson, Austin C Kozlowski, Bernard Koch, James Evans, Erik Brynjolfsson, and Michael Bernstein. 2025. Llm social simulations are a promising research method.arXiv preprint arXiv:2504.02234 (2025)

  2. [2]

    Alex Atcheson, Vinay Koshy, and Karrie Karahalios. 2024. Not What it Used to Be: Characterizing Content and User-base Changes in Newly Created Online Communities. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 738, 12 pages. doi:10....

  3. [3]

    Jinyu Cai, Yusei Ishimizu, Mingyue Zhang, Munan Li, Jialong Li, and Kenji Tei

  4. [4]

    Simulation of language evolution under regulated social media platforms: A synergistic approach of large language models and genetic algorithms.arXiv preprint arXiv:2502.19193(2025)

  5. [5]

    Eshwar Chandrasekharan, Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2022. Quarantined! Examining the effects of a community-wide moderation interven- tion on Reddit.ACM Transactions on Computer-Human Interaction (TOCHI)29, 4 (2022), 1–26

  6. [6]

    Eshwar Chandrasekharan, Umashanthi Pavalanathan, Anirudh Srinivasan, Adam Glynn, Jacob Eisenstein, and Eric Gilbert. 2017. You Can’t Stay Here: The Efficacy of Reddit’s 2015 Ban Examined Through Hate Speech.Proc. ACM Hum.-Comput. Interact.1, CSCW, Article 31 (Dec. 2017), 22 pages. doi:10.1145/3134666

  7. [7]

    Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert. 2018. The Internet’s Hidden Rules: An Empirical Study of Reddit Norm Violations at Micro, Meso, and Macro Scales.Proc. ACM Hum.-Comput. Interact.2, CSCW, Article 32 (Nov. 2018), 25 pages. doi:10.1145/3274301

  8. [8]

    Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy Rogers. 2024. Simu- lating opinion dynamics with networks of llm-based agents. InFindings of the association for computational linguistics: NAACL 2024. 3326–3346

  9. [9]

    Yun-Shiuan Chuang, Krirk Nirunwiroj, Zach Studdiford, Agam Goyal, Vincent V Frigo, Sijia Yang, Dhavan V Shah, Junjie Hu, and Timothy T Rogers. 2024. Beyond demographics: Aligning role-playing LLM-based agents using human belief networks. InFindings of the Association for Computational Linguistics: EMNLP

  10. [10]

    Yun-Shiuan Chuang, Siddharth Suresh, Nikunj Harlalka, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. 2023. The wisdom of partisan crowds: Comparing collective intelligence in humans and llm-based agents.arXiv preprint arXiv:2311.09665(2023)

  11. [11]

    Yi Feng, Chen Huang, Zhibo Man, Ryner Tan, Long P Hoang, Shaoyang Xu, and Wenxuan Zhang. 2026. MoltNet: Understanding Social Behavior of AI Agents in the Agent-Native MoltBook.arXiv preprint arXiv:2602.13458(2026). PoliSim@CHI 2026, April 16, 2026, Barcelona, Spain Goyal et al

  12. [12]

    Agam Goyal, Charlotte Lambert, and Eshwar Chandrasekharan. 2025. The language of approval: Identifying the drivers of positive feedback online.arXiv preprint arXiv:2509.10370(2025)

  13. [13]

    Agam Goyal, Charlotte Lambert, Yoshee Jain, and Eshwar Chandrasekharan. 2024. Uncovering the Internet’s Hidden Values: An Empirical Study of Desirable Be- havior Using Highly-Upvoted Content on Reddit.arXiv preprint arXiv:2410.13036 (2024)

  14. [14]

    Agam Goyal, Olivia Pal, Hari Sundaram, Eshwar Chandrasekharan, and Koustuv Saha. 2026. Social Simulacra in the Wild: AI Agent Communities on Moltbook. arXiv preprint arXiv:2603.16128(2026)

  15. [15]

    Agam Goyal, Xianyang Zhan, Yilun Chen, Koustuv Saha, and Eshwar Chan- drasekharan. 2025. Momoe: Mixture of moderation experts framework for ai- assisted online governance. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 12656–12671

  16. [16]

    Agam Goyal, Xianyang Zhan, Charlotte Lambert, Koustuv Saha, and Eshwar Chandrasekharan. 2026. VASTU: Value-Aligned Social Toolkit for Online Content Curation.arXiv preprint arXiv:2601.12491(2026)

  17. [17]

    Alexis Hannart, Judea Pearl, Friederike EL Otto, Philippe Naveau, and Michael Ghil. 2016. Causal counterfactual theory for the attribution of weather and climate-related events.Bulletin of the American Meteorological Society97, 1 (2016), 99–110

  18. [18]

    Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2019. Does transparency in moderation really matter? User behavior after content removal explanations on reddit.Proceedings of the ACM on Human-Computer Interaction3, CSCW (2019), 1–27

  19. [19]

    Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2019. Does Transparency in Moderation Really Matter?: User Behavior After Content Removal Explanations on Reddit.Proc. ACM Hum.-Comput. Interact.3, CSCW, Article 150 (Nov. 2019), 27 pages. doi:10.1145/3359252

  20. [20]

    Vinay Koshy, Frederick Choi, Yi-Shyuan Chiang, Hari Sundaram, Eshwar Chan- drasekharan, and Karrie Karahalios. 2025. Venire: A Machine Learning-Guided Panel Review System for Community Content Moderation.Proceedings of the ACM on Human-Computer Interaction9, 7 (2025), 1–35

  21. [21]

    2012.Building successful online communities: Evidence-based social design

    Robert E Kraut and Paul Resnick. 2012.Building successful online communities: Evidence-based social design. Mit Press

  22. [22]

    Deepak Kumar, Yousef Anees AbuHashem, and Zakir Durumeric. 2024. Watch your language: Investigating content moderation with large language models. In Proceedings of the International AAAI Conference on Web and Social Media, Vol. 18. 865–878

  23. [23]

    Posi- tive reinforcement helps breed positive behavior

    Charlotte Lambert, Frederick Choi, and Eshwar Chandrasekharan. 2024. " Posi- tive reinforcement helps breed positive behavior": Moderator Perspectives on Encouraging Desirable Behavior.Proceedings of the ACM on Human-Computer Interaction8, CSCW2 (2024), 1–33

  24. [24]

    Charlotte Lambert, Agam Goyal, Eunice Mok, and Eshwar Chandrasekharan

  25. [25]

    Mind Your Ps and Qs: Supporting Positive Reinforcement in Moderation Through a Positive Queue.arXiv preprint arXiv:2509.18437(2025)

  26. [26]

    Charlotte Lambert, Koustuv Saha, and Eshwar Chandrasekharan. 2025. Does Positive Reinforcement Work?: A Quasi-Experimental Study of the Effects of Positive Feedback on Reddit. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–16

  27. [27]

    Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, and Saadia Gabriel

    Genglin Liu, Vivian T. Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, and Saadia Gabriel. 2025. MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Viol...

  28. [28]

    J Nathan Matias. 2019. Preventing harassment and increasing group participation through social norms in 2,190 online science discussions.Proceedings of the National Academy of Sciences116, 20 (2019), 9785–9789

  29. [29]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th annual acm symposium on user interface software and technology. 1–22

  30. [30]

    Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022. Social simulacra: Creating populated prototypes for social computing systems. InProceedings of the 35th annual ACM symposium on user interface software and technology. 1–18

  31. [31]

    Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109 (2024)

  32. [32]

    2009.Causality

    Judea Pearl. 2009.Causality. Cambridge university press

  33. [33]

    Giuseppe Russo, Maciej Styczen, Manoel Horta Ribeiro, and Robert West. 2025. Does content moderation lead users away from fringe movements? evidence from a recovery community. InProceedings of the International AAAI Conference on Web and Social Media, Vol. 19. 1719–1734

  34. [34]

    Morgan Klaus Scheuerman, Jialun Aaron Jiang, Casey Fiesler, and Jed R Brubaker

  35. [35]

    A framework of severity for harmful content online.Proceedings of the ACM on Human-Computer Interaction5, CSCW2 (2021), 1–33

  36. [36]

    2017.The reflective practitioner: How professionals think in action

    Donald A Schön. 2017.The reflective practitioner: How professionals think in action. Routledge

  37. [37]

    Galen Cassebeer Weld, Leon Leibmann, Amy X Zhang, and Tim Althoff. 2025. Perceptions of Moderators as a Large-Scale Measure of Online Community Gov- ernance.Proceedings of the ACM on Human-Computer Interaction9, 7 (2025), 1–29

  38. [38]

    Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, et al . 2024. Oasis: Open agent social interaction simulations with one million agents.arXiv preprint arXiv:2411.11581(2024)

  39. [39]

    Xianyang Zhan, Agam Goyal, Yilun Chen, Eshwar Chandrasekharan, and Kous- tuv Saha. 2025. SLM-mod: Small language models surpass LLMs at content moderation. InProceedings of the 2025 Conference of the Nations of the Ameri- cas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 8774–8790