Pith. sign in

REVIEW 3 major objections 100 references

A genetic-algorithm method finds viable AI policy packages by balancing perceived harm reduction, expert cost, and lay priorities.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 18:55 UTC pith:BHYT7PC4

load-bearing objection Solid methodological synthesis that turns participatory SAPs + expert costs + LLM scenario deltas into a weight-sweepable GA fitness function; useful for early policy screening if you treat the outputs as deliberation starters, not validated packages. the 3 major comments →

arxiv 2605.27395 v2 pith:BHYT7PC4 submitted 2026-04-20 cs.CY cs.AI

Informing AI Policy Assessment using Large-Scale Simulation of Interventions

classification cs.CY cs.AI
keywords AI policygenetic algorithmparticipatory AIpolicy simulationharm mitigationstakeholder-action pairsscenario evaluationAI governance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Policymakers face billions of possible combinations of AI-governance actions and scarce early guidance on which to prioritize. This paper offers a simulation pipeline that rewrites short harm scenarios under different packages of stakeholder-action pairs, scores the perceived drop in severity and magnitude with an LLM aligned to lay ratings, subtracts expert-assessed implementation cost, and adds a participatory priority score. A genetic algorithm searches that combinatorial space and returns diverse local optima under different weightings of the three factors. Equal weights tend to surface only one or two popular low-cost actions; weighting only harm mitigation produces large, costly packages that often include expert-flagged nonviable options. The authors argue that the resulting diversity of packages can serve as a concrete starting point for deliberation and brings participatory input into practical policy pipelines.

Core claim

Varying the relative weights on perceived harm mitigation, expert cost, and lay-stakeholder priority systematically changes which stakeholder-action pairs a genetic algorithm selects as locally optimal policy packages for generative-AI media harms, and the diversity of those packages can orient further assessment.

What carries the argument

The scalar fitness F(S, P′) = α(M) − β(C) + γ(D), optimized by a genetic algorithm over binary subsets of stakeholder-action pairs, where M is the average LLM-scored drop in severity and magnitude after scenarios are rewritten under the package.

Load-bearing premise

That LLM scores of severity and magnitude on rewritten scenarios, aligned to average human ratings, are reliable enough proxies for perceived harm mitigation to drive the fitness ranking.

What would settle it

Have a fresh panel of lay raters score severity and magnitude on a held-out set of original versus policy-rewritten scenarios; if the human deltas reverse or scramble the ranking of packages that the LLM fitness function preferred, the optimization claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Policymakers can re-run the same search under different stated priorities and see which actions remain robust across weightings.
  • Omitting the expert cost weight frequently admits legally or technically nonviable actions, so expert veto is load-bearing.
  • Equal weighting typically recommends only one or two high-priority, low-to-moderate-cost actions rather than large regulatory suites.
  • The pipeline can be staged: filter nonviable options first, optimize for harm, then surface remaining packages for cost review and public vote.
  • Actions that appear across multiple harms (transparency labels, fact-checking, human element in journalism) repeatedly surface in both single- and multi-harm runs.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same combinatorial search could be applied to other multi-criteria policy domains (climate adaptation, public-health mandates) wherever participatory scores and expert cost estimates already exist.
  • Because the algorithm returns local optima, repeated runs effectively sample a portfolio of near-equally good packages that could feed later multi-objective tools.
  • If LLM deltas systematically mis-estimate mitigation for particular demographic groups, the method would quietly reweight which publics are protected without anyone noticing in the aggregate fitness score.
  • Pairing the weighted scalar search with a post-hoc Pareto front would let users inspect trade-offs after the fact rather than only under pre-committed weights.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper proposes a multi-criteria methodology for AI policy assessment that combines (i) lay-stakeholder participatory ratings of stakeholder-action pairs (SAPs), (ii) expert-assessed implementation costs (including non-viability flags), and (iii) LLM-simulated perceived harm mitigation obtained by rewriting narrative scenarios under candidate policy packages and scoring severity/magnitude deltas. A genetic algorithm explores the combinatorial space of SAP subsets for three generative-AI media harms (political manipulation, unemployment, media sensationalism), optimizing the scalarized fitness F(S,P') = α(M) − β(C) + γ(D) under varying component weights. Results show that equal weighting yields sparse low-cost high-participatory packages, pure harm-mitigation weighting yields large diverse packages (often including non-viable SAPs), and intermediate weights produce intermediate coverage; the authors argue the resulting diversity of local optima supplies useful starting points for deliberation and operationalizes participatory AI into policy pipelines.

Significance. If the proxies and optimization are reliable, the work supplies a concrete, reproducible pipeline that lets policymakers explicitly surface trade-offs among effectiveness, cost and public priority—something currently missing from most AI risk-management frameworks. Strengths include open source code, transparent weight sweeps and heat-maps (Figs. 2–5, Table 2), reported LLM–human correlations, robustness checks across models, and an explicit multi-impact extension. The method is positioned as an early-stage exploratory tool rather than a substitute for RCTs or final decisions, which is appropriately modest. The main contribution is therefore methodological: a practical way to integrate participatory data into large-scale policy search.

major comments (3)
  1. Sec. 3.2 and Eq. (2): The fitness term α(M) is driven exclusively by LLM-derived severity and magnitude deltas between original scenarios S and rewritten scenarios S'. The reported Pearson correlations (0.794 severity, 0.707 magnitude) are computed on absolute ratings of a 18-scenario depth test set, not on the deltas themselves. No human validation is provided that the ranking or magnitude of these deltas is preserved across packages of different size or content. The Claude comparison only shows that two LLMs produce statistically indistinguishable deltas (p>0.05); it does not establish fidelity to human-perceived mitigation. Because C and D are static lookup tables, any systematic bias in the LLM deltas directly distorts the fitness landscape and therefore the diversity of local optima claimed in Sec. 4–5 and Table 2. This validation gap is load-bearing for the central claim that the p
  2. Sec. 3.1.3: Cost values (and the hard non-viable override of 4) are assigned solely by the four-author interdisciplinary panel (average pairwise κ=0.481). While the paper correctly labels this a proof-of-concept, the method is presented as ready for “practical policy development pipelines.” Author-as-expert scoring introduces a potential circularity risk when the same team also designs the scenarios and interprets the GA outputs; independent expert panels or inter-lab reliability data would be required before the cost component can be treated as an external constraint rather than an internal modeling choice.
  3. Sec. 4.1–4.2 and Table 2: The claim that “the diversity of viable policy combinations found by the genetic algorithm could be a useful starting point” rests on the observation that different weight vectors produce different SAP sets. Because the GA is stochastic and only local optima are recovered, and because the only package-dependent term is the unvalidated LLM delta, it remains unclear whether the observed diversity reflects genuine multi-criteria trade-offs or artifacts of rewriting quality, package-size effects, or residual LLM self-preference. A minimal human re-rating of a stratified sample of (S,S') pairs under the final packages would be needed to confirm that the ranking of packages is stable.

Circularity Check

0 steps flagged

No load-bearing circularity: fitness is an explicit weighted sum of independently measured inputs; GA local optima are exploratory outputs, not forced by construction or self-citation of the result.

full rationale

The paper's central claim is methodological and empirical: a genetic algorithm optimizing the scalarized fitness F(S,P')=α(M)−β(C)+γ(D) (Eq. 5) under different explicit weights produces diverse local policy packages that can serve as deliberation starting points. M is obtained by LLM rewriting of fixed human-validated scenarios S under candidate packages P' followed by severity/magnitude scoring with an LLM aligned on a separate human-rated corpus (Sec. 3.2); C and D are static expert and participatory lookup tables (Tables 1, 3–5). Nothing in the derivation equates the output packages to the inputs by definition, nor is any parameter fitted to a target quantity and then re-presented as a prediction of that quantity. Self-citations to prior work [7,8] supply only the input scenario set S and the SAP inventory P; they do not justify the optimization result or uniqueness of any local optimum. The observed diversity under re-weighting (Table 2, Figs. 2–5) is therefore an empirical outcome of the search, not a tautology. Minor self-citation of input datasets is normal and non-load-bearing, yielding a score of 1 rather than 0.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The method rests on a handful of free hyperparameters chosen by pilot runs, standard evolutionary-algorithm assumptions, domain assumptions that SAPs are independent and that LLM scenario rewrites preserve harm semantics, and no new physical or mathematical entities.

free parameters (4)
  • severity/magnitude weights ws, wm = 0.65 / 0.35
    Fixed at 0.65/0.35 by author choice citing prior risk literature; directly scales the harm-mitigation term in Eq. 2.
  • component weights α, β, γ = multiple discrete triples
    Swept over discrete sets (all-equal, only-harm, mostly-harm, etc.); the entire result surface depends on these hand-chosen values.
  • GA population size, crossover rate, mutation rate, elitism k = Pop≈105–201; cx=0.8; mut=0.03; k=3
    Population sized by formula involving average nonzero SAPs; crossover 0.80, mutation 0.03, k=3 chosen from GA literature and pilots; control which local optima are found.
  • z-score normalization statistics = per-impact μ,σ from 1000 samples
    Mean and std of severity, magnitude, cost, D computed from 1000 random packages per impact type; rescale the three terms so they can be added.
axioms (4)
  • domain assumption LLM-rewritten scenarios under a policy package produce severity and magnitude deltas that are valid proxies for perceived harm mitigation.
    Core of the α(M) term (Sec. 3.1.2, Eq. 2); justified only by correlation with human ratings on a small test set.
  • domain assumption Expert panel cost scores (1–3 or non-viable) and participatory priority×agreement products are commensurable after z-scoring and can be linearly combined.
    Assumed in the scalarized fitness F (Eq. 5); no utility-theoretic justification given.
  • domain assumption Stakeholder-action pairs are largely independent; incompatibilities are rare and can be ignored except for explicit non-viable flags.
    Stated in Sec. 3.1.1; allows binary encoding of chromosomes.
  • standard math Standard genetic-algorithm operators (roulette selection, two-point crossover, bit-flip mutation, elitism) adequately explore the 2^n discrete space for local optima useful to policymakers.
    Invoked throughout Sec. 3.3; literature citations supplied but no proof of global optimality claimed.

pith-pipeline@v1.1.0-grok45 · 35858 in / 2838 out tokens · 34327 ms · 2026-07-12T18:55:54.232196+00:00 · methodology

0 comments
read the original abstract

As the rapid proliferation of AI systems and harms spurs efforts in AI governance around the world, prioritizing among competing policy options has become increasingly challenging for policymakers and researchers. We introduce a methodology for identifying viable policy options to mitigate specified AI harms, helping policymakers and researchers target areas that warrant greater time and resource investment. This method combines participatory evaluation of policies, expert assessment of implementation costs, and an LLM-based assessment of perceived harm mitigation under each policy option. We leverage a genetic algorithm-based simulation study to explore a vast solution space of potential policy combinations, and examine how outcomes change under different weightings of cost, participatory input, and harm mitigation. We find that this method enables exploration of different balances between participatory and expert components, allowing policymakers and researchers to assess how much weight to assign to each. We argue that the diversity of viable policy combinations found by the genetic algorithm could be a useful starting point for deliberation. This method operationalizes existing work on participatory AI by integrating it directly into practical policy development pipelines.

Figures

Figures reproduced from arXiv: 2605.27395 by Julia Barnett, Kimon Kieslich, Natali Helberger, Nicholas Diakopoulos.

Figure 1
Figure 1. Figure 1: Illustration of our methodology. Scenarios depicting harms ( [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: A heat map displaying the final suggested policy for each weight set for [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Final suggested policies identified by the genetic algorithm for different sets of weights for harms relating to political [PITH_FULL_IMAGE:figures/full_fig_p031_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Final suggested policies identified by the genetic algorithm for different sets of weights for harms relating to labor [PITH_FULL_IMAGE:figures/full_fig_p032_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Final suggested policies identified by the genetic algorithm for different sets of weights for harms relating to media [PITH_FULL_IMAGE:figures/full_fig_p033_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Final suggested policies identified by the genetic algorithm for different sets of weights for genetic algorithm runs [PITH_FULL_IMAGE:figures/full_fig_p034_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Visual representation of the “optimal” policy identified by the genetic algorithm under the weight conditions [PITH_FULL_IMAGE:figures/full_fig_p035_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Visual representation of the “optimal” policy identified by the genetic algorithm under the weight conditions [PITH_FULL_IMAGE:figures/full_fig_p036_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

100 extracted references · 14 linked inside Pith

  1. [1]

    Camilla Adelle and Sabine Weiland. 2012. Policy assessment: the state of the art.Impact Assessment and Project Appraisal30, 1 (2012), 25–33. doi:10.1080/14615517.2012.663256

  2. [2]

    Muhammad Amer, Tugrul U Daim, and Antonie Jetter. 2013. A review of scenario planning.Futures46 (2013), 23–40

  3. [3]

    Per Dannemand Andersen, Meiken Hansen, and Cynthia Selin. 2021. Stakeholder inclusion in scenario planning—A review of European projects.Technological Forecasting and Social Change169 (2021), 120802

  4. [4]

    Josh Andres, Chris Danta, Andrea Bianchi, Sungyeon Hong, Zhuying Li, Eduardo Benitez Sandoval, Charles Patrick Martin, and Ned Cooper. 2024. Understanding and Shaping Human-Technology Assemblages in the Age of Generative AI. InCompanion Publication of the 2024 ACM Designing Interactive Systems Conference. 413–416

  5. [5]

    Jacy Reese Anthis, Ryan Liu, Sean M Richardson, Austin C Kozlowski, Bernard Koch, James Evans, Erik Brynjolfsson, and Michael Bernstein. 2025. LLM social simulations are a promising research method.arXiv preprint arXiv:2504.02234(2025). Informing AI Policy Assessment FAccT ’26, June 25–28, 2026, Montreal, QC, Canada

  6. [6]

    Brhmie Balaram, Tony Greenham, and Jasmine Leonard. 2018. Artificial Intelligence: real public engagement.RSA, London. Retrieved November5 (2018), 2018

  7. [7]

    Julia Barnett, Kimon Kieslich, and Nicholas Diakopoulos. 2024. Simulating Policy Impacts: Developing a Generative Scenario Writing Method to Evaluate the Perceived Effects of Regulation. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 7. 82–93

  8. [8]

    Julia Barnett, Kimon Kieslich, Natali Helberger, and Nicholas Diakopoulos. 2025. Envisioning Stakeholder-Action Pairs to Mitigate Negative Impacts of AI: A Participatory Approach to Inform Policy Making. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 1424–1449

  9. [9]

    Julia Barnett, Kimon Kieslich, Jasmine Sinchai, and Nicholas Diakopoulos. 2025. Scenarios in Computing Research: A Systematic Review of the Use of Scenario Methods for Exploring the Future of Computing Technologies in Society.arXiv preprint arXiv:2506.05605(2025)

  10. [10]

    Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, Noémi Éltető, et al. 2024. Centaur: a foundation model of human cognition.arXiv preprint arXiv:2410.20268(2024)

  11. [11]

    James Bisbee, Joshua D Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M Larson. 2024. Synthetic replacements for human survey data? the perils of large language models.Political Analysis32, 4 (2024), 401–416

  12. [12]

    Andrea Bonaccorsi, Riccardo Apreda, and Gualtiero Fantoni. 2020. Expert biases in technology foresight. Why they are a problem and how to mitigate them.Technological Forecasting and Social Change151 (2020), 119855

  13. [13]

    Lena Börjeson, Mattias Höjer, Karl-Henrik Dreborg, Tomas Ekvall, and Göran Finnveden. 2006. Scenario types and techniques: Towards a user’s guide.Futures38, 7 (2006), 723–739

  14. [14]

    Zana Buçinca, Chau Minh Pham, Maurice Jakesch, Marco Tulio Ribeiro, Alexandra Olteanu, and Saleema Amershi. 2023. Aha!: Facilitating ai impact assessment by generating examples of harms.arXiv preprint arXiv:2306.03280(2023)

  15. [15]

    Michael Burnam-Fink. 2015. Creating narrative scenarios: Science fiction prototyping at Emerge.Futures70 (2015), 48–55

  16. [16]

    David L Carroll et al. 1996. Genetic algorithms and optimizing chemical oxygen-iodine lasers.Developments in theoretical and applied mechanics18, 3 (1996), 411–424

  17. [17]

    Carroll (Ed.)

    John M. Carroll (Ed.). 1995.Scenario–Based Design: Envisioning Work and Technology in System Development. John Wiley & Sons, New York

  18. [18]

    David Caswell and Shuwei Fang. 2024. AI in Journalism Futures.Initial Report(2024)

  19. [19]

    Tun-Jen Chang, Sang-Chin Yang, and Kuang-Jung Chang. 2009. Portfolio optimization problems in different risk measures using genetic algorithm.Expert Systems with applications36, 7 (2009), 10529–10537

  20. [20]

    Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023. CoMPosT: Characterizing and evaluating caricature in LLM simulations.arXiv preprint arXiv:2310.11501(2023)

  21. [21]

    Carlos A Coello Coello. 1999. A comprehensive survey of evolutionary-based multiobjective optimization techniques.Knowledge and Information systems1, 3 (1999), 269–308

  22. [22]

    European Commission. 2024. Proposal for a Regulation of the European Parliament and of the Council laying down harmonised rules on Artificial Intelligence (Artificial Intelligence Act) and amending certain Union legislative acts, Pub. L. No. COM(2021) 206 final

  23. [23]

    Pierre Le Coz, Jia An Liu, Debarun Bhattacharjya, Georgina Curto, and Serge Stinckwich. 2025. What Would an LLM Do? Evaluating Policymaking Capabilities of Large Language Models.arXiv preprint arXiv:2509.03827(2025)

  24. [24]

    Tatevik Davtyan. 2024. An Overview of Global Efforts Towards AI Regulation.Bulletin of Yerevan University C: Jurisprudence15, 2 (41) (2024), 158–174

  25. [25]

    Giovanni De Gregorio and Pietro Dunn. 2022. The European risk-based approaches: Connecting constitutional dots in the digital age. Common Market Law Review59, 2 (2022)

  26. [26]

    1975.An analysis of the behavior of a class of genetic adaptive systems.University of Michigan

    Kenneth Alan De Jong. 1975.An analysis of the behavior of a class of genetic adaptive systems.University of Michigan

  27. [27]

    Kenneth A De Jong and William M Spears. 1990. An analysis of the interacting roles of population size and crossover in genetic algorithms. InInternational Conference on Parallel Problem Solving from Nature. Springer, 38–47

  28. [28]

    Kalyanmoy Deb and Himanshu Jain. 2014. An Evolutionary Many-Objective Optimization Algorithm Using Reference-Point-Based Nondominated Sorting Approach, Part I: Solving Problems With Box Constraints.IEEE Transactions on Evolutionary Computation18, 4 (2014), 577–601. doi:10.1109/TEVC.2013.2281535

  29. [29]

    K. Deb, A. Pratap, S. Agarwal, and T. Meyarivan. 2002. A fast and elitist multiobjective genetic algorithm: NSGA-II.IEEE Transactions on Evolutionary Computation6, 2 (2002), 182–197. doi:10.1109/4235.996017

  30. [30]

    Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The participatory turn in ai design: Theoretical foundations and the current state of practice. InProceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. 1–23

  31. [31]

    Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2024. Questioning the survey responses of large language models.Advances in Neural Information Processing Systems37 (2024), 45850–45878

  32. [32]

    Yijiang River Dong, Tiancheng Hu, and Nigel Collier. 2024. Can llm be a personalized judge?arXiv preprint arXiv:2406.11657(2024). FAccT ’26, June 25–28, 2026, Montreal, QC, Canada Barnett et al

  33. [33]

    Larry J Eshelman, Richard A Caruana, and J David Schaffer. 1989. Biases in the crossover landscape. InProceedings of the third international conference on Genetic algorithms. 10–19

  34. [34]

    G Ezeani, A Koene, R Kumar, N Santiago, and D Wright. 2021. A survey of artificial intelligence risk assessment methodologies.The global state of play and leading practices identified(2021)

  35. [35]

    Jan Ferrer i Picó, Michelle Catta-Preta, Alex Trejo Omeñaca, Marc Vidal, and Josep Maria Monguet i Fierro. 2025. The time machine: future scenario generation through generative AI tools.Future Internet17, 1 (2025), 48

  36. [36]

    Carlos M Fonseca, Peter J Fleming, et al . 1993. Genetic algorithms for multiobjective optimization: formulationdiscussion and generalization.. InIcga, Vol. 93. 416–423

  37. [37]

    2020.The risk-based approach to data protection

    Raphaël Gellert. 2020.The risk-based approach to data protection. Oxford University Press

  38. [38]

    Tarleton Gillespie. 2024. Generative AI and the politics of visibility.Big Data & Society11, 2 (2024), 20539517241252131

  39. [39]

    David E Golberg. 1989. Genetic algorithms in search, optimization, and machine learning.Addion wesley1989, 102 (1989), 36

  40. [40]

    David E Goldberg, Kalyanmoy Deb, and James H Clark. 1991. Genetic algorithms, noise, and the sizing of populations.Complex systems 6 (1991), 333–362

  41. [41]

    2016.Multiple criteria decision analysis

    Salvatore Greco, Jose Figueira, and Matthias Ehrgott. 2016.Multiple criteria decision analysis. Vol. 37. Springer

  42. [42]

    John J Grefenstette. 1986. Optimization of control parameters for genetic algorithms.IEEE Transactions on systems, man, and cybernetics 16, 1 (1986), 122–128

  43. [43]

    David Hartmann, José Renato Laranjeira De Pereira, Chiara Streitbörger, and Bettina Berendt. 2025. Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society.AI and Ethics5, 4 (2025), 3617–3638

  44. [44]

    Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. 2024. Predicting results of social science experiments using large language models.Preprint(2024)

  45. [45]

    Guzyal Hill, Matthew Waddington, and Leon Qiu. 2025. From pen to algorithm: optimizing legislation for the future with artificial intelligence.AI & SOCIETY40, 4 (2025), 3075–3086

  46. [46]

    Michel Hohendanner, Chiara Ullstein, Dohjin Miyamoto, Emma F Huffman, Gudrun Socher, Jens Grossklags, and Hirotaka Osawa

  47. [47]

    ACM Hum.-Comput

    Metaverse Perspectives from Japan: A Participatory Speculative Design Case Study.Proc. ACM Hum.-Comput. Interact.8, CSCW2, Article 400 (Nov. 2024), 51 pages. doi:10.1145/3686939

  48. [48]

    1992.Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence

    John H Holland. 1992.Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence. MIT press

  49. [49]

    Hollywood

    John S. Hollywood. 2025.The Well-Tempered Artificial Intelligence Assistant for Policy Processes: Approximately Optimizing Artificial Intelligence Outputs to Empower Policy Discussions. RAND PEA3888-4. RAND Corporation. https://www.rand.org/content/dam/rand/ pubs/perspectives/PEA3800/PEA3888-4/RAND_PEA3888-4/RAND_PEA3888-4.pdf

  50. [50]

    2023.Large language models as simulated economic agents: What can we learn from homo silicus?Technical Report

    John J Horton. 2023.Large language models as simulated economic agents: What can we learn from homo silicus?Technical Report. National Bureau of Economic Research

  51. [51]

    Abe Bohan Hou, Hongru Du, Yichen Wang, Jingyu Zhang, Zixiao Wang, Paul Pu Liang, Daniel Khashabi, Lauren M Gardner, and Tianxing He. 2025. Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy. InSecond Conference on Language Modeling. https://openreview.net/forum?id=zSbecER9il

  52. [52]

    Tiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier, Dirk Hovy, and Paul Röttger. 2025. SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors.arXiv preprint arXiv:2510.17516(2025)

  53. [53]

    Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al. 2023. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714(2023)

  54. [54]

    Kimon Kieslich, Nicholas Diakopoulos, and Natali Helberger. 2025. Anticipating impacts: using large-scale scenario-writing to explore diverse implications of generative AI in the news environment.AI and Ethics5, 5 (2025), 4555–4577

  55. [55]

    Kimon Kieslich, Natali Helberger, and Nicholas Diakopoulos. 2026. Scenario-Based Sociotechnical Envisioning (SSE): An Approach to Enhance Systemic Risk Assessments.AI and Ethics(2026)

  56. [56]

    Run Wild a Little With Your Imagination

    Shamika Klassen and Casey Fiesler. 2022. " Run Wild a Little With Your Imagination" Ethical Speculation in Computing Education with Black Mirror. InProceedings of the 53rd ACM Technical Symposium on Computer Science Education-Volume 1. 836–842

  57. [57]

    Austin Kozlowski and James Evans. 2024. Simulating subjects: The promise and peril of ai stand-ins for social agents and interactions

  58. [58]

    Tzu-Sheng Kuo, Quan Ze Chen, Amy X Zhang, Jane Hsieh, Haiyi Zhu, and Kenneth Holstein. 2025. PolicyCraft: Supporting Collaborative and Participatory Policy Design through Case-Grounded Deliberation. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–24

  59. [59]

    James Edgar Lim and Julian Savulescu. 2025. AI preference prediction and policy making.AI & SOCIETY(2025), 1–15

  60. [60]

    Alex Jiahong Lu, Eleanor Villafranca Wikstrom, and Tawanna R Dillahunt. 2024. Contamination, Otherness, and Negotiating Bottom- Up Sociotechnical Imaginaries in Participatory Speculative Design. InProceedings of the Participatory Design Conference 2024: Full Papers-Volume 1. 173–186

  61. [61]

    Alfred D Martin Jr. 1955. Mathematical programming of portfolio selections.Management Science1, 2 (1955), 152–166. Informing AI Policy Assessment FAccT ’26, June 25–28, 2026, Montreal, QC, Canada

  62. [62]

    Donald Martin Jr, Vinodkumar Prabhakaran, Jill Kuhlberg, Andrew Smart, and William S Isaac. 2020. Participatory problem formulation for fairer machine learning through community based system dynamics.arXiv preprint arXiv:2005.07572(2020)

  63. [63]

    Anna-Katharina Meßmer and Martin Degeling. 2023. Auditing Recommender Systems–Putting the DSA into practice with a risk- scenario-based approach.arXiv preprint arXiv:2302.04556(2023)

  64. [64]

    1998.An introduction to genetic algorithms

    Melanie Mitchell. 1998.An introduction to genetic algorithms. MIT press

  65. [65]

    Jimin Mun, Liwei Jiang, Jenny Liang, Inyoung Cheong, Nicole DeCario, Yejin Choi, Tadayoshi Kohno, and Maarten Sap. 2024. Particip-ai: A democratic surveying framework for anticipating future ai use cases, harms and benefits. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 7. 997–1010

  66. [66]

    Priyanka Nanayakkara, Nicholas Diakopoulos, and Jessica Hullman. 2020. Anticipatory ethics and the role of uncertainty.arXiv preprint arXiv:2011.13170(2020)

  67. [67]

    Blagovesta Nikolova. 2014. The rise and promise of participatory foresight.European journal of futures research2, 1 (2014), 33

  68. [68]

    NIST. 2024. Artificial intelligence risk management framework: Generative artificial intelligence profile.NIST Trustworthy and Responsible AI Gaithersburg, MD, USA(2024)

  69. [69]

    Alexandra Olteanu, Solon Barocas, Su Lin Blodgett, Lisa Egede, Alicia DeVrio, and Myra Cheng. 2025. Ai automatons: Ai systems intended to imitate humans.arXiv preprint arXiv:2503.02250(2025)

  70. [70]

    Carsten Orwat, Jascha Bareis, Anja Folberth, Jutta Jahnel, and Christian Wadephul. 2024. Normative challenges of risk regulation of artificial intelligence.NanoEthics18, 2 (2024), 11

  71. [71]

    1998.Combinatorial optimization: algorithms and complexity

    Christos H Papadimitriou and Kenneth Steiglitz. 1998.Combinatorial optimization: algorithms and complexity. Courier Corporation

  72. [72]

    Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. InProceedings of the 36th annual acm symposium on user interface software and technology. 1–22

  73. [73]

    Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. 2024. Generative agent simulations of 1,000 people.arXiv preprint arXiv:2411.10109(2024)

  74. [74]

    Mark Petticrew, Zaid Chalabi, and David R Jones. 2012. To RCT or not to RCT: deciding when ‘more evidence is needed’for public health policy and practice.J Epidemiol Community Health66, 5 (2012), 391–396

  75. [75]

    Martijn Poel, Eric T Meyer, and Ralph Schroeder. 2018. Big data for policymaking: Great expectations, but with limited progress?Policy & Internet10, 3 (2018), 347–367

  76. [76]

    1991.Foundations of Genetic Algorithms 1991 (FOGA 1)

    Gregory JE Rawlins. 1991.Foundations of Genetic Algorithms 1991 (FOGA 1). Vol. 1. Elsevier

  77. [77]

    Anka Reuel, Avijit Ghosh, Jenny Chim, Andrew Tran, Yanan Long, Jennifer Mickel, Usman Gohar, Srishti Yadav, Pawan Sasanka Ammanamanchi, Mowafak Allaham, et al. 2025. Who Evaluates AI’s Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations.arXiv preprint arXiv:2511.05613(2025)

  78. [78]

    Friederike Richter, Kirsty Campbell, and Jasmin Riedl. 2025. Lost in translation: why digital twins thrive in research but falter in politics and public administration.Data & Policy7 (2025), e71. doi:10.1017/dap.2025.10027

  79. [79]

    Gernot Rieder and Judith Simon. 2016. Datatrust: Or, the political quest for numerical evidence and the epistemologies of Big Data.Big Data & Society3, 1 (2016), 2053951716649398

  80. [80]

    Huw Roberts, Josh Cowls, Emmie Hine, Jessica Morley, Vincent Wang, Mariarosaria Taddeo, and Luciano Floridi. 2023. Governing artificial intelligence in China and the European Union: Comparing aims and promoting ethical outcomes.The Information Society39, 2 (2023), 79–97

Showing first 80 references.