Pith. sign in

REVIEW 4 major objections 6 minor 93 references

Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper proposes a two-step mixed-methods framework that first measures what clinicians, patients, policymakers, and other stakeholders need from explanations of health simulations, then steers LLMs to write summaries matching each group

desk verdict A clear, honest vision paper for stakeholder-tailored LLM summaries of health simulations; the construct-validity gap between 'information needs' and the empathy measure is the main thing to fix before this becomes a validated pipeline. read the letter →

arxiv 2509.04646 v1 pith:IBEYBLQW submitted 2025-09-04 cs.AI cs.ET

classification cs.AIcs.ET
keywords personalizedsummarizationlargelanguagemodelshealthsimulationagent-basedstakeholderpreferencesmixed-methodsframeworkempathymeasurementmodel-to-textgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Health simulations are complex, and the people who could use them—clinicians, policymakers, patients, caregivers, and advocates—need different information in different styles. This paper argues that current LLM-generated explanations of models and simulations are one-size-fits-all and therefore miss the mark. It proposes a step-by-step mixed-methods framework: generate technically correct candidate summaries using a 2x2 factorial design over content and style, measure stakeholder reactions with validated empathy questionnaires, analyze the results per stakeholder group, and steer an LLM through preference optimization to produce tailored summaries. A factuality check by modelers sits between generation and elicitation. If the framework works, it gives a repeatable process for turning any health model and its simulation outputs into summaries that are correct, preference-aligned, and actionable for each audience.

What carries the argument

The mechanism is a closed loop: model decomposition into structured triplets small enough for an LLM; factorial design generating candidate summaries that vary controllable attributes such as length, tone, and topic coverage; a modeler factuality check covering knowledge, reasoning, and relevance errors; validated empathy instruments (the Toronto Empathy Questionnaire and the State Empathy Scale) capturing reader reactions; factorial analysis with effect sizes, power analysis, and repeated-measures ANOVA identifying group preferences; preference optimization (DPO and similar alignment methods) steering the LLM; and automatic LLM-based evaluation metrics that filter candidates before improved

What would settle it

Run the full loop with two stakeholder groups on the same health simulation. If the factorial analysis finds no significant between-group differences in empathy scores across the four content-by-style summaries, or if a retest one month later reverses each group's preferred combination, the central premise of stable group-level tailoring fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the barrier to using health simulations is not only technical but communicative: we lack a systematic account of what different stakeholders need from an explanation. The authors' main contribution is a vision and framework for eliciting, incorporating, and evaluating those needs. The framework has two broad steps. First, the model structure is decomposed into small structured pieces (such as RDF triplets) and translated into text, while simulation outputs are summarized via statistical analysis or multimodal LLMs; candidate summaries are then generated by varying content and style in a designed experiment, checked for factuality by modelers, and piloted wit

Load-bearing premise

The framework assumes that stakeholder groups have stable, distinct preferences for summary content and style that can be elicited reliably with a 2x2 factorial design plus empathy questionnaires, and then used to steer an LLM—an assumption the seven-person pilot does not test.

Editorial extensions

If this is right

  • Different stakeholder groups can receive summaries with the same factual core but different content coverage and style, replacing the single generic text now produced.
  • Factuality becomes a gate: no candidate summary reaches stakeholders until modelers have checked it for knowledge, reasoning, and relevance errors.
  • Because preferences are analyzed per group and the LLM is steered accordingly, the approach extends existing model-to-text pipelines without retraining from scratch; few-shot examples can suffice for domain adaptation.
  • Evaluation shifts from text quality alone to whether tailored summaries change downstream decisions, ideally tested by a randomized controlled trial comparing generic versus tailored summaries.
  • The same pipeline can be re-run for new models or new stakeholder groups, making the process repeatable rather than a one-off customization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If preferences are not stable across time or contexts, group-level tailoring may need to become individual-level or dynamically updated; the seven-person pilot does not yet rule this out.
  • Empathy is the only reaction dimension the protocol measures; trust, perceived accuracy, cognitive load, and actionability may also drive preferences and would not be captured directly.
  • A stronger test than preference alignment is behavioral: tailored summaries should change the decisions stakeholders actually make. A vignette study comparing choices after generic versus tailored summaries would be a cheap precursor to the named randomized controlled trial.
  • The reliance on empathy instruments assumes textual summaries are evaluated emotionally; administrative or technical audiences may prefer a low-empathy executive style, which would complicate the factorial interpretation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper argues that existing LLM-based explanations of health simulations are one-size-fits-all and proposes a two-step framework to elicit stakeholder-specific content and style preferences and to steer LLM summarization accordingly. Step 1 decomposes simulation models, generates candidate summaries via a 2x2 factorial design over content and style, checks factuality with modelers, and pilots the summaries with stakeholders using empathy questionnaires. Step 2 analyzes factorial data per stakeholder group, steers LLM generation via preference optimization, evaluates with LLM-based and human metrics, and shares optimized summaries back to participants. The authors explicitly position the contribution as a vision and framework, and they defer empirical validation to future work, including ablation studies and a downstream randomized controlled trial.

Significance. If the framework works as intended, it would address a real gap by making health simulation outputs accessible to diverse stakeholders in a tailored way. The paper's strengths are its clear articulation of a repeatable pipeline, the inclusion of a modeler factuality gate, the use of designed experiments for preference elicitation, and an unusually candid discussion of limitations, including an explicit call for ablation studies and downstream decision metrics. However, the claimed ability to elicit content and style needs is not demonstrated, and the key measurement choice (empathy) lacks construct-validity evidence. As submitted, the paper is best read as a research proposal rather than a validated method; its value will depend on the planned empirical studies.

major comments (4)
  1. [Step 1: Identify Information Needs and Preferred Styles From Different Stakeholder Groups] The central claim is that the framework 'elicit[s] ... style and content needs'; however, the only outcome measure specified for the stakeholder pilot is empathy (TEQ and State Empathy Scale). The paper does not justify why maximizing state empathy is equivalent to satisfying a stakeholder's information needs; indeed, the hospital-administrator example (executive summaries, bullet points, business-oriented language) suggests a summary can be appropriate without being maximally empathy-inducing. Because Step 2's factorial analysis and LLM steering optimize the measured outcome, this construct-validity gap is load-bearing. Please either add direct measures of perceived usefulness, comprehension, or decision quality, or reposition the framework as one for empathy-oriented summaries, and justify the use of a single affective outcome across heterogeneous stakeholder groups.
  2. [Step 1: Identify Information Needs and Preferred Styles From Different Stakeholder Groups] The text states: 'Since perceiving a narrative as immersive and compelling depends on the mental state of the reader, we recommend using the validated Toronto Empathy Questionnaire (TEQ; 16 items) to obtain multiple empathy measures.' TEQ is a trait empathy scale, not a state measure of narrative response; using it in this role is likely a misapplication. The State Empathy Scale is the appropriate state measure. Please clarify that TEQ is a baseline covariate and remove the implication that TEQ measures the reader's state in response to a summary.
  3. [Step 2: Optimize the Alignment of Language Models and Stakeholder Communication Needs] The optimization loop is underspecified. The text says 'An optimization process is involved, as the new summaries should be automatically assessed and the architecture adjusted if the scores are insufficient,' but no objective function, score thresholds, or adjustment rules are given. For a framework that claims to be a repeatable process, this prevents replication. Specify at a conceptual level what is optimized (e.g., which evaluation metrics serve as rewards), what 'insufficient' means, and which architectural parameters (RAG parameters, temperature, or others) are adjusted.
  4. [Discussion: On Participatory AI] The paper itself concedes that 'The scientific basis for this framework will be strengthened by collecting experimental data, performing ablation studies...' and that 'the ultimate demonstration that the pipeline works lies in its ability to affect decisions.' I agree. As submitted, the only empirical result is a pilot with n=7 reporting a median response time of 19.16 minutes, which does not test preference validity, summary quality, or decision impact. The authors should either explicitly scope the paper as a position/vision paper with no empirical claims, making the title and abstract match, or include a proof-of-concept with a small but substantive evaluation.
minor comments (6)
  1. [Proposed Framework, Step 1] The factorial design is referred to as '22 factorial design'; this should be typeset as 2^2 factorial design.
  2. [Proposed Framework, Step 2] Typo: 'repeated measures ANOV A' should read 'repeated measures ANOVA'.
  3. [Author affiliations] There is a spacing issue in 'Old Dominion University / 1030 University Blvd, Suffolk, V A 23435, USA' — 'V A' should be 'VA'.
  4. [References] The reference 'Ahrweiler et al. 2019' has 'policy advise'; likely should be 'policy advice'.
  5. [Figure 3] The feedback loop between 'automated assessment' and 'architecture adjustment' would be easier to follow if the figure annotated the specific metrics and parameters involved.
  6. [Proposed Framework, Step 1] The internal loop that asks participants for their preferences and then evaluates whether summaries match those preferences is not inherently circular, because modeler factuality checks and the proposed downstream RCT provide independent anchors. The paper would benefit from stating this explicitly to preempt concerns about self-reference.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the paper is a framework proposal with an iterative feedback loop and an external RCT anchor; self-citations are used as ordinary building blocks, not as load-bearing self-justification.

full rationale

The paper does not derive quantitative predictions or present equations that reduce to fitted parameters; it proposes a mixed-methods framework. The central loop—generating candidate summaries, measuring participant reactions, factorial-analyzing those reactions, steering the LLM, and re-assessing—is an iterative user-centered design process, not a claim that a prediction is validated by its own inputs. The only pilot (n=7) is reported as response time, not as evidence that empathy scores validate content/style preferences, and the paper explicitly defers causal validation to a future randomized controlled trial measuring downstream decisions, providing an external anchor. Self-citations (e.g., Gandee and Giabbanelli 2024 for model-to-text decomposition; Giabbanelli et al. 2024a for few-shot saturation; Giabbanelli et al. 2025 for empathy-based storytelling) are used as prior empirical building blocks with alternatives listed (e.g., RDF Walks), so they are not load-bearing uniqueness arguments. The construct-validity concern—that empathy questionnaires are used as a proxy for information needs—is a substantive measurement assumption, but it is not circular in the logical sense: 'preferred' is operationalized via empathy, but this is an explicit modeling choice, not a derivation that makes the framework's output equivalent to its input. The paper itself acknowledges its limitations, stating that 'The scientific basis for this framework will be strengthened by collecting experimental data, performing ablation studies...' (Discussion). Thus no circular step meets the evidentiary bar required by the review rules.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities; four domain assumptions carry the framework, all untested. The most fragile is that stakeholder preferences are stable and elicitable with the chosen instruments.

assumptions (4)
  • domain assumption Distinct stakeholder groups have stable, divergent preferences for summary content and style that can be elicited and generalized.
    The framework's Step 1 and Step 2 depend on stable group-level preferences; the paper motivates this with the hospital layout example but provides no data.
  • domain assumption The Toronto Empathy Questionnaire and State Empathy Scale are valid proxies for whether a summary is understandable and actionable.
    Step 1 recommends TEQ and State Empathy Scale as time-efficient instruments; their predictive validity for summary-driven decision-making is assumed.
  • ad hoc to paper A 2x2 factorial design over content and style captures the variation in stakeholder needs.
    Step 1 states 'if each controllable aspect is simplified by two options, then we have 22 factorial design'; this restricts the preference space to two binary attributes.
  • domain assumption LLMs can produce factually correct summaries if model decomposition, few-shot fine-tuning, and modeler review are applied.
    The framework leans on prior model-to-text work by the authors and others; factuality is checked after generation, not guaranteed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization." pith.science (2026). https://pith.science/paper/IBEYBLQW

@misc{pith2026250904646,
  author       = {Pith},
  title        = {Pith review of: Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IBEYBLQW}},
  note         = {Machine review of arXiv:2509.04646}
}
read the original abstract

Modeling & Simulation (M&S) approaches such as agent-based models hold significant potential to support decision-making activities in health, with recent examples including the adoption of vaccines, and a vast literature on healthy eating behaviors and physical activity behaviors. These models are potentially usable by different stakeholder groups, as they support policy-makers to estimate the consequences of potential interventions and they can guide individuals in making healthy choices in complex environments. However, this potential may not be fully realized because of the models' complexity, which makes them inaccessible to the stakeholders who could benefit the most. While Large Language Models (LLMs) can translate simulation outputs and the design of models into text, current approaches typically rely on one-size-fits-all summaries that fail to reflect the varied informational needs and stylistic preferences of clinicians, policymakers, patients, caregivers, and health advocates. This limitation stems from a fundamental gap: we lack a systematic understanding of what these stakeholders need from explanations and how to tailor them accordingly. To address this gap, we present a step-by-step framework to identify stakeholder needs and guide LLMs in generating tailored explanations of health simulations. Our procedure uses a mixed-methods design by first eliciting the explanation needs and stylistic preferences of diverse health stakeholders, then optimizing the ability of LLMs to generate tailored outputs (e.g., via controllable attribute tuning), and then evaluating through a comprehensive range of metrics to further improve the tailored generation of summaries.

Figures

Figures reproduced from arXiv: 2509.04646 by the authors.

Figure 1
Figure 1. A model consists of elements and interrelationships from the problem domain, exemplified here as suicide prevention. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. A simulation model has (1) a static structure, which can be decomposed and transformed into text using existing [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Our second step analyses the survey data to find what each group needs in a summary. Then, we steer LLMs in pro [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

93 extracted references · 70 canonical work pages

  1. [1]

    Ahmed, R.; and Hemanth, D. J. 2025. Hybrid text summarization: Integrating extractive and abstractive models for enhanced cross-domain summarization. Intelligent Decision Technologies, 18724981251322745

  2. [2]

    M.; Capellas, B

    Ahrweiler, P.; Sp \"a th, E.; Siqueiros Garc \' a, J. M.; Capellas, B. L.; and Wurster, D. 2025. Inclusive technology co-design for participatory AI. Participatory Artificial Intelligence in Public Social Services: From Bias to Fairness in Assessing Beneficiaries, 35--62

  3. [3]

    Ahrweiler, P.; et al. 2019. Co-designing social simulation models for policy advise: lessons learned from the INFSO-SKIN study. In 2019 Spring simulation conference (SpringSim), 1--12. IEEE

  4. [4]

    Aminpour, P.; Schwermer, H.; and Gray, S. 2021. Do social identity and cognitive diversity correlate in environmental stakeholders? A novel approach to measuring cognitive distance within and between groups. Plos one, 16(11): e0244907

  5. [5]

    D.; Papandrianos, N

    Apostolopoulos, I. D.; Papandrianos, N. I.; Papathanasiou, N. D.; and Papageorgiou, E. I. 2024. Fuzzy cognitive map applications in medicine over the last two decades: A review study. Bioengineering, 11(2): 139

  6. [6]

    L.; and Giabbanelli, P

    Baniukiewicz, M.; Dick, Z. L.; and Giabbanelli, P. J. 2018. Capturing the fast-food landscape in England using large-scale network analysis. EPJ Data Science, 7(1): 39

  7. [7]

    A.; Wornow, M.; Swaminathan, A.; Lehmann, L

    Bedi, S.; Liu, Y.; Orr-Ewing, L.; Dash, D.; Koyejo, S.; Callahan, A.; Fries, J. A.; Wornow, M.; Swaminathan, A.; Lehmann, L. S.; et al. 2025. Testing and evaluation of health care applications of large language models: a systematic review. Jama

  8. [8]

    K.; Zaghir, J.; Zheng, Y.; Bensahla, A.; Bjelogrlic, M.; and Lovis, C

    Bednarczyk, L.; Reichenpfader, D.; Gaudet-Blavignac, C.; Ette, A. K.; Zaghir, J.; Zheng, Y.; Bensahla, A.; Bjelogrlic, M.; and Lovis, C. 2025. Scientific evidence for clinical text summarization using large language models: scoping review. Journal of Medical Internet Research, 27: e68998

Show all 93 references
  1. [9]

    Belfrage, M.; et al. 2024. Simulating change: A systematic literature review of agent-based models for policy-making. In 2024 Annual Modeling and Simulation Conference (ANNSIM), 1--13. IEEE

  2. [10]

    J.; and Tobias, A

    Brooks, R. J.; and Tobias, A. M. 1996. Choosing the best model: Level of detail, complexity, and model performance. Mathematical and computer modelling, 24(4): 1--14

  3. [11]

    H.; Kader, R.; Ortiz-Prado, E.; Makowski, M

    Busch, F.; Hoffmann, L.; Rueger, C.; van Dijk, E. H.; Kader, R.; Ortiz-Prado, E.; Makowski, M. R.; Saba, L.; Hadamitzky, M.; Kather, J. N.; et al. 2025. Current applications and challenges in large language models for patient care: a systematic review. Communications Medicine,...

  4. [12]

    J.; and Gotz, D

    Caban, J. J.; and Gotz, D. 2015. Visual analytics in healthcare--opportunities and research challenges. Journal of the American Medical Informatics Association, 22(2): 260--262

  5. [13]

    Y.; et al

    Chu, S. Y.; et al. 2025. Think together and work better: Combining humans' and LLMs' think-aloud outcomes for effective text evaluation. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1--23

  6. [14]

    N.; and Buchmann, R

    Dolha, D. N.; and Buchmann, R. A. 2024. Generative AI for BPMN process analysis: experiments with multi-modal process representations. In International Conference on Business Informatics Research, 19--35. Springer

  7. [15]

    Dong, H.; Xiong, W.; Goyal, D.; Zhang, Y.; Chow, W.; Pan, R.; Diao, S.; Zhang, J.; Shum, K.; and Zhang, T. 2023. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767

  8. [16]

    C.; et al

    Dos Santos, V. C.; et al. 2025. Enhancing healthcare operations: a systematic literature review on approaches for hospital facility layout planning. Journal of Health Organization and Management, 39(1): 22--45

  9. [17]

    Fahland, D.; Fournier, F.; Limonad, L.; Skarbovsky, I.; and Swevels, A. J. 2024. How well can large language models explain business processes? arXiv preprint arXiv:2401.12846

  10. [18]

    Fedeli, A.; and Manrique Negrin, D. A. 2024. Towards a collaborative approach for Digital Twin simulation models comprehension. In Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, MODELS Companion '24, 660–664. New Yo...

  11. [19]

    Ferrand, N.; Hassenforder, E.; and Girard, S. 2024. Engineering participation: Preparing and designing a participatory process. Transformative Participation for Socio-Ecological Sustainability-Around the CoOPLAGE pathways, 109--121

  12. [20]

    W.; Triplett, Z.; Asif, N.; Susanto, A.; Chowdhury, A.; Azcoaga Lorenzo, A.; Dras, M.; and Berkovsky, S

    Fraile Navarro, D.; Coiera, E.; Hambly, T. W.; Triplett, Z.; Asif, N.; Susanto, A.; Chowdhury, A.; Azcoaga Lorenzo, A.; Dras, M.; and Berkovsky, S. 2025. Expert evaluation of large language models for clinical dialogue summarization. Scientific reports, 15(1): 1195

  13. [21]

    J.; and Giabbanelli, P

    Gandee, T. J.; and Giabbanelli, P. J. 2024. Combining natural language generation and graph algorithms to explain causal maps through meaningful paragraphs. In International Conference on Conceptual Modeling, 359--376. Springer

  14. [22]

    J.; et al

    Gandee, T. J.; et al. 2024. A Visual Analytics Environment for Navigating Large Conceptual Models by Leveraging Generative Artificial Intelligence. Mathematics, 12(13): 1946

  15. [23]

    Gao, M.; Ruan, J.; Sun, R.; Yin, X.; Yang, S.; and Wan, X. 2023. Human-like summarization evaluation with chatgpt. arXiv preprint arXiv:2304.02554

  16. [24]

    Ghaffarzadegan, N.; Lyneis, J.; and Richardson, G. P. 2011. How small system dynamics models can help the public policy process. System Dynamics Review, 27(1): 22--44

  17. [25]

    Giabbanelli, P.; Phatak, A.; Mago, V.; and Agrawal, A. 2024 a . Narrating Causal Graphs with Large Language Models. In Hawaii International Conference on System Sciences 2024 (HICSS-57)

  18. [26]

    J.; and Baniukiewicz, M

    Giabbanelli, P. J.; and Baniukiewicz, M. 2018. Navigating complex systems for policymaking using simple software tools. In Advanced data analytics in health, 21--40. Springer

  19. [27]

    J.; and Baniukiewicz, M

    Giabbanelli, P. J.; and Baniukiewicz, M. 2019. Visual analytics to identify temporal patterns and variability in simulations from cellular automata. ACM Transactions on Modeling and Computer Simulation (TOMACS), 29(1): 1--26

  20. [28]

    J.; Daumas, C.; Flandre, N

    Giabbanelli, P. J.; Daumas, C.; Flandre, N. Y.; Pitkar, A.; and Vazquez-Estrada, J. 2025. Promoting empathy in decision-making by turning agent-based models into stories using large-language models. Journal of Simulation

  21. [29]

    J.; and Vesuvala, C

    Giabbanelli, P. J.; and Vesuvala, C. X. 2023. Human factors in leveraging systems science to shape public policy for obesity: A usability study. Information, 14(3): 196

  22. [30]

    J.; et al

    Giabbanelli, P. J.; et al. 2024 b . Broadening Access to Simulations for End-Users via Large Language Models: Challenges and Opportunities. In 2024 winter simulation conference (wsc), 2535--2546. IEEE

  23. [31]

    Gierend, K.; Kr \"u ger, F.; Genehr, S.; Hartmann, F.; Siegel, F.; Waltemath, D.; Ganslandt, T.; and Zeleke, A. A. 2024. Provenance information for biomedical data and workflows: Scoping review. Journal of medical Internet research, 26: e51297

  24. [32]

    Gravel, J.; D’Amours-Gravel, M.; and Osmanlliu, E. 2023. Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clinic Proceedings: Digital Health, 1(3): 226--234

  25. [33]

    A.; Gray, S.; Cox, L

    Gray, S. A.; Gray, S.; Cox, L. J.; and Henly-Shepard, S. 2013. Mental modeler: a fuzzy-logic cognitive mapping modeling tool for adaptive environmental management. In 2013 46th Hawaii international conference on system sciences, 965--973. IEEE

  26. [34]

    Haddad, E.; and Bugarin, K. 2020. Crisis Control: The Use of Simulations for Policy Decisionmaking. Policy Brief, PB 20, 38

  27. [35]

    a m \"a l \

    H \"a m \"a l \"a inen, R. P.; Luoma, J.; and Saarinen, E. 2013. On the importance of behavioral operational research: The case of understanding and communicating about dynamic systems. European Journal of Operational Research, 228(3): 623--634

  28. [36]

    Hassan, S.; Thompson, C.; Adams, J.; Chang, M.; Derbyshire, D.; Keeble, M.; Liu, B.; Mytton, O.; Rahilly, J.; Savory, B.; et al. 2024. The adoption and implementation of local government planning policy to manage hot food takeaways near schools in England: A qualitative proces...

  29. [37]

    Hayashi, H.; Budania, P.; Wang, P.; Ackerson, C.; Neervannan, R.; and Neubig, G. 2021. Wikiasp: A dataset for multi-domain aspect-based summarization. Transactions of the Association for Computational Linguistics, 9: 211--225

  30. [38]

    He, J.; Yang, Y.; Long, W.; Xiong, D.; Gutierrez-Basulto, V.; and Pan, J. Z. 2025. Evaluating and Improving Graph to Text Generation with Large Language Models. arXiv preprint arXiv:2501.14497

  31. [39]

    Hinrichs, M.; Wang, J.; Roe, C.; and Johnston, E. W. 2025. AI Integration in Mental Health Services: Examining Trends in the USA and Peoria, Illinois. In Participatory Artificial Intelligence in Public Social Services: From Bias to Fairness in Assessing Beneficiaries, 255--275...

  32. [40]

    C.; Ghumrawi, K

    Huddleston, J.; Galgoczy, M. C.; Ghumrawi, K. A.; Giabbanelli, P. J.; Rice, K. L.; Nataraj, N.; Brown, M. M.; Harper, C. R.; and Florence, C. S. 2022. Design and Deployment of a Simulation Platform: Case Study of an Agent-Based Model for Youth Suicide Prevention. In 2022 Winte...

  33. [41]

    Janssen, M.; and Helbig, N. 2018. Innovating and changing the policy-cycle: Policy-makers be prepared! Government Information Quarterly, 35(4): S99--S105

  34. [42]

    Jeong, D.; Aggarwal, S.; Robinson, J.; Kumar, N.; Spearot, A.; and Park, D. S. 2023. Exhaustive or exhausting? Evidence on respondent fatigue in long surveys. Journal of Development Economics, 161: 102992

  35. [43]

    Kammler, C.; et al. 2023. Towards a Social Simulation Interaction Tool for Policy Makers—A New Research Agenda to Enable Usage of More Complex Social Simulations. In Conference of the European Social Simulation Association, 163--176. Springer

  36. [44]

    Keeble, M.; Adams, J.; Amies-Cull, B.; Chang, M.; Cummins, S.; Derbyshire, D.; Hammond, D.; Hassan, S.; Liu, B.; Medina-Lara, A.; et al. 2024. Public acceptability of proposals to manage new takeaway food outlets near schools: cross-sectional analysis of the 2021 International...

  37. [45]

    J.; Timmons, S.; Luo, C.; and Shi, L

    Khademi, A.; Zhang, D.; Giabbanelli, P. J.; Timmons, S.; Luo, C.; and Shi, L. 2018. An agent-based model of healthy eating with applications to hypertension. In Advanced Data Analytics in Health, 43--58. Springer

  38. [46]

    P.; Gipp, B.; and Ruas, T

    Kirstein, F.; Wahle, J. P.; Gipp, B.; and Ruas, T. 2025. Cads: A systematic literature review on the challenges of abstractive dialogue summarization. Journal of Artificial Intelligence Research, 82: 313--365

  39. [47]

    A.; Giabbanelli, P

    Lavin, E. A.; Giabbanelli, P. J.; Stefanik, A. T.; Gray, S. A.; and Arlinghaus, R. 2018. Should we simulate mental models to assess whether they agree? In Proceedings of the annual simulation symposium, 1--12

  40. [48]

    Lee, H.; Phatale, S.; Mansoor, H.; Mesnard, T.; Ferret, J.; Lu, K.; Bishop, C.; Hall, E.; Carbune, V.; Rastogi, A.; et al. 2023. Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267

  41. [49]

    Lima, F. F. d.; and Os \'o rio, F. d. L. 2021. Empathy: assessment instruments and psychometric quality--a systematic literature review with a meta-analysis of the past ten years. Frontiers in psychology, 12: 781346

  42. [50]

    Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634

  43. [51]

    F.; Martinez-Moyano, I

    Luna-Reyes, L. F.; Martinez-Moyano, I. J.; Pardo, T. A.; Cresswell, A. M.; Andersen, D. F.; and Richardson, G. P. 2006. Anatomy of a group model-building intervention: Building dynamic theory from case study research. System Dynamics Review: The Journal of the System Dynamics ...

  44. [52]

    Manellanga, R.; and David, I. 2024. Participatory and collaborative modeling of sustainable systems: A systematic review. In Proceedings of the ACM/IEEE 27th international conference on model driven engineering languages and systems, 645--654

  45. [53]

    Mavridis, A.; Tegos, S.; Anastasiou, C.; Papoutsoglou, M.; and Meditskos, G. 2025. Large language models for intelligent RDF knowledge graph construction: results from medical ontology mapping. Frontiers in Artificial Intelligence, 8: 1546179

  46. [54]

    Montibeller, G. 2018. Behavioral challenges in policy analysis with conflicting objectives. In Recent advances in optimization and modeling of contemporary problems, 85--108. INFORMS

  47. [55]

    Mussa, O.; Rana, O.; Goossens, B.; Orozco-terWengel, P.; and Perera, C. 2024. Towards Enhancing Linked Data Retrieval in Conversational UIs Using Large Language Models. In International Conference on Web Information Systems Engineering, 246--261. Springer

  48. [56]

    Nezhad, B.; et al. 2025. Fair Summarization: Bridging Quality and Diversity in Extractive Summaries. In Prabhakaran, V.; Dev, S.; Benotti, L.; Hershcovich, D.; Cao, Y.; Zhou, L.; Cabello, L.; and Adebara, I., eds., Proceedings of the 3rd Workshop on Cross-Cultural Consideratio...

  49. [57]

    Olabisi, O.; and Agrawal, A. 2024. Understanding Position Bias Effects on Fairness in Social Multi-Document Summarization. In Scherrer, Y.; Jauhiainen, T.; Ljube s i \'c , N.; Zampieri, M.; Nakov, P.; and Tiedemann, J., eds., Proceedings of the Eleventh Workshop on NLP for Sim...

  50. [58]

    U.; Soroush, A.; Sakhuja, A.; Freeman, R.; Horowitz, C

    Omar, M.; Sorin, V.; Agbareia, R.; Apakama, D. U.; Soroush, A.; Sakhuja, A.; Freeman, R.; Horowitz, C. R.; Richardson, L. D.; Nadkarni, G. N.; et al. 2025. Evaluating and addressing demographic disparities in medical large language models: a systematic review. International Jo...

  51. [59]

    C.; Burgoine, T.; Sharp, S

    Patterson, R.; Ogilvie, D.; Hoenink, J. C.; Burgoine, T.; Sharp, S. J.; Hajna, S.; and Panter, J. 2025. Combined associations of takeaway food availability and walkability with adiposity: Cross-sectional and longitudinal analyses. Health & Place, 91: 103405

  52. [60]

    Peters, U.; and Chin-Yee, B. 2025. Generalization Bias in Large Language Model Summarization of Scientific Research. arXiv:2504.00025

  53. [61]

    D.; Mancino, M

    Ponzo, V.; Goitre, I.; Favaro, E.; Merlo, F. D.; Mancino, M. V.; Riso, S.; and Bo, S. 2024. Is ChatGPT an effective tool for providing dietary advice? Nutrients, 16(4): 469

  54. [62]

    D.; and Finn, C

    Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct preference optimization: your language model is secretly a reward model. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23. Red Hoo...

  55. [63]

    F.; Bansal, M.; and Dreyer, M

    Ribeiro, L. F.; Bansal, M.; and Dreyer, M. 2023. Generating summaries with controllable readability levels. arXiv preprint arXiv:2310.10623

  56. [64]

    Robinson, S.; and Brooks, R. 2024. Assumptions and simplifications in discrete-event simulation modelling. Journal of Simulation, 1--18

  57. [65]

    Z.; et al

    Sarmiento, I.; Cockcroft, A.; Dion, A.; Belaid, L.; Silver, H.; Pizarro, K.; Pimentel, J.; Tratt, E.; Skerritt, L.; Ghadirian, M. Z.; et al. 2024. Fuzzy cognitive mapping in participatory research and decision making: a practice review. Archives of Public Health, 82(1): 76

  58. [66]

    Savory, B.; Thompson, C.; Hassan, S.; Adams, J.; Amies-Cull, B.; Chang, M.; Derbyshire, D.; Keeble, M.; Liu, B.; Medina-Lara, A.; et al. 2025. ``It does help but there's a limit...'': Young people's perspectives on policies to manage hot food takeaways opening near schools. So...

  59. [67]

    Schaaff, K.; Reinig, C.; and Schlippe, T. 2023. Exploring ChatGPT’s empathic abilities. In 2023 11th international conference on affective computing and intelligent interaction (ACII), 1--8. IEEE

  60. [68]

    B.; Zhao, Z.; Sayin, B.; Flek, L.; and Rosso, P

    Schlicht, I. B.; Zhao, Z.; Sayin, B.; Flek, L.; and Rosso, P. 2025. Do LLMs provide consistent answers to health-related questions across languages? In European Conference on Information Retrieval, 314--322. Springer

  61. [69]

    W.; Breazeal, C.; and Sap, M

    Shen, J.; Mire, J.; Park, H. W.; Breazeal, C.; and Sap, M. 2024. Heart-felt narratives: Tracing empathy and narrative style in personal stories with llms. arXiv preprint arXiv:2405.17633

  62. [70]

    Shool, S.; Adimi, S.; Saboori Amleshi, R.; Bitaraf, E.; Golpira, R.; and Tara, M. 2025. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Medical Informatics and Decision Making, 25(1): 117

  63. [71]

    A.; and Giabbanelli, P

    Shrestha, A.; Mielke, K.; Nguyen, T. A.; and Giabbanelli, P. J. 2022. Automatically explaining a model: Using deep neural networks to generate text from causal maps. In 2022 Winter simulation conference (WSC), 2629--2640. IEEE

  64. [72]

    R.; Cole-Lewis, H.; et al

    Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Amin, M.; Hou, L.; Clark, K.; Pfohl, S. R.; Cole-Lewis, H.; et al. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, 31(3): 943--950

  65. [73]

    N.; McKinnon, M

    Spreng, R. N.; McKinnon, M. C.; Mar, R. A.; and Levine, B. 2009. The Toronto Empathy Questionnaire: Scale development and initial validation of a factor-analytic solution to multiple empathy measures. Journal of personality assessment, 91(1): 62--71

  66. [74]

    St-Aubin, B.; Wainer, G.; and Loor, F. 2023. A survey of visualization capabilities for simulation environments. In 2023 Annual Modeling and Simulation Conference (ANNSIM), 13--24. IEEE Computer Society

  67. [75]

    D.; Lauf, S.; Magliocca, N

    Sun, Z.; Lorscheid, I.; Millington, J. D.; Lauf, S.; Magliocca, N. R.; Groeneveld, J.; Balbi, S.; Nolzen, H.; M \"u ller, B.; Schulze, J.; et al. 2016. Simple or complicated agent-based models? A complicated issue. Environmental Modelling & Software, 86: 56--67

  68. [76]

    Tran, D.; Dolgun, A.; and Demirhan, H. 2020. Weighted inter-rater agreement measures for ordinal outcomes. Communications in Statistics-Simulation and Computation, 49(4): 989--1003

  69. [77]

    Urlana, A.; Mishra, P.; Roy, T.; and Mishra, R. 2023. Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects--A Survey. arXiv preprint arXiv:2311.09212

  70. [78]

    van der Zee, D.-J. 2017. Approaches for simulation model simplification. In 2017 Winter Simulation Conference (WSC), 4197--4208. IEEE

  71. [79]

    H.; and Scheffer, M

    Van Nes, E. H.; and Scheffer, M. 2005. A strategy to improve the contribution of complex simulation models to ecological theory. Ecological modelling, 185(2-4): 153--164

  72. [80]

    P.; Seehofnerov \'a , A.; et al

    Van Veen, D.; Van Uden, C.; Blankemeier, L.; Delbrouck, J.-B.; Aali, A.; Bluethgen, C.; Pareek, A.; Polacin, M.; Reis, E. P.; Seehofnerov \'a , A.; et al. 2024. Adapted large language models can outperform medical experts in clinical text summarization. Nature medicine, 30(4):...

  73. [81]

    Wang, J.; Zhang, C.; Zhang, D.; Tong, H.; Yan, C.; and Jiang, C. 2025. A recent survey on controllable text generation: A causal perspective. Fundamental Research, 5(3): 1194--1203

  74. [82]

    Welivita, A.; and Pu, P. 2024. Is ChatGPT more empathetic than humans? arXiv preprint arXiv:2403.05572

  75. [83]

    K.; and Hosseinichimeh, N

    Wittenborn, A. K.; and Hosseinichimeh, N. 2022. Exploring personalized psychotherapy for depression: A system dynamics approach. Plos one, 17(10): e0276441

  76. [84]

    Ye, H.; Jin, J.; Xie, Y.; Zhang, X.; and Song, G. 2025. Large language model psychometrics: A systematic review of evaluation, validation, and enhancement. arXiv preprint arXiv:2505.08245

  77. [85]

    Yu, T.; Lin, T.-E.; Wu, Y.; Yang, M.; Huang, F.; and Li, Y. 2025. Diverse AI Feedback For Large Language Model Alignment. Transactions of the Association for Computational Linguistics, 13: 392--407

  78. [86]

    Yu, Y.-L. 2024. Disparities by race/ethnicity and immigration status in perceived importance of and access to culturally competent health care in the United States. Journal of Racial and Ethnic Health Disparities, 11(3): 1829--1841

  79. [87]

    L.; Massey, D.; Laboy, M.; O’Brien, D

    Zellner, M. L.; Massey, D.; Laboy, M.; O’Brien, D. T.; Mueller, A.; and Engelberg, D. 2025. Enhancing digital twin technology with community-led, science-driven participatory modeling: A case in green infrastructure planning. Environment and Planning B: Urban Analytics and Cit...

  80. [88]

    L.; Milz, D.; Lyons, L.; Hoch, C.; and Radinsky, J

    Zellner, M. L.; Milz, D.; Lyons, L.; Hoch, C.; and Radinsky, J. 2022. Finding the balance between simplicity and realism in participatory modeling for environmental planning. Environmental Modelling & Software, 157: 105481

  81. [89]

    S.; and Zhang, J

    Zhang, H.; Yu, P. S.; and Zhang, J. 2025. A systematic survey of text summarization: From statistical methods to large language models. ACM Computing Surveys, 57(11): 1--41

  82. [90]

    Zhao, Q.; Santos, E.; Nguyen, H.; and Mohamed, A. 2009. What makes a good summary? In Computational Methods for Counterterrorism, 33--50. Springer

  83. [91]

    Zhao, Y.; Zhang, J.; Chern, I.; Gao, S.; Liu, P.; He, J.; et al. 2023. Felm: Benchmarking factuality evaluation of large language models. Advances in Neural Information Processing Systems, 36: 44502--44523

  84. [92]

    Zhong, M.; Liu, Y.; Ge, S.; Mao, Y.; Jiao, Y.; Zhang, X.; Xu, Y.; Zhu, C.; Zeng, M.; and Han, J. 2022 a . Unsupervised multi-granularity summarization. arXiv preprint arXiv:2201.12502

  85. [93]

    Zhong, M.; Liu, Y.; Yin, D.; Mao, Y.; Jiao, Y.; Liu, P.; Zhu, C.; Ji, H.; and Han, J. 2022 b . Towards a unified multi-dimensional evaluator for text generation. arXiv preprint arXiv:2210.07197

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.