REVIEW 4 major objections 5 minor 2 cited by
Selecting the Right LLM for eGov Explanations
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The perceived quality of LLM-generated explanations can be measured and used to rank alternative LLMs for e-government services.
desk verdict A useful applied comparison of LLMs for eGov explanations, but the data only separate the worst model, so the promised ranking method is not yet demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is the adapted six-construct explanation-quality scale: six self-reported 1–7 Likert items grouped into fidelity (completeness, soundness, causability) and interpretability (clarity, compactness, comprehensibility), with trust and curiosity included as covariates after the original scale development. The scale converts a subjective impression of an explanation into a numeric vector that can be compared across LLMs through a multivariate analysis of covariance, turning 'which model explains best' into a testable difference in means. Explanations were generated with prompts embedding the tax-refund process model, causal execution dependencies, and, for an in-flight case, the executed trace log, so the comparison holds the generation context fixed and isolates the linguistic output of each model.
What would settle it
Have an independent team write different ground-truth explanations for the same three tax cases and repeat the 128-person survey; if the resulting ranking of the four LLMs changes materially, the elicited scores are an artifact of the reference narrative rather than a stable property of the models.
Extended reading notes
Core claim
The central claim, stated as Hypothesis 1, is that the perceived quality of LLM-generated explanations can be empirically elicited to determine the ranking among a set of alternative LLM model types. To test it, the authors adapted their earlier explanation-quality scale, rewording 24 statements to the tax-refund context, and had 128 citizens rate explanations generated by granite-3-8b-instruct, llama-3-1-70b, GPT-4o, and flan-ul2-20b for three inquiry cases. Controlling for trust, curiosity, digital literacy, and business-process expertise, the analysis showed a significant effect of LLM type on both fidelity and interpretability; contrast analysis found a significant gap only between the top and bottom models on fidelity and a mildly significant gap on interpretability. The authors therefore present the scale as a practical instrument for ranking LLMs, while noting that non-extreme pairs may not be clearly separable and that the final choice depends on a provider's weighting of the two quality dimensions.
Load-bearing premise
The comparison depends on the researchers' own 'ground truth' narratives being an unbiased reference standard for fidelity, and on the reused 24-statement scale still measuring the same six constructs after only wording changes.
Editorial extensions
If this is right
- An e-government provider can use the 24-statement survey to identify which LLM produces explanations citizens find most faithful and understandable, at least at the level of separating the weakest model from the stronger ones.
- The scale transfers to other e-government processes by rewording the statements, so a public authority that already runs citizen surveys does not need to build a new questionnaire from scratch.
- Trust and curiosity, not digital literacy or BPM expertise, moderated perceived quality, meaning providers should monitor citizens' general trust in AI rather than their technical skills.
- The final choice among closely rated models is an explicit trade-off between fidelity and interpretability; the scale makes that trade-off visible instead of forcing a single winner.
- Automated replacement of the survey is not yet reliable: on the current dataset, embeddings-based regression explained little variance and LLM-as-a-judge ratings correlated only weakly with human scores.
Reading between the lines
- Editorial extension: the reported contrasts suggest the strongest practical role of the scale is screening out a clearly weaker LLM, while ties among strong models would still require a provider's own weighting of fidelity versus interpretability.
- Editorial extension: the paper changes only the wording when adapting the scale, so a natural follow-up is to re-validate the six-construct structure on a different e-government service, such as a benefits or permit decision, where the stakes and vocabulary differ.
- Editorial extension: the weak LLM-as-a-judge correlations may still support a two-stage pipeline—machine scoring to shortlist candidates, human panels to calibrate finalists—though the current sparse data cannot establish that pipeline.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript adapts a previously developed six-construct scale (fidelity and interpretability subscales) to the e-government domain of tax-refund explanations, and conducts an online survey in which 128 participants rate explanations generated by four LLMs across three inquiry cases. The authors report MANCOVA results, pairwise contrasts, and exploratory attempts to automate human ratings via regression on embeddings and LLM-as-a-judge prompting. The central claim is that perceived explanation quality can be empirically elicited to rank alternative LLM model types, and that the resulting comparison provides a basis for choosing an LLM for e-government explanation generation.
Significance. If the statistical and measurement concerns are resolved, the paper would offer a concrete, openly documented template for comparing LLM-generated explanations in a public-service context. Its strengths include the use of a scale with prior validation in business-process explanation research, public availability of survey materials, prompts, and data on Zenodo, and transparent reporting of the limited predictive power of the automation attempts. The study also demonstrates large overall differences in perceived fidelity and interpretability across LLM types. However, the paper's inferential results currently support only a partial ordering, and the experiment's statistical analysis does not fully account for its mixed design, so the central ranking claim requires revision or additional evidence.
major comments (4)
- [Section IV, Section VI, Table II] The results do not support the full ranking stated in Hypothesis 1. The pairwise contrast analysis reported in Section VI shows significant differences only between the best and worst model for fidelity (p < .001) and only a mildly significant best-versus-worst difference for interpretability (p < .1). The adjusted means for granite-3-8b-instruct, llama-3-1-70b, and GPT-4o differ by roughly 0.1–0.2 on a 7-point scale, and no pairwise contrast among these three is reported as significant. The data therefore establish at most that flan-ul2-20b is perceived as worse on fidelity, not a full ranking among the four models. The authors should either revise Hypothesis 1 and the conclusions to state the achievable resolution as a partial ordering, or supply additional evidence such as equivalence tests, Bayesian model ranking, or a power analysis showing that the design could detect smaller differences.
- [Section V-B, Table II] The MANCOVA appears to treat the 256 ratings as independent observations (residual df = 248), although each of the 128 participants contributed two ratings for two explanations generated by different LLMs. Under a mixed design, ratings from the same participant are likely correlated, so the reported F-tests and pairwise contrasts may have biased standard errors. The authors should re-analyze the data with a repeated-measures or mixed-effects model that includes participant-level random effects, or at least report cluster-robust standard errors, and indicate whether the pattern of pairwise comparisons changes.
- [Section V-C and Section V-B] Fidelity is operationalized as agreement with a researcher-written 'ground truth' narrative, and the six-construct scale from prior work [7] is adapted with only wording changes. The manuscript does not report re-validation evidence for the adapted scale in the tax-refund domain, such as internal consistency, factor structure, or measurement invariance, nor any check on the adequacy of the ground-truth narratives, such as expert review or inter-rater agreement. Since the entire model comparison rests on these measurement decisions, the paper should provide such evidence or explicitly temper the claim that the scale is directly transferable to new e-government contexts.
- [Section V-A, Section VI] The sample is not representative of the general taxpayer population: participants were recruited through AI4GOV project partners and their networks, 76.6% hold graduate degrees, and 93% rated their digital literacy at 5 or higher on a 7-point scale. The statement in Section V-A that the sample's representativeness is 'ensured' by the fact that 97.7% are taxpayers is not sufficient. This limitation should be stated explicitly in the discussion, and the authors should consider a sensitivity analysis or an explicit argument about how selection on education and digital literacy could affect the LLM ranking.
minor comments (5)
- [Section VI, Figure 3] The text refers to 'interoperability' where 'interpretability' is meant; this typo also recurs in Section VII and should be corrected throughout.
- [Table II] The table is labeled 'MANCOVA' but reports univariate F-tests for fidelity and interpretability only; the multivariate test statistics (e.g., Pillai's trace, Wilks' lambda) should be reported, or the table relabeled as ANCOVA.
- [Section VII-B] The sentence reporting 'overall predictive power ... (i.e., 0.34 and 0.37, respectively)' is ambiguous because these values do not match the preceding R2 values of 0.132 and 0.117; please clarify what statistic these numbers represent.
- [Section VI] The phrase 'mildly significant' for the interpretability contrast is non-standard; please report the exact p-value and a confidence interval for the best-versus-worst comparison.
- [Figure 2] The prompt example contains the phrase 'national task refund process', which appears to be a typo for 'national tax refund process'.
Circularity Check
No significant circularity: fresh survey data drive the LLM-selection claim; the self-cited scale is an externally validated instrument, not an input that determines the outcome.
full rationale
The paper's central claim—that perceived explanation quality can be elicited to compare LLMs—is supported by a new controlled survey of 128 participants rating explanations generated by four LLM types, with the results analyzed via MANCOVA and pairwise contrasts. The adapted scale from [7] is self-cited, but that citation is not the evidence for the new comparison; the evidence is the newly collected ratings, which are reported transparently, including the limitation that pairwise contrasts separate only the best from the worst model. Because [7] is a published, externally validated instrument and the current paper does not fit scale parameters to the current outcomes and then reuse those fits to rank the same models, the self-citation does not create a definitional loop. The researcher-written 'ground truth' narratives are an explicit operationalization of fidelity, not a hidden reduction, and the LLM outputs were generated from process descriptions and causal dependencies rather than from the ground truth texts. The predictive-modeling section uses human ratings as supervised targets and honestly reports weak R-squared values; no fitted parameter is renamed as a prediction. The skeptic's observation that only the worst model is significantly separable is a statistical-strength and correctness concern, not a circularity concern. No load-bearing step reduces, by the paper's own equations or by self-citation, to its own inputs, so no circular step is identified.
Assumptions & free parameters
assumptions (4)
- domain assumption The explanation quality scale developed and validated in [7] remains valid after wording adaptations for the tax refund domain.
- domain assumption The researcher-written 'ground truth' narratives are the correct reference for fidelity.
- domain assumption Perceived explanation quality proxies trust and adoption of e-government services.
- standard math Standard MANCOVA assumptions (normality, homogeneity of variances) hold.
Cite this review
Pith. "Pith review of Selecting the Right LLM for eGov Explanations." pith.science (2026). https://pith.science/paper/46ADOU7U
@misc{pith2026250421032,
author = {Pith},
title = {Pith review of: Selecting the Right LLM for eGov Explanations},
year = {2026},
howpublished = {\url{https://pith.science/paper/46ADOU7U}},
note = {Machine review of arXiv:2504.21032}
}
read the original abstract
The perceived quality of the explanations accompanying e-government services is key to gaining trust in these institutions, consequently amplifying further usage of these services. Recent advances in generative AI, and concretely in Large Language Models (LLMs) allow the automation of such content articulations, eliciting explanations' interpretability and fidelity, and more generally, adapting content to various audiences. However, selecting the right LLM type for this has become a non-trivial task for e-government service providers. In this work, we adapted a previously developed scale to assist with this selection, providing a systematic approach for the comparative analysis of the perceived quality of explanations generated by various LLMs. We further demonstrated its applicability through the tax-return process, using it as an exemplar use case that could benefit from employing an LLM to generate explanations about tax refund decisions. This was attained through a user study with 128 survey respondents who were asked to rate different versions of LLM-generated explanations about tax refund decisions, providing a methodological basis for selecting the most appropriate LLM. Recognizing the practical challenges of conducting such a survey, we also began exploring the automation of this process by attempting to replicate human feedback using a selection of cutting-edge predictive techniques.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
On the Hybrid Nature of ABPMS Process Frames and its Implications on Automated Process Discovery
A process frame for AI-augmented BPM can be formalized as a set of overlapping procedural/declarative specifications, and 14 of 16 common procedural behaviors can be rediscovered as Declare constraints.
-
XABPs: Towards eXplainable Autonomous Business Processes
The paper defines a taxonomy of explainability for autonomous business processes and lists challenges that must be solved before such systems can be trusted and audited.
Reference graph
Works this paper leans on
-
[7]
How well can large language models explain business processes as perceived by users?
D. Fahland, F. Fournier, L. Limonad, I. Skarbovsky, and A. J. E. Swevels, “How well can large language models explain business processes as perceived by users?” Data & Knowledge Engineering , vol. 157, p. 102416, 2 2025. [Online]. Available: https://www.scienced irect.com/science/article/pii/S0169023X25000114
work page 2025
-
[1]
Factors affecting employees’ adoption of e- government in the Iraqi public education sector,
K. Alminshid and M. Omar, “Factors affecting employees’ adoption of e- government in the Iraqi public education sector,”Electronic Government, vol. 17, no. 2, 2021
work page 2021
-
[2]
Trust , Felt Trust and E- Government Adoption : A Theoretical Perspective,
A. Dashti, I. Benbasat, and A. Burton-jones, “Trust , Felt Trust and E- Government Adoption : A Theoretical Perspective,” Working Papers on Information Systems, vol. 10, no. 83, 2010
work page 2010
-
[3]
The utilization of e-government services: Citizen trust, innovation and acceptance factors,
L. Carter and F. B ´elanger, “The utilization of e-government services: Citizen trust, innovation and acceptance factors,” Information Systems Journal, vol. 15, no. 1, 2005
work page 2005
-
[4]
AI- augmented Business Process Management Systems: A Research Mani- festo,
M. Dumas, F. Fournier, L. Limonad, A. Marrella, and et al., “AI- augmented Business Process Management Systems: A Research Mani- festo,” ACM Transactions on Management Information Systems, vol. 14, no. 1, 2023
work page 2023
-
[5]
Trust in AI: progress, challenges, and future directions,
S. Afroogh, A. Akbari, E. Malone, M. Kargar, and H. Alambeigi, “Trust in AI: progress, challenges, and future directions,” Humanities and Social Sciences Communications , vol. 11, no. 1, p. 1568, 11 2024
work page 2024
-
[6]
N. S. Abdul Wahi and L. Berenyi, “Examining the Effect of Social Influence and Facilitating Conditions on E-government Adoption by Employees in Mandatory Condition,” in Proceedings of the Central and Eastern European eDem and eGov Days 2024 . New York, NY , USA: ACM, 9 2024, pp. 104–110
work page 2024
-
[8]
The WHY in Business Processes: Discovery of Causal Execution Dependencies,
F. Fournier, L. Limonad, I. Skarbovsky, and Y . David, “The WHY in Business Processes: Discovery of Causal Execution Dependencies,” K¨unstliche Intelligenz , 1 2025. [Online]. Available: https://rdcu.be/d52Qz
work page 2025
Show all 19 references
-
[9]
Business Process Management Architectures,
M. Weske, “Business Process Management Architectures,” in Business Process Management. Berlin, Heidelberg: Springer Berlin Heidelberg, 2019, pp. 351–384. [Online]. Available: http://link.springer.com/10.100 7/978-3-662-59432-2 8
2019
-
[10]
van der Aalst, Process Mining
W. van der Aalst, Process Mining. Berlin, Heidelberg: Springer, 2016. [Online]. Available: http://link.springer.com/10.1007/978-3-662-49851 -4
2016 doi
-
[11]
Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI),
A. Adadi and M. Berrada, “Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI),” IEEE Access, vol. 6, 2018
2018
-
[12]
Explainable Artificial Intelligence: Objectives, Stakeholders, and Future Research Opportuni- ties,
C. Meske, E. Bunde, J. Schneider, and M. Gersch, “Explainable Artificial Intelligence: Objectives, Stakeholders, and Future Research Opportuni- ties,” Information Systems Management , vol. 39, no. 1, pp. 53–63, 1 2022
2022
-
[13]
Towards Explainable Process Predictions for Industry 4.0 in the DFKI-Smart-Lego-Factory,
J. R. Rehse, N. Mehdiyev, and P. Fettke, “Towards Explainable Process Predictions for Industry 4.0 in the DFKI-Smart-Lego-Factory,” KI - Kunstliche Intelligenz, vol. 33, no. 2, 2019
2019
-
[14]
A survey of methods for explaining black box models,
R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, and D. Pedreschi, “A survey of methods for explaining black box models,” ACM Computing Surveys , vol. 51, no. 5, 2018
2018
-
[15]
Definition of Generative AI - Gartner Information Technology Glossary
“Definition of Generative AI - Gartner Information Technology Glossary.” [Online]. Available: https://www.gartner.com/en/information -technology/glossary/generative-ai
-
[16]
Hype Cycle for Generative AI (ID G00812271),
A. Chandrasekaran and L. Ramos, “Hype Cycle for Generative AI (ID G00812271),” Gartner, Tech. Rep., 7 2024
2024
-
[17]
Definition of Large Language Models (LLMs) - Gartner Information Technology Glossary
“Definition of Large Language Models (LLMs) - Gartner Information Technology Glossary.” [Online]. Available: https://www.gartner.com/en /information-technology/glossary/large-language-models-llm
-
[18]
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing,
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing,” ACM Computing Surveys, vol. 55, no. 9, 2023
2023
-
[19]
Explanation Quality Survey - the Tax Refund Case,
L. Limonad and F. Fournier, “Explanation Quality Survey - the Tax Refund Case,” 1 2025. [Online]. Available: https://zenodo.org/records/1 4637610
2025
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.