{"id":"26bef68b-d9ca-41b6-b536-5f0fcd85b714","arxiv_id":"2502.00015","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic map of 39 papers shows LLM ethics concerns cluster into five dimensions, and most mitigation strategies remain unevaluated.","lead":"This paper maps 39 studies on ethical problems with large language models, grouping them into five concerns: safety, privacy, transparency, bias, and accountability. It finds that most proposed fixes have not been tested in practice, and implementation lags in areas like healthcare and government.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5/39 'fully evaluated' statistic is load-bearing for the central claim, but the evaluation-status coding has no reported inter-rater reliability and conflicts with the contribution statement's 13-study count; the count needs independent verification.","rationale":"The reader's weakest_assumption identifies the five-dimension coding frame as the main risk. That is a real concern for RQ1 counts, since forcing fairness under bias, autonomy under accountability, and consent under privacy may distort the reported distributions. However, the central claim's distinct contribution rests more directly on the evaluation-status gap and the implementation challenges. The 5/39 fully-evaluated statistic is the only quantitative anchor for the abstract's conclusion, and the paper's internal inconsistency between Section 1's 13-study count and Section 4.3's 5-study count is evidence that this coding is not stable. I do not claim the authors are dishonest; mapping-study coding is inherently subjective, and the paper includes some validity-mitigation efforts, such as librarian consultation, pilot tests, and consensus meetings. The broad qualitative conclusion that many mitigation strategies are unevaluated is plausible and consistent with prior work. Still, because the central quantitative claim is not auditable without the coding matrix and an inter-rater reliability check, the appropriate verdict remains conditional: the paper should be accepted only after the authors provide the per-study evaluation codes, reconcile the 13 vs 5 counts, and demonstrate coding reliability. This does not move the reader's verdict, which already called for conditional acceptance, but it sharpens the specific condition that must be met.","tokens_in":29483,"tokens_out":4391,"duration_ms":46561,"concrete_test":"Release the per-study evaluation-status coding matrix and have two independent raters, blind to the authors' labels, re-code all 39 primary studies using the Section 4.3 rubric. Compute Cohen's kappa and compare the fully-evaluated count. Also require a reconciliation table for the 13 vs 5 counts, e.g. identifying the 8 partial evaluations treated as not evaluated. If the re-coded fully-evaluated count is not 5, or if rater agreement is below 0.6, the abstract's quantitative claim should be revised or restricted to 'reported evaluation status' rather than presented as an established finding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central conclusion, that ethical issues hinder practical implementation of mitigation strategies and that existing frameworks lack adaptability, is anchored by the Section 4.3 quantitative claim that mitigation strategies and recommendations were fully evaluated in only 5/39 (12.8%) studies. This statistic is load-bearing: if it is wrong, the paper's distinctive contribution over prior mapping studies, e.g. Atlam et al., is substantially weakened. The paper reports no inter-rater reliability for this coding. Section 3.2.3 only states that the first and third authors extracted data, other authors reviewed the extractions, and discrepancies were resolved through consensus. Section 4.3 defines a rubric with 'Not evaluated', 'Fully evaluated', and partial treated as not evaluated, but no per-study coding table or extraction artifact is provided, so a reader cannot verify which 5 studies are fully evaluated. There is also an internal tension: Section 1 states that 13 of 39 papers conducted some form of empirical evaluation, while Section 4.3 reports only 5 fully evaluated. These could be reconciled if 8 studies were partial, but the paper never shows that reconciliation explicitly. Without a reproducible coding matrix, the headline statistic and the downstream claims about evaluation gaps are not auditable. This is a threat to construct validity, not an allegation of misconduct; coding in mapping studies is subjective and can be corrected, but the current manuscript leaves the load-bearing number unverified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic mapping study of 39 primary studies (2020–2024) on ethical concerns of large language models. The authors define five ethical dimensions (safety, privacy, transparency, bias, accountability) derived from four industry/governmental frameworks, map the selected studies and the 130 identified ethical issues onto these dimensions, categorize mitigation strategies into five themes, and analyze implementation challenges grouped into six themes. The paper's headline findings are that only 5/39 (12.8%) mitigation strategies were fully evaluated and that ethical issues hinder practical implementation, especially in healthcare and public governance, with existing frameworks lacking adaptability.","tokens_in":29703,"tokens_out":8056,"duration_ms":70600,"significance":"If the results hold, the study makes a useful contribution to the AI-ethics secondary literature: it is cross-domain, includes both peer-reviewed studies and grey literature, and explicitly codes the empirical evaluation status of mitigation strategies and implementation challenges. The authors are transparent about their search protocol, inclusion/exclusion criteria, and use of snowballing. The most distinctive quantitative claim, the 5/39 fully-evaluated rate, would differentiate this mapping from earlier reviews such as Atlam et al. However, because that statistic and the RQ1 dimension counts depend on coding decisions that are not fully auditable, the strength of the contribution rests on the revisions described below.","major_comments":[{"comment":"The paper's distinctive claim that mitigation strategies are rarely evaluated is internally inconsistent and not auditable. Section 1 states that 13 of the 39 studies \"conducted some form of empirical evaluation,\" while Section 4.3 reports that strategies were \"fully evaluated only in 5/39 (12.8%) studies\" and that partially evaluated strategies are treated as not evaluated. The difference of 8 studies is never explicitly reconciled, and no appendix or table identifies which 5 studies were fully evaluated, which 8 partially evaluated, and which 26 not evaluated. Because the abstract and conclusion use the evaluation gap as a load-bearing result, the manuscript should provide a per-study evaluation-status table and an explicit reconciliation of the 13-study and 5-study counts.","section":"§1 vs §4.3"},{"comment":"No inter-rater reliability is reported for the coding that produces the headline statistics. The data-analysis description states that the first and third authors extracted data and that discrepancies were resolved by consensus, and Section 4.3 says the authors independently applied the evaluation rubric, but no agreement metric (e.g., Cohen's kappa) or coding artifact is provided. For a mapping study whose main contributions are counts of ethical dimensions and evaluation status, the absence of reliability evidence is a substantial construct-validity threat; at minimum, the authors should report agreement rates and provide the coding matrix as an appendix or repository.","section":"§3.2.3, §4.3"},{"comment":"The RQ1 dimension counts are partly circular. The authors state that the five dimensions were selected because they are \"the most frequently emphasized across the four major frameworks\" and \"also emerged as the most recurrent themes during our coding,\" and they then use these dimensions as the coding frame for all 39 studies. This makes the prominence of the five dimensions in Table 3 and Table 5 an artifact of the chosen frame rather than an independent finding about the literature. The folding of fairness into bias, autonomy into accountability, and consent into privacy is reasonable, but the paper should either present the coding frame as a sensitivity analysis or explicitly temper claims such as \"accountability is rarely addressed in the education and public safety domains\" and \"transparency ... completely absent from the cybersecurity literature\" (Section 4.1) so that they are stated relative to the chosen framework, not as absolute properties of the field.","section":"§4.2"},{"comment":"Table 3 is inconsistent with the domain list in Section 4.1. Section 4.1 describes an Economics domain, but Table 3 contains no Economics row; the table also does not list several primary studies that appear in the paper (e.g., P12 and P36), and some papers appear in more than one row (P1, P2, P27) without any note on how multiple domain assignments were handled. Because Section 5.1 draws conclusions about the contextual significance of ethical dimensions from these per-domain counts, the domain mapping needs to be completed and made reproducible.","section":"Table 3 vs §4.1"}],"minor_comments":[{"comment":"The internal-validity paragraph states that the timeframe for the SMS was \"between 2023 and July 2024,\" but Section 3.2.2 reports that the search was conducted between April and May 2024 and Figure 3 shows selected studies from 2020 to 2024. This should be corrected to reflect the actual publication window.","section":"§6"},{"comment":"The figure references appear to be off by one: \"as illustrated in Figures 2 and 3\" and \"The bar chart in Figure 2\" should refer to Figure 3 (distribution by year) and Figure 4 (distribution by publication), respectively; Figure 2 is already used for the data-analysis process in Section 3.2.3.","section":"§4.1"},{"comment":"The cross-reference \"Table 4.2.1 shows the primary studies addressing each dimension\" should be to Table 5 (Ethical Dimensions Identified from papers).","section":"§4.2"},{"comment":"In the list for \"User Empowerment and Transparency in AI Interactions,\" the entry \"P1, P2 P2, P5\" contains a duplicated \"P2\"; please correct.","section":"§4.3"},{"comment":"The phrase \"shown in 9\" should be \"shown in Appendix B\" (the database search strings are presented in Section 9).","section":"§3.2.1"},{"comment":"The guidelines author is Petersen et al., not \"Peterson et al.,\" in the methodology text; the reference list already uses the correct spelling.","section":"§3 / references"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an AI-ethics or software-engineering journal, but the methodological transparency expectations for a systematic mapping study are high. The missing coding matrix and the unresolved 13-vs-5 evaluation count are the main reasons for major revision; these are fixable within the manuscript's scope. No concerns about citation practices or novelty disclosure beyond what is stated in the report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the synthesis: 39 studies across domains, plus industry and government guidelines, with a systematic look at whether mitigation strategies are actually evaluated. Prior reviews were healthcare-only (Li), high-level (Atlam), or tool-focused without real-world assessment (Morley). This paper closes that gap, and the main qualitative conclusion—ethical concerns are multi-dimensional and context-dependent, and mitigation is rarely backed by empirical evidence—is consistent with what the reported counts show.\n\nThe method is mostly sound: six databases, snowballing, grey literature, a clear rubric for evaluation status, and the authors are honest that 26 of 39 studies are conceptual and that the field skews Western. The thematic coding (17 themes under five dimensions) is reasonable and the RQ3 discussion of implementation challenges is a useful practical contribution.\n\nNow the soft spots. The stress-test note lands: 5/39 (12.8%) fully evaluated is load-bearing for the central claim, but the paper gives no per-study coding table, no inter-rater reliability figure, and the introduction says 13 of 39 conducted some empirical evaluation. Those numbers can be reconciled (8 partial, treated as not evaluated), but the paper never shows that. Without a reproducible coding matrix, the headline statistic is not auditable. This is a construct-validity problem, not fraud, and it is fixable by shipping the extraction artifact and the reconciliation. Minor: the timeframe is inconsistent—search was April–May 2024 in one place, July 2024 cutoff in the validity section, June 2024 in Figure 3. Also, the five dimensions are pre-selected from frameworks and then used to code the literature, so the prominence counts are partly circular; a reliability check would address the worst of that.\n\nWho is this for? Researchers and policymakers wanting a current map of LLM ethics concerns, mitigation strategies, and where the evaluation gaps are. It does not resolve an open question or change practice, but it is a competent synthesis. It deserves a serious referee: the method is largely sound, the synthesis is genuinely useful, and the flaws are correctable. Send it to review, with a request for the coding data and clarification of the evaluation-status counts.","headline":"Useful cross-domain mapping of LLM ethics, but the headline 5/39 evaluation statistic needs an audit trail before it can carry the weight the paper puts on it.","tokens_in":30258,"tokens_out":1640,"would_cite":true,"duration_ms":17571,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic mapping of 39 studies finds that LLM ethics mitigation strategies are mostly conceptual, with only 5 of 39 fully evaluated in practice.","keywords":["generative AI ethics","large language models","systematic mapping study","mitigation strategies","ethical dimensions","empirical evaluation","AI governance","AI ethics frameworks"],"falsifier":"Re-code the same 39 papers with an independent team using the same five dimensions and report inter-rater agreement; or re-run the search in the same six databases and check whether the 39-study set is reproduced. If agreement is low, if a sixth dimension such as fairness or autonomy materially changes the domain counts, or if the snowballing step finds many additional eligible studies, then the reported gaps are properties of the coding choices rather than of the literature itself.","tokens_in":29280,"feed_emoji":"⚖️","tokens_out":6417,"duration_ms":57670,"temperature":0.7,"pith_summary":"The paper sets out to map, across application domains, the ethical concerns raised by generative AI and large language models, the strategies proposed to mitigate them, and the barriers to putting those strategies into practice. It reviews 39 primary studies and codes their concerns through five ethical dimensions: safety, privacy, transparency, bias, and accountability. The central finding is that ethical concerns are multi-dimensional and context-dependent, and that proposed mitigation strategies are overwhelmingly unevaluated: only 5 of 39 studies (12.8 percent) fully evaluate their proposed strategy or recommendation. The paper argues that implementation challenges are most acute in high-stakes domains such as healthcare and public governance, and that existing ethical frameworks lack the adaptability to track evolving societal expectations. A sympathetic reader would care because the map identifies where the field's remedies are untested and where future empirical work is most needed.","feed_headline":"Only 5 in 39 AI ethics fixes see full testing","feed_subtitle":"A systematic map of 39 studies finds LLM safeguards are mostly conceptual and domain gaps are wide.","key_machinery":"The central instrument is a five-dimensional coding scheme—safety, privacy, transparency, bias, and accountability—derived from four selected international guidelines and regulatory frameworks and used to classify every ethical concern, mitigation strategy, and implementation challenge in the 39 primary studies. A second instrument is an evaluation-status rubric that distinguishes strategies that were not evaluated, partially evaluated, or fully evaluated, applied to determine whether a proposed mitigation has real empirical support. These two instruments together produce the paper's counts, the domain-by-dimension gap table, and the conclusion that most mitigation strategies remain conceptual.","core_discovery":"On the paper's own terms, the discovery is a structured evidence-base: across 39 studies published from 2020 to 2024, the ethical concerns of LLM use cluster into five dimensions—safety, privacy, transparency, bias, and accountability—and every proposed mitigation strategy or recommendation can be classified by whether it was not evaluated, partially evaluated, or fully evaluated. The result is that 26 of 39 studies are conceptual, only 13 conduct some empirical evaluation, and only 5 of 39 fully evaluate a mitigation strategy. Domain-level coding shows uneven coverage, with accountability absent from the cybersecurity and public-safety papers, transparency absent from cybersecurity, and minimal transparency attention in education, societal impact, legal, and public-safety papers. The authors conclude that ethical issues themselves hinder practical implementation of mitigation strategies, especially in healthcare and public governance, and that existing frameworks are not adaptable enough for evolving societal expectations and diverse contexts. The paper also flags in its validity section that, because most included studies are conceptual, author perspective bias and publication bias may skew the prominence of particular ethical dimensions.","pith_inferences":["If the 12.8 percent fully-evaluated rate generalizes beyond the 39 studies, then current AI-ethics guidelines should be read as hypotheses rather than validated best practices.","The five-dimension frame folds neighbouring values (fairness into bias, autonomy into accountability, consent into privacy); a finer-grained ontology might change the reported domain gaps, and the absence of inter-rater reliability means the counts should be treated as indicative.","A testable extension is to require an evaluation-status label in ethics papers that propose mitigation strategies, similar to preregistration in medical research, which would make the field's evidence base auditable.","The domain-gap analysis suggests concrete empirical priorities: study transparency and accountability in cybersecurity LLM use, and accountability in education and public safety, where the map currently shows near silence."],"forward_implications":["Adopting a mitigation strategy from this literature without independent validation is risky: fewer than one in eight studies fully evaluates what it proposes.","Domain-specific gaps are actionable: cybersecurity research on LLM ethics rarely discusses transparency or accountability, and education and public-safety research under-weights accountability.","High-stakes domains such as healthcare and public governance report the most severe implementation barriers, so pilots in those settings need the most careful empirical assessment.","Ethical frameworks should be treated as living documents with scheduled review cycles, not one-time checklists, because regulations and societal expectations change on multi-year cycles.","The evaluation rubric gives future reviewers and practitioners a shared way to classify whether an ethics strategy is a proposal or a validated fix."],"supporting_citations":[{"why":"supplies the six-step systematic review protocol the study follows","marker":"[12]"},{"why":"supplies the systematic mapping guidelines for classification and gap identification","marker":"[13]"},{"why":"the selected regulatory framework that grounds the safety and transparency dimensions","marker":"[38]"},{"why":"the selected design guideline that grounds the safety and bias dimensions","marker":"[46]"},{"why":"the selected risk-management framework that grounds the accountability and privacy dimensions","marker":"[47]"},{"why":"the selected industry standard that grounds the accountability dimension","marker":"[48]"},{"why":"the prior healthcare-focused systematic review whose domain scope this study extends","marker":"[14]"},{"why":"the earlier high-level mapping whose lack of evaluation-status analysis this study fills","marker":"[16]"},{"why":"the prior review of objective AI-ethics metrics that this study contrasts with human-centred frameworks","marker":"[53]"},{"why":"the snowballing guideline used to add the final two studies to the review","marker":"[54]"}],"fun_headline_variants":["Only 13% of AI ethics fixes get real-world tests","AI ethics: 39 studies, 5 with tested fixes","Mapping AI ethics: most safeguards never tested","Why AI ethics fixes fail: systematic map of 39 studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on the assumption that the five ethical dimensions chosen before coding—safety, privacy, transparency, bias, and accountability—are the correct and complete lens for labelling every study's concerns, and that the labelling was applied consistently across all 39 papers without measuring inter-rater reliability.","fun_headline_variants_meta":{"raw":{"variants":["Only 13% of AI ethics fixes get real-world tests","AI ethics: 39 studies, 5 with tested fixes","Mapping AI ethics: most safeguards never tested","Why AI ethics fixes fail: systematic map of 39 studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000279,"raw_usage":{"total_tokens":1666,"prompt_tokens":966,"completion_tokens":700,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":633}},"tokens_in":582,"tokens_out":700,"duration_ms":6272,"temperature":1.0,"reasoning_tokens":633,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:30:52.083323+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the same 39 papers with an independent team using the same five dimensions and report inter-rater agreement; or re-run the search in the same six databases and check whether the 39-study set is reproduced. If agreement is low, if a sixth dimension such as fairness or autonomy materially changes the domain counts, or if the snowballing step finds many additional eligible studies, then the reported gaps are properties of the coding choices rather than of the literature itself.","supporting_citations":[{"cited_title":"Kitchenham, L","cited_arxiv_id":null,"evidence_quote":"supplies the six-step systematic review protocol the study follows"},{"cited_title":"Madiega, Artificial intelligence act, European Parliament: European Parlia- mentary Research Service (2021)","cited_arxiv_id":null,"evidence_quote":"the selected regulatory framework that grounds the safety and transparency dimensions"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the selected design guideline that grounds the safety and bias dimensions"},{"cited_title":"AI, Artificial intelligence risk management framework (ai rmf 1.0) (2023)","cited_arxiv_id":null,"evidence_quote":"the selected risk-management framework that grounds the accountability and privacy dimensions"},{"cited_title":"pdf (2022)","cited_arxiv_id":null,"evidence_quote":"the selected industry standard that grounds the accountability dimension"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the prior healthcare-focused systematic review whose domain scope this study extends"},{"cited_title":"Atlam, M","cited_arxiv_id":null,"evidence_quote":"the earlier high-level mapping whose lack of evaluation-status analysis this study fills"},{"cited_title":"Palumbo, D","cited_arxiv_id":null,"evidence_quote":"the prior review of objective AI-ethics metrics that this study contrasts with human-centred frameworks"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the snowballing guideline used to add the final two studies to the review"}],"review_version":1}