Pith. sign in

REVIEW 4 major objections 5 minor 42 references

A new Persian-Islamic trust benchmark finds safety is the weakest dimension across most major LLMs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Useful first Persian trustworthiness benchmark with genuinely cultural prompts, but the headline numbers rest on an unvalidated judge pipeline; treat the rankings as provisional. the 4 major comments →

arxiv 2509.06838 v1 pith:RILI4D2Z submitted 2025-09-08 cs.CL cs.CR

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

classification cs.CL cs.CR
keywords large language modelstrustworthinessPersian languagecultural alignmentsafety evaluationbenchmarkethicsalignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces the EPT Benchmark, a 1,200-prompt Persian-Islamic evaluation set that tests large language models on truthfulness, safety, fairness, robustness, privacy, and ethics. Using a two-stage scoring procedure—ChatGPT similarity matching followed by native-expert majority vote—it ranks eight commercial and open models, reporting that Claude 3.7 Sonnet reaches an average 89.6% compliance while Qwen 3 falls to 70.4%. The paper's central claim is that current LLMs are markedly unsafe and culturally misaligned for Persian speakers: safety is the weakest dimension for most models, with several scoring below 57%. If true, this matters because Persian-speaking users of mainstream LLMs face disproportionate risks of harmful, biased, or culturally inappropriate output, and benchmark design itself needs to be culturally grounded.

Core claim

The central claim is that a culturally tailored Persian-Islamic benchmark can reveal trustworthiness gaps that English-centric evaluation misses. The authors curated 200 prompts per dimension (1,200 total), each with one expected answer reflecting Persian norms such as family privacy, religious sensitivity, and community equity, and scored model outputs as compliant or non-compliant. Their results show safety as the weakest dimension across most models—Qwen 3 at 48.8% and several others below 57.2%—while robustness is the strongest (93.0% for Gemini 2.5 Pro and GPT-4o). Claude 3.7 Sonnet is the most consistent overall (average 89.6%, SD 4.05), and Qwen 3 the weakest (70.4%, SD 11.46). The pa

What carries the argument

The EPT Benchmark itself: a labeled dataset of 1,200 Persian prompts (200 for each of six trustworthiness dimensions), evaluated by a binary compliance metric defined as the percentage of responses judged aligned with an expected answer. The scoring pipeline combines an automated first pass, in which ChatGPT measures similarity between expected and generated answers, with a second pass in which native Persian experts independently judge each response and a majority vote produces the final label. This two-stage design is what converts raw model outputs into the compliance percentages that drive all rankings.

Load-bearing premise

The benchmark's scores rest on the unvalidated assumption that ChatGPT's similarity judgments match human expert labels and that a single expected answer per prompt captures Persian-Islamic trustworthiness.

What would settle it

Take a random sample of EPT prompts, have a fresh panel of native Persian experts score model responses without seeing the expected answers, and compare their binary labels with the ChatGPT-based labels; low agreement would show the reported compliance rates are evaluation artifacts.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the benchmark measures what it claims, safety alignment for Persian users is the most urgent gap: several major models comply on safety in fewer than six of ten prompts.
  • Culturally grounded evaluation becomes a precondition for deploying LLMs in Persian-language applications such as education, healthcare, and public services.
  • Model selection for Persian users should weight safety and cultural-ethical alignment rather than general ability alone: the overall leader and the last-place model differ by nearly 20 percentage points.
  • The publicly released dataset gives developers a concrete target for fine-tuning: the prompts and expected answers can serve as a Persian safety and ethics training set.
  • The authors' planned extensions to multimodal inputs, adversarial scenarios, and domain-expert refinement would likely expose additional failure modes beyond text-only safety.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The published compliance rates are only as trustworthy as the scoring pipeline, and the paper reports no validation of the ChatGPT judge against human labels and no inter-annotator agreement; a human-only rescoring of a random subset could shift the rankings.
  • Each prompt has a single expected answer authored by the EPT team, so the scores encode one particular Persian-Islamic normative stance; a different expert panel might set different expectations and produce different dimension-level results.
  • The appendix's safety examples show models responding to prompts for a realistic murder plot with detailed methods; a stricter operationalization of safety, counting refusals and warnings rather than any compliant answer, would likely lower several models' safety scores further.
  • Comparing EPT-style scores across languages could test whether the safety deficits are a Persian-specific alignment problem or a general weakness revealed most clearly by culturally specific prompts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces the EPT (Evaluation of Persian Trustworthiness) benchmark, a culturally grounded Persian-language benchmark with 1,200 prompts across six dimensions: ethics, fairness, privacy, robustness, safety, and truthfulness. Eight LLMs (GPT-4o, Claude 3.7 Sonnet, Gemini 2.5 Pro, DeepSeek v3, Grok 3, Llama 3.3, Mistral 3, and Qwen 3) are evaluated using a binary compliance metric. The evaluation uses a two-stage pipeline in which ChatGPT first measures similarity between expected and model-generated responses, followed by a human expert majority vote. The paper reports that Claude 3.7 Sonnet achieves the highest mean compliance (89.6%) and Qwen 3 the lowest (70.4%), and concludes that safety is the weakest dimension across most models. The dataset is publicly released.

Significance. If the evaluation pipeline is valid, EPT would fill a real gap: there is no widely used Persian-Islamic trustworthiness benchmark covering these six dimensions. The paper's strengths include a publicly available dataset, 200 prompts per dimension, evaluation of eight proprietary and open-weight models, and raw appendix outputs that lend qualitative support to the safety concern. However, the central quantitative claims currently rest on an unvalidated labeling procedure. The manuscript must demonstrate label reliability and statistical robustness before the compliance rates can be accepted as measuring model trustworthiness. With those additions, the benchmark would be a useful contribution to culturally aware LLM evaluation.

major comments (4)
  1. [Section 4] The two-stage evaluation is underspecified. The paper states that ChatGPT 'measure[s] the similarity between expected and model-generated responses' and that 'qualified experts' then review and majority-vote, but it does not report the automated judge's prompt, model version, temperature, or number of runs, the number of experts, the annotation instructions, or how the two stages were merged into the published binary labels. No inter-annotator agreement or ChatGPT-versus-human agreement is reported. Since the headline numbers (Section 4.1: Claude 3.7 Sonnet 89.6%, Qwen 3 70.4%) are produced by this pipeline, the central claim requires evidence that the labels are stable. The Appendix items make this concrete: the ethics item expects a single 'bad' classification for abortion, the fairness item requires a particular answer about a politically charged sentence, and the robustness item trea
  2. [Section 4.2] The justification 'Inferential statistical tests were not conducted due to the categorical nature of the data' is incorrect: binary compliance data are exactly the kind of data for which chi-square tests, Fisher's exact tests, logistic regression, or bootstrap confidence intervals are appropriate. With n=200 per model-dimension, the reported percentages carry substantial sampling uncertainty. For example, a 92.0% rate has an approximate 95% CI of roughly ±4 percentage points; differences of a few points between models may not be significant. Report confidence intervals and at least pairwise comparisons (with multiple-testing correction) or a mixed-effects model with prompt and model as factors.
  3. [Section 4.1 and Appendix (Safety)] The qualitative safety finding is partially supported by the Appendix, where multiple models provide detailed methods of murder. However, the numerical rates are not auditable without item-level labels. In particular, the Appendix safety example includes Claude 3.7 Sonnet giving a concrete plot (slow poisoning, tampering with brakes); Claude is reported to have 92.0% Safety. If that response was counted as compliant, the safety metric is inconsistent; if non-compliant, a single item is at stake. The paper should release item-level labels for all 1,200 items, at least in the supplement, and clarify how the displayed Appendix responses were scored.
  4. [Section 4 (evaluator bias)] Using ChatGPT as the automated judge while evaluating GPT-4o is an evaluator-bias risk. If the judge and GPT-4o share alignment behavior, the comparison can systematically favor or disfavor OpenAI models. The human stage does not remove this risk unless the automated stage is shown to agree with human labels. Please either use an independent judge (e.g., a different model family or a second human panel) or report per-model ChatGPT–human agreement and disagreement patterns.
minor comments (5)
  1. [Title and Section 2.3] Typographical spacing: 'T rustworthiness' appears in the title and body headings; also 'F airness' and 'T ruthfulness' in Section 2.3 headings.
  2. [Section 4.1] The sentence 'GPT-4o, Grok 3, Llama 3.3, and Mistral 3 all scoring below 57.2%' is imprecise because Grok 3 is exactly 57.2; use 'at or below'.
  3. [Appendix] The appendix refers to the 'Honesty' category, while the rest of the paper uses 'Truthfulness'; unify terminology.
  4. [References] Reference [33] (SafetyBench) lists authors 'Zhuo Zheng, Ronghui Zhang, ...'; this appears to be an incorrect author set for arXiv:2305.13656. Please verify and correct all citations.
  5. [Section 4] The paper says 'The dataset was created by a team of experts...' without specifying the number of experts and their relevant credentials; this detail should be added for transparency.

Circularity Check

0 steps flagged

No circularity: EPT is an empirical benchmark; its compliance rates are measured outcomes, not outputs recycled from fitted inputs or self-citation.

full rationale

The paper's central claim is that the EPT benchmark measures Persian-Islamic trustworthiness and that reported compliance rates reflect model differences. This is an empirical measurement claim, not a derivation. The compliance metric is explicitly defined as the number of compliant responses divided by the total number of responses (Section 4), and the per-dimension percentages are computed from model outputs judged against pre-authored expected answers. Those expected answers are the benchmark's input labels, not quantities fitted from the same model outputs and then renamed as predictions. No equation in the paper constructs the reported rates from a parameter that was previously fitted to those rates, and no result is obtained by substituting the conclusion into the premise. The two-stage evaluation pipeline (ChatGPT similarity followed by expert majority vote) is under-specified and is a genuine validation risk: agreement between the automated judge and human experts, the number and protocol of experts, and the merging rule are not reported. However, an unvalidated measurement pipeline is a correctness or reliability concern, not circularity: the model responses being scored are independent artifacts, and the benchmark labels are not derived from the scoring rule. The appendix does show that some safety responses contain detailed harmful content, which independently supports the qualitative safety finding; the numeric rankings remain an empirical claim that could be wrong if the labeling is unstable, but being falsifiable in that way is the opposite of being circular. The paper contains no load-bearing self-citations, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via citation. Consequently, under the hard rule that circularity must be demonstrated by a specific reduction, no circular step can be identified.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The central claim rests on three pillars: the cultural ground truth is correctly encoded in the expected answers, the ChatGPT judge is reliable, and the prompts are representative. None of these is independently verified in the paper, but none is a fitted parameter or an invented physical entity.

axioms (4)
  • domain assumption A single Persian-Islamic normative stance can serve as the ground truth for trustworthiness in each dimension.
    Expected answers are authored by a team of experts, and the examples in Section 2.3 and the Appendix (riba, abortion, Persian Gulf naming) treat one cultural and religious position as correct. If this stance is disputed or unrepresentative, compliance scores measure conformity rather than trustworthiness.
  • domain assumption ChatGPT's similarity-based binary judgments are a valid approximation of human expert judgments.
    Section 4 states that ChatGPT measures similarity between expected and model responses, but no validation, agreement statistics, or error analysis is reported.
  • domain assumption A binary compliant/non-compliant label captures the relevant behavior in each dimension.
    Section 4 defines compliance as the ratio of correct aligned responses to total responses. It is unclear how partial answers, equivocation, or nuance are classified, as in DeepSeek's long abortion answer in the Appendix.
  • domain assumption 200 prompts per dimension provide a representative sample of that dimension.
    Section 4 states that 1,200 prompts were crafted by experts, but no sampling procedure, prompt diversity metrics, or coverage analysis is provided.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models." pith.science (2026). https://pith.science/paper/RILI4D2Z

@misc{pith2026250906838,
  author       = {Pith},
  title        = {Pith review of: EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RILI4D2Z}},
  note         = {Machine review of arXiv:2509.06838}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs), trained on extensive datasets using advanced deep learning architectures, have demonstrated remarkable performance across a wide range of language tasks, becoming a cornerstone of modern AI technologies. However, ensuring their trustworthiness remains a critical challenge, as reliability is essential not only for accurate performance but also for upholding ethical, cultural, and social values. Careful alignment of training data and culturally grounded evaluation criteria are vital for developing responsible AI systems. In this study, we introduce the EPT (Evaluation of Persian Trustworthiness) metric, a culturally informed benchmark specifically designed to assess the trustworthiness of LLMs across six key aspects: truthfulness, safety, fairness, robustness, privacy, and ethical alignment. We curated a labeled dataset and evaluated the performance of several leading models - including ChatGPT, Claude, DeepSeek, Gemini, Grok, LLaMA, Mistral, and Qwen - using both automated LLM-based and human assessments. Our results reveal significant deficiencies in the safety dimension, underscoring the urgent need for focused attention on this critical aspect of model behavior. Furthermore, our findings offer valuable insights into the alignment of these models with Persian ethical-cultural values and highlight critical gaps and opportunities for advancing trustworthy and culturally responsible AI. The dataset is publicly available at: https://github.com/Rezamirbagheri110/EPT-Benchmark.

Figures

Figures reproduced from arXiv: 2509.06838 by Ali Javeri, Amir Mahdi Sadeghzadeh, Mohammad Mahdi Mirkamali, Mohammad Reza Mirbagheri, Rasool Jalili, Zahra Motoshaker Arani.

Figure 1
Figure 1. Figure 1: EPT Benchmark diagram sensitive content handling, or public interaction. This multifaceted concept encompasses ethics, fairness, privacy, robustness, safety, and truthfulness, each evaluated through standardized benchmarks, expert human evaluation, and interpretability mechanisms to ensure consistent performance across adversarial, ambiguous, and real-world scenarios [11, 12]. The inherent complexity of LL… view at source ↗
Figure 2
Figure 2. Figure 2: Bar plot of compliance rates across six aspects for eight LLMs. Safety is the weakest dimen￾sion (e.g., Qwen 3: 48.8%), while Robustness is the strongest (e.g., Gemini 2.5 Pro: 93.0%). This visual￾ization highlights performance gaps and strengths. Ethics Fairness Privacy Robustness Safety Truthfulness Category Claude DeepSeek GPT-4o Gemini Grok LLaMA Mistral Qwen Model 82.5 93.0 87.1 91.5 92.0 91.5 91.0 85… view at source ↗
Figure 3
Figure 3. Figure 3: Heatmap of compliance rates, with color intensity indicating performance disparities. Low compliance is evident in Qwen 3’s Safety (48.8%), while high compliance is seen in Claude 3.7 Sonnet’s Fairness (93.0%). This figure identifies trends and outliers. These findings highlight a critical need for targeted improvements in model trustworthiness, with par￾ticular emphasis on addressing severe weaknesses in … view at source ↗
Figure 4
Figure 4. Figure 4: Violin plot showing compliance rate distri [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Combined radar chart of compliance pro￾files for all LLMs. It underscores Safety deficits (e.g., Qwen 3: 48.8%) and Robustness strengths (e.g., Gem￾ini 2.5 Pro: 93.0%), enabling rapid comparison of model performance. as text, images, and audio, which are increasingly prevalent in real-world applications. This will enable comprehensive assessments across varied linguistic, cultural, and contextual aspects, … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 22 canonical work pages · 3 internal anchors

  1. [1]

    Emergent abilities in large language mod- els: A survey.arXiv preprint arXiv:2503.05788, Feb 2025

    Leonardo Berti, Flavio Giorgi, and Gjergji Kas- neci. Emergent abilities in large language mod- els: A survey.arXiv preprint arXiv:2503.05788, Feb 2025

  2. [2]

    Multimodal large language models: A survey.arXiv preprint arXiv:2506.10016, May 2025

    Longzhen Han, Awes Mubarak, Almas Baimagambetov, Nikolaos Polatidis, and Thar Baker. Multimodal large language models: A survey.arXiv preprint arXiv:2506.10016, May 2025

  3. [3]

    Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Ka- terina Sedova

    Josh A. Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Ka- terina Sedova. Generative language models and automated influence operations: Emerg- ing threats and potential mitigations.CoRR, abs/2301.04246, 2023

  4. [4]

    Training language mod- els to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Sandipan Slama, Alex Ray, et al. Training language mod- els to follow instructions with human feedback. Advances in Neural Information Processing Sys- tems, 35:27730–27744, 2022

  5. [5]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xi- aoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. A survey on evaluation of large language models.CoRR, abs/2307.03109, 2023

  6. [6]

    Bowman, Zac Hatfield- Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernan- dez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Lan- dau, Kamal Ndouss...

  7. [7]

    A com- prehensive survey of llm alignment techniques: Rlhf, rlaif, ppo, dpo and more.arXiv preprint arXiv:2407.16216, jul 2024

    Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Zixu (James) Zhu, Xiang-Bo Mao, Sitaram Asur, and Na (Claire) Cheng. A com- prehensive survey of llm alignment techniques: Rlhf, rlaif, ppo, dpo and more.arXiv preprint arXiv:2407.16216, jul 2024

  8. [8]

    Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan

    Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. A general lan- guage assistant a...

  9. [9]

    AI alignment: A comprehensive survey.CoRR, abs/2310.19852, 2023

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xue- hai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao. AI alignment: A comprehensive survey.CoRR...

  10. [10]

    Large language model alignment: A survey.CoRR, abs/2309.15025, 2023

    TianhaoShen, RenrenJin, YufeiHuang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. Large language model alignment: A survey.CoRR, abs/2309.15025, 2023

  11. [11]

    Trustworthy llms: a survey and guideline for evaluating large language models’ alignment

    Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xi- aoyingZhang, RuochengGuo, HaoCheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy llms: a survey and guideline for evaluating large language models’ alignment. CoRR, abs/2308.05374, 2023

  12. [12]

    In chatgpt we trust? measur- ing and characterizing the reliability of chatgpt

    Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. In chatgpt we trust? measur- ing and characterizing the reliability of chatgpt. CoRR, abs/2304.08979, 2023

  13. [13]

    Edward Y. Chang. Modeling emotions and ethics with large language models.arXiv preprint arXiv:2404.13071, Apr 2024

  14. [14]

    Deconstructing The Ethics of Large Language Models from Long-standing Issues to New-emerging Dilemmas: A Survey

    Chengyuan Deng, Yiqun Duan, Xin Jin, Heng Chang, and et al. Deconstructing the ethics of large language models from long-standing issues to new-emerging dilemmas: A survey.arXiv preprint arXiv:2406.05392, Jun 2024

  15. [15]

    Navigating LLM ethics: Ad- vancements, challenges, and future directions

    Junfeng Jiao, Saleh Afroogh, Yiming Xu, and Connor Phillips. Navigating LLM ethics: Ad- vancements, challenges, and future directions. CoRR, abs/2406.18841, 2024

  16. [16]

    On the trustworthiness landscape of state-of-the-art generative models: A survey and outlook.International Journal of Computer Vision, Feb 2025

    Mingyuan Fan, Chengyu Wang, Cen Chen, Yang Liu, and Jun Huang. On the trustworthiness landscape of state-of-the-art generative models: A survey and outlook.International Journal of Computer Vision, Feb 2025. Received: 1 April 2024 / Accepted: 4 February 2025 / Published online: 28 February 2025

  17. [17]

    Gallegos, Ryan A

    Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md. Mehrab Tanjim, Sungchul Kim, Franck Der- noncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. Bias and fairness in large language mod- els: A survey.Comput. Linguistics, 50(3):1097– 1179, 2024

  18. [18]

    A survey on privacy risks and protection in large language models.arXiv preprint arXiv:2505.01976, May 2025

    Kang Chen, Xiuze Zhou, Yuanguo Lin, Shibo Feng, Li Shen, and Pengcheng Wu. A survey on privacy risks and protection in large language models.arXiv preprint arXiv:2505.01976, May 2025

  19. [19]

    Privacy issues in large language models: A survey.Com- put

    Hareem Kibriya, Wazir Zada Khan, Ayesha Sid- diqa, and Muhammad Khurrum Khan. Privacy issues in large language models: A survey.Com- put. Electr. Eng., 120:109698, 2024

  20. [20]

    Trustllm: Trustworthiness in large language models

    Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, et al. Trustllm: Trustworthiness in large language models. InInternational Conference on Machine Learning, pages 20166–20270. PMLR, 2024

  21. [21]

    Assessing hidden risks of llms: An empirical study on robustness, consistency, and credibility.arXiv preprint, May 2023

    Wentao Ye, Mingfeng Ou, Tianyi Li, Xuetao Ma, Yifan Yanggong, Sai Wu, Jie Fu, Gang Chen, Junbo Zhao, et al. Assessing hidden risks of llms: An empirical study on robustness, consistency, and credibility.arXiv preprint, May 2023

  22. [22]

    Zico Kolter

    Eric Wong, Leslie Rice, and J. Zico Kolter. Im- proving robustness and generalization of deep learning models via adversarial training. InIn- ternational Conference on Learning Representa- tions (ICLR), 2021

  23. [23]

    Bo Liu, Li-Ming Zhan, Zexin Lu, Yujie Feng, Lei Xue, and Xiao-Ming Wu. How good are llms at out-of-distribution detection? In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors,Proceedings of the 2024 Joint International Conference on Computa- tional Linguistics, Language Resources and Eval- uatio...

  24. [24]

    Cold-attack: Jailbreaking llms with stealthiness and controllability

    Xingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu. Cold-attack: Jailbreaking llms with stealthiness and controllability. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024

  25. [25]

    Jail- break attacks and defenses against large lan- guage models: A survey.CoRR, abs/2407.04295, 2024

    Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li. Jail- break attacks and defenses against large lan- guage models: A survey.CoRR, abs/2407.04295, 2024

  26. [26]

    Factuality challenges in the era of large language models.arXiv preprint, 2023

    Isabelle Augenstein, Timothy Baldwin, Meey- oung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, et al. Factuality challenges in the era of large language models.arXiv preprint, 2023

  27. [27]

    LaToza, Kevin Moran, and Wing Lam

    Sajed Jalil, Suzzana Rafi, Thomas D. LaToza, Kevin Moran, and Wing Lam. ChatGPT and software testing education: Promises & perils. In2023 IEEE International Conference on Soft- ware Testing, Verification and Validation Work- shops (ICSTW), pages 4130–4137, Dublin, Ire- land, April 2023. IEEE

  28. [28]

    Survey of hallucination in natural language generation

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Comput. Surv., 55(12):248:1–248:38, 2023

  29. [29]

    Halueval: A large-scale hallucination evaluation benchmark for large lan- guage models

    Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. Halueval: A large-scale hallucination evaluation benchmark for large lan- guage models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 6449–6464. ...

  30. [30]

    Yujia Zhou, Yan Liu, Xiaoxi Li, Jiajie Jin, Hongjin Qian, Zheng Liu, Chaozhuo Li, Zhicheng Dou, Tsung-Yi Ho, and Philip S. Yu. Trustworthiness in retrieval -augmented generation systems: A survey.arXiv preprint arXiv:2409.10102, Sep 2024

  31. [31]

    Truthfulqa: Measuring how models mimic hu- man falsehoods

    Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic hu- man falsehoods. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors,Proceed- ings of the 60th Annual Meeting of the Associa- tion for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 3214–3252. Assoc...

  32. [32]

    Trustgpt: A bench- mark for trustworthy and responsible large lan- guage models.arXiv preprint arXiv:2306.11507, 2023

    Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Tianyu Li, and Danqi Chen. Trustgpt: A bench- mark for trustworthy and responsible large lan- guage models.arXiv preprint arXiv:2306.11507, 2023

  33. [33]

    Link Prediction without Graph Neural Networks

    Zhuo Zheng, Ronghui Zhang, Junxiao Zhang, Wei Zhao, Yuxiang You, Yuxin Bian, Jun Li, Chuanyang Zhou, Xiao Zhang, Jie Yang, et al. Safetybench: Evaluating the safety of large lan- guage models with multiple choice questions. arXiv preprint arXiv:2305.13656, 2023

  34. [34]

    Llm ethics benchmark: A three- dimensional assessment system for evaluating moral reasoning in large language models.arXiv preprint, May 2025

    Junfeng Jiao, Saleh Afroogh, Abhejay Murali, Kevin Chen, David Atkinson, and Amit Dhu- randhar. Llm ethics benchmark: A three- dimensional assessment system for evaluating moral reasoning in large language models.arXiv preprint, May 2025. Available at arXiv

  35. [35]

    Cvalues: Measuring the values of chinese large language models from safety to responsibility.arXiv preprint arXiv:2307.09705, 2023

    Xianjun Xu, Xiao Zhang, Tianlin Zhao, Yufei Zhang, Rui Yan, Wenyu Yang, Linda Xu, and Haifeng Jiang. Cvalues: Measuring the values of chinese large language models from safety to responsibility.arXiv preprint arXiv:2307.09705, 2023

  36. [36]

    Safety assessment of chinese large language models.arXiv preprint arXiv:2304.10436, 2023

    Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. Safety assessment of chinese large language models.arXiv preprint arXiv:2304.10436, 2023

  37. [37]

    Percy Liang, Rishi Bommasani, Tony Lee, Dim- itris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cos- grove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew A. Hudson, Eric Ze- likman, Esin Durmus, Faisal Ladhak, Frieda Rong, H...

  38. [38]

    ROBBIE: Ro- bust bias evaluation of large generative language models

    David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Michael Smith. ROBBIE: Ro- bust bias evaluation of large generative language models. InThe 2023 Conference on Empirical Methods in Natural Language Processing, 2023

  39. [39]

    Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. Decodingtrust: A comprehensive assessment of trustworthiness inGPTmodels. InThirty-seventh Conference on N...

  40. [40]

    Flames: Bench- marking value alignment of chinese large lan- guage models, 2023

    Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, YingchunWang, andDahuaLin. Flames: Bench- marking value alignment of chinese large lan- guage models, 2023

  41. [41]

    CLEVA: Chinese language models EVAl- uation platform

    Yanyang Li, Jianqiao Zhao, Duo Zheng, Zi-Yuan Hu, Zhi Chen, Xiaohui Su, Yongfeng Huang, Shi- jia Huang, Dahua Lin, Michael Lyu, and Liwei Wang. CLEVA: Chinese language models EVAl- uation platform. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Singapore, December 2023. Association for Com-...

  42. [42]

    Parameswaran

    Reya Vir, Shreya Shankar, Harrison Chase, Will Fu-Hinthorn, and Aditya G. Parameswaran. Promptevals: A dataset of assertions and guardrails for custom production large lan- guage model pipelines.arXiv preprint arXiv:2504.14738, 2025. Appendix In this section, we analyze the performance of large language models across six key dimensions of trustworthiness ...

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.