Pith. sign in

REVIEW 8 cited by

DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11698 v5 pith:QLMKB2HH submitted 2023-06-20 cs.CL cs.AIcs.CR

classification cs.CLcs.AIcs.CR
keywords modelstrustworthinessgpt-4comprehensivedecodingtrusthttpsrobustnesswork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike. Yet, while the literature on the trustworthiness of GPT models remains limited, practitioners have proposed employing capable GPT models for sensitive applications such as healthcare and finance -- where mistakes can be costly. To this end, this work proposes a comprehensive trustworthiness evaluation for large language models with a focus on GPT-4 and GPT-3.5, considering diverse perspectives -- including toxicity, stereotype bias, adversarial robustness, out-of-distribution robustness, robustness on adversarial demonstrations, privacy, machine ethics, and fairness. Based on our evaluations, we discover previously unpublished vulnerabilities to trustworthiness threats. For instance, we find that GPT models can be easily misled to generate toxic and biased outputs and leak private information in both training data and conversation history. We also find that although GPT-4 is usually more trustworthy than GPT-3.5 on standard benchmarks, GPT-4 is more vulnerable given jailbreaking system or user prompts, potentially because GPT-4 follows (misleading) instructions more precisely. Our work illustrates a comprehensive trustworthiness evaluation of GPT models and sheds light on the trustworthiness gaps. Our benchmark is publicly available at https://decodingtrust.github.io/ ; our dataset can be previewed at https://huggingface.co/datasets/AI-Secure/DecodingTrust ; a concise version of this work is at https://openreview.net/pdf?id=kaHpo8OZw2 .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  2. RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    RLearner-LLM's Hybrid-DPO fuses DeBERTa NLI and LLM verifier scores to deliver up to 6x higher NLI entailment than standard SFT while preserving answer coverage across academic domains.

  3. MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MAGPIE is a 158-scenario benchmark showing large language model agents misclassify and leak contextually private information in multi-agent collaboration, even under explicit privacy instructions.

  4. Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)

    cs.CY 2025-05 reject novelty 6.0 of 10

    Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.

  5. Online Aggregation of Trajectory Predictors

    cs.RO 2025-02 conditional novelty 5.0 of 10

    An online learning rule, based on SQUINT, mixes multiple trajectory predictors and tracks the best expert under distribution shift.

  6. Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

    cs.AI 2026-01 conditional novelty 4.0 of 10

    Small open-weight LLMs endorse prohibited actions 24% of the time under affirmative framing but 77-100% under negated framings, a polarity swing that threatens high-stakes AI deployment.

  7. The Science of Evaluating Foundation Models

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.

  8. Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.

Pith tools