Pith. sign in

REVIEW 7 cited by

Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12867 v4 pith:L6PL36PL submitted 2023-01-30 cs.CL cs.SE

classification cs.CLcs.SE
keywords llmschatgptethicaltextitfuturepracticaltoxicityapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent breakthroughs in natural language processing (NLP) have permitted the synthesis and comprehension of coherent text in an open-ended way, therefore translating the theoretical algorithms into practical applications. The large language models (LLMs) have significantly impacted businesses such as report summarization software and copywriters. Observations indicate, however, that LLMs may exhibit social prejudice and toxicity, posing ethical and societal dangers of consequences resulting from irresponsibility. Large-scale benchmarks for accountable LLMs should consequently be developed. Although several empirical investigations reveal the existence of a few ethical difficulties in advanced LLMs, there is little systematic examination and user study of the risks and harmful behaviors of current LLM usage. To further educate future efforts on constructing ethical LLMs responsibly, we perform a qualitative research method called ``red teaming'' on OpenAI's ChatGPT\footnote{In this paper, ChatGPT refers to the version released on Dec 15th.} to better understand the practical features of ethical dangers in recent LLMs. We analyze ChatGPT comprehensively from four perspectives: 1) \textit{Bias} 2) \textit{Reliability} 3) \textit{Robustness} 4) \textit{Toxicity}. In accordance with our stated viewpoints, we empirically benchmark ChatGPT on multiple sample datasets. We find that a significant number of ethical risks cannot be addressed by existing benchmarks, and hence illustrate them via additional case studies. In addition, we examine the implications of our findings on AI ethics and harmal behaviors of ChatGPT, as well as future problems and practical design considerations for responsible LLMs. We believe that our findings may give light on future efforts to determine and mitigate the ethical hazards posed by machines in LLM applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Biased or Personalized? The Impact of Personal Information on AI-driven Development

    cs.SE 2026-07 conditional novelty 6.0 of 10

    Changing only the prompter's age and gender in AI coding prompts produces statistically significant differences in generated website interface design, template content, and code structure across 800 generated websites...

  2. LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A systematic review that categorizes LLM threats, severity scores, and mitigations across development and operation life cycles and multiple deployment scenarios.

  3. MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MAGPIE is a 158-scenario benchmark showing large language model agents misclassify and leak contextually private information in multi-agent collaboration, even under explicit privacy instructions.

  4. Sword and Shield: Uses and Strategies of LLMs in Navigating Disinformation

    cs.HC 2025-06 conditional novelty 6.0 of 10

    In a 25-participant Werewolf-style game, all roles used an LLM chatbot strategically, as a sword for disinformation and a shield against it.

  5. Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A name-substituted variant of the BBQ benchmark shows that LLMs retain nationality stereotypes even when explicit labels are removed, with smaller models showing more bias and lower accuracy.

  6. The Role of Generative AI in Facilitating Social Interactions: A Scoping Review

    cs.HC 2025-06 conditional novelty 5.0 of 10

    A scoping review of 30 studies finds GAI social-interaction tools are mostly GPT-based text chatbots targeting vulnerable users, with limited participatory design and mostly early-stage evaluations.

  7. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools