Pith. sign in

REVIEW 8 cited by

Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.02128 v1 pith:BPXK55MD submitted 2022-09-05 cs.CL

classification cs.CL
keywords adversarialattacksmodelsplmspre-traineddevelopmentfine-tuninglanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in the development of large language models have resulted in public access to state-of-the-art pre-trained language models (PLMs), including Generative Pre-trained Transformer 3 (GPT-3) and Bidirectional Encoder Representations from Transformers (BERT). However, evaluations of PLMs, in practice, have shown their susceptibility to adversarial attacks during the training and fine-tuning stages of development. Such attacks can result in erroneous outputs, model-generated hate speech, and the exposure of users' sensitive information. While existing research has focused on adversarial attacks during either the training or the fine-tuning of PLMs, there is a deficit of information on attacks made between these two development phases. In this work, we highlight a major security vulnerability in the public release of GPT-3 and further investigate this vulnerability in other state-of-the-art PLMs. We restrict our work to pre-trained models that have not undergone fine-tuning. Further, we underscore token distance-minimized perturbations as an effective adversarial approach, bypassing both supervised and unsupervised quality measures. Following this approach, we observe a significant decrease in text classification quality when evaluating for semantic similarity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents

    cs.CR 2026-07 conditional novelty 7.0 of 10

    A trained attack model generates single emails that silently inject false memories into persistent AI agents, achieving 87.5% end-to-end success on GPT-5.4 and transferring across architectures and memory backends.

  2. Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels

    cs.CR 2025-10 conditional novelty 6.0 of 10

    Indirect prompt injection through ads, webviews, and notifications reliably diverts mobile LLM agents into leaking data and installing malware across eight evaluated agents.

  3. When Compression Becomes an Attack Surface: Black-Box Attacks on Prompt-Compressed LLM Agents

    cs.CR 2025-10 reject novelty 6.0 of 10

    The paper claims prompt compression is a new attack surface, but the abstract's COMA attack never appears in the body and the body's SoftCom requires white-box access.

  4. Understanding the Ability of LLMs to Handle Character-Level Perturbation

    cs.CL 2025-10 conditional novelty 6.0 of 10

    LLMs remain surprisingly accurate on math and coding when invisible Unicode noise is inserted after every character, with robustness driven by implicit internal denoising and, for some models, explicit rewriting in ch...

  5. When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A systematic evaluation shows that masking the interlocutor's persona lowers target speaker identification accuracy, and that zero-shot models often copy biography details, making identification easier but dialogues m...

  6. We Can't Understand AI Using our Existing Vocabulary

    cs.CL 2025-02 conditional novelty 4.0 of 10

    AI interpretability is reframed as building a shared human-machine language in which each new word is a learned token embedding trained by preference optimization.

  7. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

  8. Prompt Injection 2.0: Hybrid AI Threats

    cs.CR 2025-07 reject novelty 2.0 of 10

    A structured taxonomy of hybrid prompt injection attacks shows how XSS, CSRF, and SQL injection vectors converge with LLM manipulation to bypass traditional controls.

Pith tools