Pith. sign in

REVIEW 3 cited by

Triggerless Backdoor Attack for NLP Tasks with Clean Labels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.07970 v2 pith:D2OIY4ZV submitted 2021-11-15 cs.CL cs.AIcs.CR

classification cs.CLcs.AIcs.CR
keywords labelstrategybackdoorattacksclean-labeleddetectedeasilypoisoned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Backdoor attacks pose a new threat to NLP models. A standard strategy to construct poisoned data in backdoor attacks is to insert triggers (e.g., rare words) into selected sentences and alter the original label to a target label. This strategy comes with a severe flaw of being easily detected from both the trigger and the label perspectives: the trigger injected, which is usually a rare word, leads to an abnormal natural language expression, and thus can be easily detected by a defense model; the changed target label leads the example to be mistakenly labeled and thus can be easily detected by manual inspections. To deal with this issue, in this paper, we propose a new strategy to perform textual backdoor attacks which do not require an external trigger, and the poisoned samples are correctly labeled. The core idea of the proposed strategy is to construct clean-labeled examples, whose labels are correct but can lead to test label changes when fused with the training set. To generate poisoned clean-labeled examples, we propose a sentence generation model based on the genetic algorithm to cater to the non-differentiable characteristic of text data. Extensive experiments demonstrate that the proposed attacking strategy is not only effective, but more importantly, hard to defend due to its triggerless and clean-labeled nature. Our work marks the first step towards developing triggerless attacking strategies in NLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PCAP-Backdoor: Backdoor Poisoning Generator for Network Traffic in CPS/IoT Environments

    cs.LG 2025-01 conditional novelty 6.0 of 10

    PCAP-Backdoor shows that an attacker who supplies only benign raw PCAP traffic can poison a deep learning IDS so that triggered attack traffic is classified as benign.

  2. When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations

    cs.CR 2024-11 conditional novelty 6.0 of 10

    Backdoored LLMs produce more diverse, less coherent explanations on triggered inputs, and this difference can be used to detect the backdoor.

  3. Neutralizing Backdoors through Information Conflicts for Large Language Models

    cs.CL 2024-11 conditional novelty 4.0 of 10

    A trigger-agnostic defense that merges a backdoored LLM with a clean-data LoRA model and adds contradictory prompt evidence, reducing attack success while keeping most clean-task accuracy.

Pith tools