Pith. sign in

REVIEW 27 cited by

Label-Consistent Backdoor Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02771 v2 pith:IZE2KCJ3 submitted 2019-12-05 stat.ML cs.CRcs.LG

Label-Consistent Backdoor Attacks

classification stat.ML cs.CRcs.LG
keywords backdoorattacksinputsinjectingmodeladversarylabel-consistentrely
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Deep neural networks have been demonstrated to be vulnerable to backdoor attacks. Specifically, by injecting a small number of maliciously constructed inputs into the training set, an adversary is able to plant a backdoor into the trained model. This backdoor can then be activated during inference by a backdoor trigger to fully control the model's behavior. While such attacks are very effective, they crucially rely on the adversary injecting arbitrary inputs that are---often blatantly---mislabeled. Such samples would raise suspicion upon human inspection, potentially revealing the attack. Thus, for backdoor attacks to remain undetected, it is crucial that they maintain label-consistency---the condition that injected inputs are consistent with their labels. In this work, we leverage adversarial perturbations and generative models to execute efficient, yet label-consistent, backdoor attacks. Our approach is based on injecting inputs that appear plausible, yet are hard to classify, hence causing the model to rely on the (easier-to-learn) backdoor trigger.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks

    cs.LG 2026-05 unverdicted novelty 8.0

    In the proportional high-dimensional regime, stronger backdoor training triggers improve clean accuracy and make attack success non-monotonic for regularized GLMs on Gaussian mixtures, with closed-form proofs for squa...

  2. Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction

    cs.CR 2026-04 conditional novelty 7.5

    Scene-conditioned spatial-misdirection and duration-inflation backdoors succeed at 2.5–10% poison ratios on multimodal scanpath predictors and resist five adapted defenses.

  3. Lilith: Backdoor Generalization under Training-Inference Trigger Shift

    cs.CR 2026-07 conditional novelty 7.0

    A single poisoned training trigger can create a backdoor that fires for a whole family of unseen inference-time triggers, provided the variants preserve the anchor's representation geometry.

  4. Fast and Lightweight Backdoor Detection via Head Random Probing

    cs.CR 2026-05 unverdicted novelty 7.0

    HTell detects backdoors by random probing of the model head, reporting 99.03% true positive rate and 2.11% false positive rate at 12.69 ms per model on a benchmark of over 6700 models.

  5. Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction

    cs.CR 2026-04 conditional novelty 7.0

    Backdoor attacks on VLM-based scanpath predictors can redirect fixations toward chosen objects or inflate durations using input-conditioned triggers that evade cluster detection, and no tested defense blocks them with...

  6. Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning

    cs.CR 2026-03 unverdicted novelty 7.0

    SABLE shows that semantics-aware natural triggers enable effective backdoor attacks in federated learning against multiple aggregation rules while preserving benign accuracy.

  7. Inevitable Encounters: Backdoor Attacks Involving Lossy Compression

    cs.CR 2026-03 unverdicted novelty 7.0

    ROI coding enables backdoor triggers to survive lossy compression by embedding malicious information into binary bitstreams via sample-specific or customized masks for both learned and traditional codecs.

  8. From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense

    cs.LG 2026-01 conditional novelty 7.0

    PRISM uses a frozen VLM as an evolving semantic gatekeeper, reporting average attack success below 1% on CIFAR-10 while preserving clean accuracy.

  9. Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor

    cs.LG 2026-07 conditional novelty 6.5

    Learning a backdoored reference model as a poisonous-sample oracle enables near-perfect training-time backdoor removal with negligible natural-accuracy loss.

  10. Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling

    cs.LG 2026-08 conditional novelty 6.0

    Backdoors can be trapped by a decoy shortcut branch trained alongside the main network and removed simply by discarding that branch, no extra data needed.

  11. Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs

    cs.CR 2026-07 conditional novelty 6.0

    Retiming target-class neuromorphic events installs clean-label SNN backdoors with ASR up to 1.0 while leaving rate frames identical, and rate-collapsed defenses miss them.

  12. Benign on Label, Malicious by Design: Clean-Label Dormant-to-Activated Backdoor via Machine Unlearning with Removable Camouflage

    cs.CR 2026-07 conditional novelty 6.0

    Dual generators jointly learn persistent clean-label triggers and removable camouflage so a backdoor stays dormant until attacker-submitted camouflage samples are unlearned.

  13. Anti-Backdoor Coreset Selection via Cumulative Entropy

    cs.LG 2026-07 conditional novelty 6.0

    ABCS selects a high-entropy coreset from poisoned training data and trains on it, mitigating eight backdoor attacks with minimal natural-accuracy loss.

  14. How Context Attribution Handles What the Model Already Knows

    cs.CL 2026-07 conditional novelty 6.0

    Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.

  15. Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

    cs.CV 2026-06 unverdicted novelty 6.0

    Mirage achieves 73% misclassification success on LiDAR 3DOD models with 0.5% poisoning rate via label-consistent trigger injection.

  16. Density-aware Sample-specific Attack

    cs.LG 2026-05 unverdicted novelty 6.0

    A density-aware sample-specific backdoor attack steers triggers into low-density regions via bilevel optimization to achieve high post-defense success rates on image datasets.

  17. Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget

    cs.CR 2026-05 unverdicted novelty 6.0

    Checkerboard derives a closed-form checkerboard trigger for clean-label backdoor attacks that achieves over 94% ASR with poisoning rates as low as 0.46% on ImageNet-100 and 99.99% ASR with 20 samples on CIFAR-10.

  18. CSC: Turning the Adversary's Poison against Itself

    cs.CR 2026-04 unverdicted novelty 6.0

    CSC identifies backdoored samples via early-epoch latent clustering and conceals them by relabeling to a virtual class, driving attack success rates near zero on benchmarks with little clean accuracy loss.

  19. TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

    cs.CV 2025-09 conditional novelty 6.0

    TokenSwap poisons LVLMs so that triggered images produce captions with subject and object roles reversed, achieving high attack success while evading a perplexity-based detector.

  20. Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

    cs.CR 2026-07 conditional novelty 5.5

    World-model-based embodied AI creates a predictive security boundary where attacks on data, sensors, imagination, ranking, and feedback can turn into unsafe physical action and false safety certificates.

  21. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  22. Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

    cs.CR 2026-07 unverdicted novelty 5.0

    Pmeta-TLA combines a frame-level timbre leakage trigger with meta-learning and PCGrad to inject multiple backdoors into speech models in one training run, claiming better attack success, stealth, and lower cost than b...

  23. TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models

    cs.CR 2026-06 unverdicted novelty 5.0

    TEMPO-Diffusion is a targeted backdoor attack framework for diffusion models that uses time-conditioned triggers to poison class-specific synthetic data, achieving high attack success in downstream classifiers.

  24. Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks

    cs.CR 2026-06 unverdicted novelty 5.0

    Trigger color significantly affects semantic backdoor attack success in federated learning on CelebA hair-color classification, with white triggers better for blond targets and black for black targets.

  25. Dataset Poisoning Attacks on Behavioral Cloning Policies

    cs.LG 2025-11 conditional novelty 5.0

    A few doctored demonstrations with a small red patch give attackers near-complete hidden control over behavior-cloning policies without lowering the policy's ordinary task reward.

  26. Cryptographic Backdoor for Neural Networks: Boon and Bane

    cs.CR 2025-09 conditional novelty 5.0

    A signature-verification circuit attached to a neural network yields an undetectable backdoor attack and three privacy defenses: watermarking, user authentication, and IP-leak tracking.

  27. Certification of Machine Learning Models via Directional Sharpness

    cs.LG 2026-06 unverdicted novelty 4.0

    Directional sharpness is introduced as a metric that correlates more strongly with generalization, identifies poor generalization more reliably, and supports efficient auditing and zero-knowledge certification.