REVIEW 28 cited by
Label-Consistent Backdoor Attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Label-Consistent Backdoor Attacks
read the original abstract
Deep neural networks have been demonstrated to be vulnerable to backdoor attacks. Specifically, by injecting a small number of maliciously constructed inputs into the training set, an adversary is able to plant a backdoor into the trained model. This backdoor can then be activated during inference by a backdoor trigger to fully control the model's behavior. While such attacks are very effective, they crucially rely on the adversary injecting arbitrary inputs that are---often blatantly---mislabeled. Such samples would raise suspicion upon human inspection, potentially revealing the attack. Thus, for backdoor attacks to remain undetected, it is crucial that they maintain label-consistency---the condition that injected inputs are consistent with their labels. In this work, we leverage adversarial perturbations and generative models to execute efficient, yet label-consistent, backdoor attacks. Our approach is based on injecting inputs that appear plausible, yet are hard to classify, hence causing the model to rely on the (easier-to-learn) backdoor trigger.
Forward citations
Cited by 28 Pith papers
-
When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks
In the proportional high-dimensional regime, stronger backdoor training triggers improve clean accuracy and make attack success non-monotonic for regularized GLMs on Gaussian mixtures, with closed-form proofs for squa...
-
Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction
Scene-conditioned spatial-misdirection and duration-inflation backdoors succeed at 2.5–10% poison ratios on multimodal scanpath predictors and resist five adapted defenses.
-
Lilith: Backdoor Generalization under Training-Inference Trigger Shift
A single poisoned training trigger can create a backdoor that fires for a whole family of unseen inference-time triggers, provided the variants preserve the anchor's representation geometry.
-
Fast and Lightweight Backdoor Detection via Head Random Probing
HTell detects backdoors by random probing of the model head, reporting 99.03% true positive rate and 2.11% false positive rate at 12.69 ms per model on a benchmark of over 6700 models.
-
Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction
Backdoor attacks on VLM-based scanpath predictors can redirect fixations toward chosen objects or inflate durations using input-conditioned triggers that evade cluster detection, and no tested defense blocks them with...
-
Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning
SABLE shows that semantics-aware natural triggers enable effective backdoor attacks in federated learning against multiple aggregation rules while preserving benign accuracy.
-
Inevitable Encounters: Backdoor Attacks Involving Lossy Compression
ROI coding enables backdoor triggers to survive lossy compression by embedding malicious information into binary bitstreams via sample-specific or customized masks for both learned and traditional codecs.
-
From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense
PRISM uses a frozen VLM as an evolving semantic gatekeeper, reporting average attack success below 1% on CIFAR-10 while preserving clean accuracy.
-
Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor
Learning a backdoored reference model as a poisonous-sample oracle enables near-perfect training-time backdoor removal with negligible natural-accuracy loss.
-
Mitigating Backdoors via Decoy Shortcuts and Knowledge Decoupling
Backdoors can be trapped by a decoy shortcut branch trained alongside the main network and removed simply by discarding that branch, no extra data needed.
-
Temporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs
Retiming target-class neuromorphic events installs clean-label SNN backdoors with ASR up to 1.0 while leaving rate frames identical, and rate-collapsed defenses miss them.
-
Benign on Label, Malicious by Design: Clean-Label Dormant-to-Activated Backdoor via Machine Unlearning with Removable Camouflage
Dual generators jointly learn persistent clean-label triggers and removable camouflage so a backdoor stays dormant until attacker-submitted camouflage samples are unlearned.
-
Anti-Backdoor Coreset Selection via Cumulative Entropy
ABCS selects a high-entropy coreset from poisoned training data and trains on it, mitigating eight backdoor attacks with minimal natural-accuracy loss.
-
How Context Attribution Handles What the Model Already Knows
Context attribution methods cannot disentangle in-context from in-weight knowledge and assign unfaithful scores under overlap; new metrics and WMDP-Cyber++ quantify the failure.
-
Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection
Mirage achieves 73% misclassification success on LiDAR 3DOD models with 0.5% poisoning rate via label-consistent trigger injection.
-
Density-aware Sample-specific Attack
A density-aware sample-specific backdoor attack steers triggers into low-density regions via bilevel optimization to achieve high post-defense success rates on image datasets.
-
Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget
Checkerboard derives a closed-form checkerboard trigger for clean-label backdoor attacks that achieves over 94% ASR with poisoning rates as low as 0.46% on ImageNet-100 and 99.99% ASR with 20 samples on CIFAR-10.
-
CSC: Turning the Adversary's Poison against Itself
CSC identifies backdoored samples via early-epoch latent clustering and conceals them by relabeling to a virtual class, driving attack success rates near zero on benchmarks with little clean accuracy loss.
-
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
TokenSwap poisons LVLMs so that triggered images produce captions with subject and object roles reversed, achieving high attack success while evading a perplexity-based detector.
-
BadFU: Backdoor Federated Learning through Adversarial Machine Unlearning
A malicious federated-learning client can hide a backdoor by adding trigger-labeled samples plus camouflage samples, then activate it by requesting unlearning of the camouflage samples.
-
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
World-model-based embodied AI creates a predictive security boundary where attacks on data, sensors, imagination, ranking, and feedback can turn into unsafe physical action and false safety certificates.
-
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.
-
Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack
Pmeta-TLA combines a frame-level timbre leakage trigger with meta-learning and PCGrad to inject multiple backdoors into speech models in one training run, claiming better attack success, stealth, and lower cost than b...
-
TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models
TEMPO-Diffusion is a targeted backdoor attack framework for diffusion models that uses time-conditioned triggers to poison class-specific synthetic data, achieving high attack success in downstream classifiers.
-
Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks
Trigger color significantly affects semantic backdoor attack success in federated learning on CelebA hair-color classification, with white triggers better for blond targets and black for black targets.
-
Dataset Poisoning Attacks on Behavioral Cloning Policies
A few doctored demonstrations with a small red patch give attackers near-complete hidden control over behavior-cloning policies without lowering the policy's ordinary task reward.
-
Cryptographic Backdoor for Neural Networks: Boon and Bane
A signature-verification circuit attached to a neural network yields an undetectable backdoor attack and three privacy defenses: watermarking, user authentication, and IP-leak tracking.
-
Certification of Machine Learning Models via Directional Sharpness
Directional sharpness is introduced as a metric that correlates more strongly with generalization, identifies poor generalization more reliably, and supports efficient auditing and zero-knowledge certification.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.