Pith. sign in

REVIEW 3 cited by

Diversifying the High-level Features for better Adversarial Transferability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.10136 v2 pith:PB3QQINO submitted 2023-04-20 cs.CV

classification cs.CV
keywords attacksfeaturestransferabilityadversarialhigh-levelbetterdiversifyingdnns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given the great threat of adversarial attacks against Deep Neural Networks (DNNs), numerous works have been proposed to boost transferability to attack real-world applications. However, existing attacks often utilize advanced gradient calculation or input transformation but ignore the white-box model. Inspired by the fact that DNNs are over-parameterized for superior performance, we propose diversifying the high-level features (DHF) for more transferable adversarial examples. In particular, DHF perturbs the high-level features by randomly transforming the high-level features and mixing them with the feature of benign samples when calculating the gradient at each iteration. Due to the redundancy of parameters, such transformation does not affect the classification performance but helps identify the invariant features across different models, leading to much better transferability. Empirical evaluations on ImageNet dataset show that DHF could effectively improve the transferability of existing momentum-based attacks. Incorporated into the input transformation-based attacks, DHF generates more transferable adversarial examples and outperforms the baselines with a clear margin when attacking several defense models, showing its generalization to various attacks and high effectiveness for boosting transferability. Code is available at https://github.com/Trustworthy-AI-Group/DHF.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving

    cs.CV 2025-01 conditional novelty 6.0 of 10

    CAD is a transfer-based black-box attack using CLIP embeddings and ChatGPT-generated deceptive reasoning text to make vision-language autonomous driving models take unsafe actions.

  2. CogMorph: Cognitive Morphing Attacks for Text-to-Image Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    CogMorph escalates the toxicity of text-to-image outputs by contextually rewriting prompts with retrieved harmful features, claiming higher emotional harm than prior jailbreak attacks.

  3. Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Environmental illusions cause 5-7% accuracy drops in lane detection models and can trigger collisions in closed-loop simulation, with a proposed defense (MIDA) recovering ~4% robustness.

Pith tools