REVIEW 2 major objections 4 minor 48 references
ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
T0 review · 2 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Adversarially augmenting the surrogate ViTs rather than reweighting their outputs is what unlocks stronger attack transferability.
desk verdict Plausible and useful ViT-specific ensemble augmentation with a serious presentation flaw: Algorithm 2 omits the L-infinity projection, which could inflate the headline gains unless the code clips. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the augmented ViT ensemble built from three stochastic structural perturbations, each an analogue of a data augmentation: masking (multi-head dropping), scaling (attention-score scaling), and mixup (MLP feature mixing). These create different backpropagation paths for the same input, so the accumulated gradient is an average over diverse network states rather than over one fixed architecture; the paper identifies that diversity as the source of transferability. Bayesian optimization tunes the randomization parameters against the remaining surrogates, automatic reweighting ($w_i \propto (L_{\max}/L_i)^b$) equalizes loss magnitudes across models so no single surrogate dominates the gradient, and step-size enlargement ($\alpha = q\epsilon/T$) drives the attack closer to convergence within ten iterations. Formally, the update is $x_{t+1}=x_t+\alpha\,\mathrm{sign}(g_{t+1})$ with momentum $g_{t+1}=\mu g_t + \nabla_{x_t}\sum_i w_i L_i$, where $L_i$ sums the losses of the three augmented variants of surrogate $i$ over two stochastic inference passes.
What would settle it
Run the released code at the paper's settings and measure the $\ell_\infty$ norm of every final adversarial image; if many exceed 16, re-run the Table 1 and Table 2 comparisons with per-iteration projection onto $\{\delta:\|\delta\|_\infty\le 16\}$ and check whether the reported 4.6-point and 15.3-point margins survive.
Extended reading notes
Core claim
The central claim is that the transferability bottleneck for ensemble attacks on ViTs lies in the surrogate models' lack of diversity, not in the ensemble combination rule. To remove that bottleneck, each original surrogate $f_i$ is replaced by three parameterised variants: Multi-head dropping ($f^{\mathrm{MHD}}_{\tau_i}$) zeroes attention heads whose random threshold falls below $\tau_i$; Attention score scaling ($f^{\mathrm{ASS}}_{s_i,\xi_i}$) multiplies each attention-score matrix elementwise by a random factor in $[s_i-\xi_i, s_i+\xi_i]$; and MLP feature mixing ($f^{\mathrm{MFM}}_{\rho_i}$) interpolates the MLP output with a randomly permuted copy using weight $\rho_i$. Bayesian optimization chooses the parameters by generating MI-FGSM adversarial examples on each variant and measuring their success against the other original surrogates. At attack time all variants contribute to a momentum gradient update whose ensemble weights are recomputed from per-model losses, and the step size is enlarged to $q\epsilon/T$ with $q=3$. The paper reports that this pipeline outperforms SVRE, AdaEA, and SMER on eight ViT targets and eight CNN targets under I-FGSM, MI-FGSM, DI-FGSM, and TI-FGSM, and that the model-augmentation module is the largest single contributor to the gain.
Load-bearing premise
The load-bearing premise is that the enlarged step size still respects the stated per-pixel perturbation budget, because Algorithm 2 prints the update $x_{t+1}=x_t+\alpha\,\mathrm{sign}(g_{t+1})$ with $\alpha=4.8$ over ten iterations and no explicit projection or final clipping, which would allow a raw perturbation of up to 48 against the declared $\epsilon=16$.
Editorial extensions
If this is right
- With DI-FGSM integration, the method reports near-saturating transfer to ViT targets, with attack success rates at or above 99% on most of the eight tested ViTs.
- Across eight CNN targets—including adversarially trained Inception ensembles and a hybrid ViT-CNN model—the method reports an average attack success rate of 88.3%, which is 15.3 percentage points above the SMER baseline.
- Ablation results show that model augmentation alone raises average success from 70.5% to 93.4% on ViTs and from 48.1% to 78.8% on CNNs, making it the dominant module.
- Automatic reweighting and step-size enlargement each improve over the baseline when used alone and are complementary to augmentation; combining all three modules gives the best reported results, 98.4% on ViTs and 88.3% on CNNs.
- Because model augmentation is a one-time preprocessing phase, the per-image attack cost in the second phase is comparable to or lower than that of prior ensemble baselines.
Reading between the lines
- If the missing projection in Algorithm 2 is not a typo, the unclipped update with $\alpha=4.8$ and $T=10$ can reach a per-pixel change of 48, three times the stated $\epsilon=16$ budget; re-running the comparisons with explicit clipping is the minimal check on whether the 15.3-point CNN margin is budget-fair.
- The paper's stated rationale predicts that the gain comes from diversity of the gradient-computation paths, not from anything attention-specific; a natural test is to apply analogous random channel dropping and feature-map mixing to CNN surrogates and see whether the transferability gain persists.
- Bayesian optimization here only needs the other surrogates as validation targets, so the same machinery could be used with a single surrogate by holding out some of its own augmented variants, making the attack usable when only one surrogate is available.
- The step-size enlargement result suggests a broader recipe for transfer attacks: within a fixed iteration count, larger-than-classical steps with a stabilizing mechanism can help escape adversarial overfitting, a hypothesis the paper only tests through its $q$ sweep.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes ViT-EnsembleAttack, an ensemble-based adversarial attack for Vision Transformers. Instead of only reweighting fixed surrogate models, the method augments each surrogate with three stochastic structural perturbations: Multi-head dropping (MHD), Attention score scaling (ASS), and MLP feature mixing (MFM). The augmentation parameters are tuned by Bayesian optimization using transfer performance to the other surrogate models, and the final attack ensembles the augmented variants with an automatic loss reweighting and an enlarged step size. Experiments on 1,000 ImageNet images report attack success rates against eight ViT targets and eight CNN targets, with claimed average gains over the SMER baseline of roughly 4.6 percentage points on ViTs and 15.3 percentage points on CNNs, together with ablations, hyperparameter sensitivity studies, and computational cost comparisons.
Significance. If the reported margins survive a properly constrained rerun, the paper would make a useful contribution to the adversarial-transferability literature: it reframes ensemble attacks as a problem of generating diverse surrogate variants rather than only combining fixed models, and it provides three ViT-specific augmentation strategies with a Bayesian model-selection procedure. The evaluation is broad, covering normally trained, robust, and adversarially trained ViT and CNN targets, and the authors state that code is released, which are strengths. The Bayesian optimization protocol is not circular: parameters are selected on sibling surrogate models and final evaluation is on disjoint held-out targets, which is legitimate model selection. However, two load-bearing issues currently prevent acceptance: Algorithm 2 does not enforce the L-infinity budget stated in Eq. (1), and the stochastic pipeline is reported without seeds, error bars, or confidence intervals. The central quantitative claims are therefore not yet supported.
major comments (2)
- [Algorithm 2, Section 4.1, Eq. (1)] The update in Algorithm 2 (line 21), x_{t+1} = x_t + alpha * sign(g_{t+1}), does not project or clip the perturbation back to the L-infinity ball of radius epsilon. With the hyperparameters in Section 4.1, alpha = q * epsilon / T = 3 * 16 / 10 = 4.8, so over T = 10 iterations the unconstrained L-infinity norm can reach 48, three times the stated epsilon = 16. The baselines use alpha = 1.6 and therefore remain within budget even without projection, so the margins in Tables 1 and 2 may reflect a larger effective perturbation budget rather than genuinely stronger transferability. The authors must either add an explicit projection/clipping step and state it in the pseudocode, or rerun all comparisons at the same constrained budget. As printed, the attack does not solve the constrained optimization problem defined in Eq. (1).
- [Section 4.1, Tables 1-2] The attack is stochastic in several places: MHD/ASS/MFM draw random masks and scales at each forward pass, Bayesian optimization includes randomness, and the inference loop samples multiple paths. Yet the paper reports no random seeds, confidence intervals, error bars, or repeated-run statistics. Since the claimed margins over SMER are 4.6% on ViTs and 15.3% on CNNs, and the strongest baselines are themselves stochastic, the reader cannot judge whether these margins are statistically meaningful. Please report mean and standard deviation (or equivalent) over at least three seeds for the proposed method and the baselines in Tables 1-3 and Figure 4.
minor comments (4)
- [Section 3.3, Attention score scaling] The text says random scaling factors follow a "uniform contribution"; this should presumably be "uniform distribution."
- [Algorithm 2, lines 5-7] The pseudocode is unclear about how the chosen strategy c and parameter p are passed to the objective function OF; clarify the signatures of gp_minimize calls so the three optimization loops are unambiguous.
- [Table 4] The table header "FLOPs (P)" calls FLOPs "floating-point operations per second," but the reported values appear to be total operations (peta-FLOPs); please correct the terminology.
- [Figure 3(d)] The caption states that computational cost grows "exponentially" with the inference loop count, but only a few discrete points are shown; either provide a precise scaling expression or soften the claim.
Circularity Check
No significant circularity: the tuned augmentation parameters are selected on sibling surrogate models with a separate image set, while the reported attack success rates are evaluated on disjoint held-out target models.
full rationale
ViT-EnsembleAttack's derivation chain is self-contained. The Bayesian-optimization objective (Algorithm 1) tunes each augmentation strategy's parameters by attacking the other sibling surrogates on a separate 4000-image set; the final ASR tables are evaluated on 1000 disjoint attack images and on target models (CaiT-S/24, TNT-S, LeViT-256, ConViT-B, RVT-S*, Drvit, Vit+DAT, ViT-B/16AT; Inception/ResNet variants, MobileViTv2) that are not part of the surrogate set (ViT-B/16, PiT-B, DeiT-B-Dis, Visformer-S). Thus no fitted parameter is folded into the reported result by construction, and no prediction is statistically forced by an input. Automatic Reweighting (Eq. 3) and Step Size Enlargement are algorithmic modules, not definitions that re-import the target quantity. The self-citations to the authors' earlier transfer-attack work are background methodology, not load-bearing support for the paper's central claim, and no uniqueness theorem from the authors' prior work is invoked. The only substantive concern found is an experimental-protocol one outside circularity: Algorithm 2's update step (line 21) does not explicitly clip x_{t+1} to the ε=16 perturbation ball, and with α=qε/T=4.8 the printed pseudocode could permit out-of-budget perturbations. This affects fairness and validity of the comparison but is not a circularity of the claimed derivation.
Assumptions & free parameters
free parameters (6)
- MHD threshold tau_i =
Optimized per surrogate via Bayesian optimization, values not reported
- ASS center s_i and half-range xi_i =
Optimized per surrogate, values not reported
- MFM mixing ratio rho_i =
Optimized per surrogate, values not reported
- Reweighting exponent b =
2
- Step size enlargement q =
3
- Inference loop count =
2
assumptions (3)
- domain assumption The L-infinity norm of the adversarial perturbation does not exceed epsilon=16 during the attack.
- domain assumption Randomized model augmentation improves surrogate diversity and reduces adversarial overfitting.
- domain assumption Average attack success rate on sibling surrogate models is a valid proxy for transferability to unseen targets.
Cite this review
Pith. "Pith review of ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers." pith.science (2026). https://pith.science/paper/TL2SYUWY
@misc{pith2026250812384,
author = {Pith},
title = {Pith review of: ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers},
year = {2026},
howpublished = {\url{https://pith.science/paper/TL2SYUWY}},
note = {Machine review of arXiv:2508.12384}
}
read the original abstract
Ensemble-based attacks have been proven to be effective in enhancing adversarial transferability by aggregating the outputs of models with various architectures. However, existing research primarily focuses on refining ensemble weights or optimizing the ensemble path, overlooking the exploration of ensemble models to enhance the transferability of adversarial attacks. To address this gap, we propose applying adversarial augmentation to the surrogate models, aiming to boost overall generalization of ensemble models and reduce the risk of adversarial overfitting. Meanwhile, observing that ensemble Vision Transformers (ViTs) gain less attention, we propose ViT-EnsembleAttack based on the idea of model adversarial augmentation, the first ensemble-based attack method tailored for ViTs to the best of our knowledge. Our approach generates augmented models for each surrogate ViT using three strategies: Multi-head dropping, Attention score scaling, and MLP feature mixing, with the associated parameters optimized by Bayesian optimization. These adversarially augmented models are ensembled to generate adversarial examples. Furthermore, we introduce Automatic Reweighting and Step Size Enlargement modules to boost transferability. Extensive experiments demonstrate that ViT-EnsembleAttack significantly enhances the adversarial transferability of ensemble-based attacks on ViTs, outperforming existing methods by a substantial margin. Code is available at https://github.com/Trustworthy-AI-Group/TransferAttack.
Figures
Reference graph
Works this paper leans on
-
[1]
An adaptive model ensemble adversarial attack for boosting adversarial transferability
Bin Chen, Jiali Yin, Shukai Chen, Bohao Chen, and Xi- meng Liu. An adaptive model ensemble adversarial attack for boosting adversarial transferability. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 4489–4498, 2023. 1, 3, 5
work page 2023
-
[2]
Visformer: The vision-friendly transformer
Zhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu, Longhui Wei, and Qi Tian. Visformer: The vision-friendly transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 589–598, 2021. 5
2021
-
[3]
Boosting adversarial at- tacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial at- tacks with momentum. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 9185–9193, 2018. 1, 2, 5
work page 2018
-
[4]
Evading defenses to transferable adversarial examples by translation-invariant attacks
Yinpeng Dong, Tianyu Pang, Hang Su, and Jun Zhu. Evading defenses to transferable adversarial examples by translation-invariant attacks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4312–4321, 2019. 1, 3, 5
work page 2019
-
[5]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. CoRR, abs/2010.11929, 2020. 1, 5
arXiv 2010
-
[6]
Convit: Improv- ing vision transformers with soft convolutional inductive bi- ases
St ´ephane d’Ascoli, Hugo Touvron, Matthew L Leavitt, Ari S Morcos, Giulio Biroli, and Levent Sagun. Convit: Improv- ing vision transformers with soft convolutional inductive bi- ases. In International conference on machine learning, pages 2286–2296. PMLR, 2021. 5
work page 2021
-
[7]
Fda: Feature disruptive attack
Aditya Ganeshan, Vivek BS, and R Venkatesh Babu. Fda: Feature disruptive attack. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8069– 8079, 2019. 2
work page 2019
-
[8]
Boosting Adversarial Transferability by Achieving Flat Local Maxima
Zhijin Ge, Hongying Liu, Xiaosen Wang, Fanhua Shang, and Yuanyuan Liu. Boosting Adversarial Transferability by Achieving Flat Local Maxima. In Proceedings of the Ad- vances in Neural Information Processing Systems, 2023. 2
work page 2023
Show all 48 references
-
[9]
Improving the Transferability of Adversarial Examples with Arbitrary Style Transfer
Zhijin Ge, Fanhua Shang, Hongying Liu, Yuanyuan Liu, Liang Wan, Wei Feng, and Xiaosen Wang. Improving the Transferability of Adversarial Examples with Arbitrary Style Transfer. In Proceedings of the ACM International Confer- ence on Multimedia, page 4440–4449, 2023. 1
2023
-
[10]
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014. 1, 2
2014 arXiv
-
[11]
Levit: a vision transformer in convnet’s clothing for faster inference
Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herv ´e J ´egou, and Matthijs Douze. Levit: a vision transformer in convnet’s clothing for faster inference. In Proceedings of the IEEE/CVF interna- tional conference on computer vision , pages 122...
-
[12]
Countering adversarial images using input transformations
Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens Van Der Maaten. Countering adversarial images using input transformations. arXiv preprint arXiv:1711.00117, 2017. 3
2017 arXiv
-
[13]
Transformer in transformer
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang. Transformer in transformer. Advances in neural information processing systems, 34:15908–15919,
-
[14]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 1, 5
2016
-
[15]
Rethinking spa- tial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh. Rethinking spa- tial dimensions of vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11936–11945, 2021. 5
2021
-
[16]
Accelerating stochastic gradi- ent descent using predictive variance reduction
Rie Johnson and Tong Zhang. Accelerating stochastic gradi- ent descent using predictive variance reduction. Advances in neural information processing systems, 26, 2013. 1, 3
2013
-
[17]
Ad- versarial examples in the physical world
Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. Ad- versarial examples in the physical world. In Artificial in- telligence safety and security , pages 99–112. Chapman and Hall/CRC, 2018. 2, 5
2018
-
[18]
Improving adversarial transferability via intermediate-level perturbation decay
Qizhang Li, Yiwen Guo, Wangmeng Zuo, and Hao Chen. Improving adversarial transferability via intermediate-level perturbation decay. Advances in Neural Information Pro- cessing Systems, 36, 2024. 1
2024
-
[19]
Learning transferable adversarial examples via ghost networks
Yingwei Li, Song Bai, Yuyin Zhou, Cihang Xie, Zhishuai Zhang, and Alan Yuille. Learning transferable adversarial examples via ghost networks. In Proceedings of the AAAI conference on artificial intelligence , pages 11458–11465,
-
[20]
Nesterov accelerated gradient and scale invariance for adversarial attacks
Jiadong Lin, Chuanbiao Song, Kun He, Liwei Wang, and John E Hopcroft. Nesterov accelerated gradient and scale invariance for adversarial attacks. arXiv preprint arXiv:1908.06281, 2019. 1, 2
1908 arXiv
-
[21]
Boosting adversarial transferability across model genus by deformation-constrained warping
Qinliang Lin, Cheng Luo, Zenghao Niu, Xilin He, We- icheng Xie, Yuanbo Hou, Linlin Shen, and Siyang Song. Boosting adversarial transferability across model genus by deformation-constrained warping. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 3459– ...
2024
-
[22]
Delving into transferable adversarial examples and black- box attacks
Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black- box attacks. arXiv preprint arXiv:1611.02770, 2016. 1, 2
2016 arXiv
-
[23]
Discrete representations strengthen vision transformer robustness
Chengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl V ondrick, Rahul Sukthankar, and Irfan Essa. Discrete representations strengthen vision transformer robustness. arXiv preprint arXiv:2111.10493, 2021. 5
2021 arXiv
-
[24]
Enhance the visual representation via discrete adversarial training
Xiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu, Gege Qi, Xiaodan Li, Rong Zhang, Hui Xue, et al. Enhance the visual representation via discrete adversarial training. Ad- vances in Neural Information Processing Systems, 35:7520– 7533, 2022. 5
2022
-
[25]
Towards robust vision transformer
Xiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li, Ranjie Duan, Shaokai Ye, Yuan He, and Hui Xue. Towards robust vision transformer. In Proceedings of the IEEE/CVF con- ference on Computer Vision and Pattern Recognition, pages 12042–12051, 2022. 5
2022
-
[26]
Separable self- attention for mobile vision transformers
Sachin Mehta and Mohammad Rastegari. Separable self- attention for mobile vision transformers. arXiv preprint arXiv:2206.02680, 2022. 5
2022 arXiv
-
[27]
When adversarial training meets vision trans- formers: Recipes from training to architecture
Yichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo, and Yisen Wang. When adversarial training meets vision trans- formers: Recipes from training to architecture. Advances in Neural Information Processing Systems, 35:18599–18611,
-
[28]
A self-supervised approach for adversarial robustness
Muzammal Naseer, Salman Khan, Munawar Hayat, Fa- had Shahbaz Khan, and Fatih Porikli. A self-supervised approach for adversarial robustness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 262–271, 2020. 3
2020
-
[29]
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211–252, 2015. 5
2015
-
[30]
Rethinking the inception archi- tecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception archi- tecture for computer vision. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 2818–2826, 2016. 5
2016
-
[31]
Inception-v4, inception-resnet and the im- pact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander Alemi. Inception-v4, inception-resnet and the im- pact of residual connections on learning. In Proceedings of the AAAI conference on artificial intelligence, 2017. 5
2017
-
[32]
Ensemble diversity facilitates adversarial transferability
Bowen Tang, Zheng Wang, Yi Bin, Qi Dou, Yang Yang, and Heng Tao Shen. Ensemble diversity facilitates adversarial transferability. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24377– 24386, 2024. 1, 3, 5, 6
2024
-
[33]
Training data-efficient image transformers & distillation through at- tention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv ´e J´egou. Training data-efficient image transformers & distillation through at- tention. In International conference on machine learning , pages 10347–10357. PMLR, 2021. 5
2021
-
[34]
Going deeper with im- age transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herv´e J´egou. Going deeper with im- age transformers. In Proceedings of the IEEE/CVF interna- tional conference on computer vision, pages 32–42, 2021. 5
2021
-
[35]
Ensemble adversarial training: Attacks and defenses
Florian Tram `er, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. Ensemble adversarial training: Attacks and defenses. arXiv preprint arXiv:1705.07204, 2017. 3, 5
2017 arXiv
-
[36]
Boosting adversarial transferability by block shuffle and rotation
Kunyu Wang, Xuanran He, Wenxuan Wang, and Xiaosen Wang. Boosting adversarial transferability by block shuffle and rotation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24336– 24346, 2024. 1
2024
-
[37]
Enhancing the Transferability of Adversarial Attacks through Variance Tuning
Xiaosen Wang and Kun He. Enhancing the Transferability of Adversarial Attacks through Variance Tuning. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1924–1933, 2021. 1
1924
-
[38]
Boosting Adversarial Transferability through En- hanced Momentum
Xiaosen Wang, Jiadong Lin, Han Hu, Jingdong Wang, and Kun He. Boosting Adversarial Transferability through En- hanced Momentum. In Proceedings of the British Machine Vision Conference, 2021. 1
2021
-
[39]
Struc- ture Invariant Transformation for better Adversarial Trans- ferability
Xiaosen Wang, Zeliang Zhang, and Jianping Zhang. Struc- ture Invariant Transformation for better Adversarial Trans- ferability. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4607–4619, 2023. 1
2023
-
[40]
Better diffusion models further improve adversarial training
Zekai Wang, Tianyu Pang, Chao Du, Min Lin, Weiwei Liu, and Shuicheng Yan. Better diffusion models further improve adversarial training. InInternational Conference on Machine Learning, pages 36246–36263. PMLR, 2023. 3
2023
-
[41]
Towards transferable adversarial attacks on vision transformers
Zhipeng Wei, Jingjing Chen, Micah Goldblum, Zuxuan Wu, Tom Goldstein, and Yu-Gang Jiang. Towards transferable adversarial attacks on vision transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2668– 2676, 2022. 1, 3
2022
-
[42]
Rethinking the backward propagation for adversarial transferability
Wang Xiaosen, Kangheng Tong, and Kun He. Rethinking the backward propagation for adversarial transferability. Ad- vances in Neural Information Processing Systems, 36:1905– 1922, 2023. 3
1905
-
[43]
Mitigating adversarial effects through random- ization
Cihang Xie, Jianyu Wang, Zhishuai Zhang, Zhou Ren, and Alan Yuille. Mitigating adversarial effects through random- ization. arXiv preprint arXiv:1711.01991, 2017. 3
2017 arXiv
-
[44]
Improving transferabil- ity of adversarial examples with input diversity
Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L Yuille. Improving transferabil- ity of adversarial examples with input diversity. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2730–2739, 2019. 2, 5
2019
-
[45]
Stochastic variance reduced ensemble adver- sarial attack for boosting the adversarial transferability
Yifeng Xiong, Jiadong Lin, Min Zhang, John E Hopcroft, and Kun He. Stochastic variance reduced ensemble adver- sarial attack for boosting the adversarial transferability. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 14983–14992,...
2022
-
[46]
Improving adversarial transferability via neuron attribution-based at- tacks
Jianping Zhang, Weibin Wu, Jen-tse Huang, Yizhan Huang, Wenxuan Wang, Yuxin Su, and Michael R Lyu. Improving adversarial transferability via neuron attribution-based at- tacks. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 14993–15002,
-
[47]
Jianping Zhang, Jen tse Huang, Wenxuan Wang, Yichen Li, Weibin Wu, Xiaosen Wang, Yuxin Su, and Michael R. Lyu. Improving the Transferability of Adversarial Samples by Path-Augmented Method. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , ...
2023
-
[48]
Bag of Tricks to Boost Adversarial Transfer- ability
Zeliang Zhang, Rongyi Zhu, Wei Yao, Xiaosen Wang, and Chenliang Xu. Bag of Tricks to Boost Adversarial Transfer- ability. 2024. 1
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.