Pith. sign in

REVIEW 4 major objections 7 minor 154 references

SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity

T0 review · 4 major / 7 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read This systematization-of-knowledge paper claims that deep reinforcement learning research in cybersecurity is systematically undermined by 11 recurring methodological pitfalls—every one of the 66 papers it reviews exhibits at least two, with

desk verdict A genuinely useful SoK with a first DRL-specific pitfall taxonomy and honest case studies; just don't over-trust the exact prevalence numbers, and the author-overlap in the coded corpus needs addressing. read the letter →

arxiv 2602.08690 v2 pith:2PLDSIEV submitted 2026-02-09 cs.LG cs.CR

classification cs.LGcs.CR
keywords deepreinforcementlearningcybersecuritysystematicreviewmethodologicalpitfallsMDPspecificationevaluationbiasreproducibilitypolicyconvergence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Deep reinforcement learning is being applied to cybersecurity tasks such as network defense, malware evasion, fuzzing, and attack simulation, but this review argues that much of that literature is methodologically unsound. The authors define 11 recurring pitfalls spanning how security problems are modeled as Markov decision processes, how agents are trained, how results are evaluated, and how systems are assumed to behave in deployment. Reading 66 papers published from 2018 to 2025, they find that every paper exhibits at least two pitfalls, with an average of 5.8 per paper. They then run controlled experiments in three representative security domains to show that these pitfalls can degrade performance or inflate apparent success. If the review is right, a large share of published DRL-for-cybersecurity results do not currently support the deployable-security conclusions drawn from them.

What carries the argument

The central object is an 11-pitfall taxonomy organized by the four stages of applying DRL—environment modeling, agent training, performance evaluation, and system deployment—with each pitfall given a short definition, such as 'the environment is inherently a POMDP but is treated as fully observable.' The taxonomy does the work of making methodological quality measurable: two independent reviewers coded each of the 66 papers for each pitfall as present, partially present, or not present, yielding the prevalence statistics. A second mechanism is a set of controlled ablations in three security domains—autonomous network defense, adversarial malware creation, and web security testing—where the a

What would settle it

An independent team re-codes the same 66 papers with a pre-registered coding manual for the 11 pitfalls; if the average pitfall count drops well below 5.8 per paper, or inter-rater agreement falls below substantial, the central prevalence claim does not reproduce. A narrower check: in a case-study environment, replace the trained policy with random actions; if random actions already match trained performance, the paper's own gain-attribution warning is confirmed.

Watch

Extended reading notes

Core claim

The paper's central claim is that the DRL-for-cybersecurity literature contains systematic, quantifiable methodological failure modes rather than isolated mistakes. It identifies 11 pitfalls: incomplete MDP specification, incorrect MDP modeling, unaddressed partial observability, missing hyperparameter reporting, absent variance analysis, undemonstrated policy convergence, weak motivation for using DRL, misattributed performance gains, oversimplified environments, unrealistic deployment assumptions, and unhandled non-stationarity. Across 66 significant papers, 71.2% lack clear evidence of policy convergence, 66.7% neglect variance analysis, 60.6% fail to address partial observability, and 40

Load-bearing premise

The whole quantitative argument depends on the authors' 11 pitfall definitions being applied consistently and correctly by two reviewers across 66 papers; if the definitions are unstable across coders, the average of 5.8 pitfalls per paper loses its quantitative meaning.

Editorial extensions

If this is right

  • Reported gains in DRL-for-cybersecurity papers should not be taken at face value until convergence, variance, and baseline performance are shown; otherwise improvements may come from environment design rather than from the learned policy.
  • A minimum reporting standard follows: complete MDP definitions, hyperparameters, training curves, multiple-seed variance with confidence intervals, and ablations against random or simpler baselines.
  • Treating security tasks as fully observable when they are POMDPs is widespread; recurrent policies or explicit POMDP formulations should become the default in such settings.
  • Deployment claims should be stress-tested, because agents trained against fixed adversaries or fixed action orders degrade sharply when those assumptions change, so training should include non-stationarity.
  • The review implies that the field's empirical evidence base for deployable DRL-based security is weaker than the volume of publications suggests.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same pitfalls likely apply to adjacent work the review excluded, such as multi-agent RL security systems and non-deep RL security agents, so the findings probably generalize beyond the 66-paper corpus.
  • A concrete extension would be to convert the 11 pitfalls into a pre-registered reporting checklist or model card for DRL-for-cybersecurity papers, then measure whether adoption changes measured pitfall prevalence over time.
  • The authors' lenient coding rule—ambiguous cases were scored as less severe—means the headline prevalence figures are plausibly lower bounds; a stricter reviewer would likely count more pitfalls, not fewer.
  • The case-study pattern suggests a cheap validity test for any new DRL-for-cybersecurity result: compare the trained agent against a random-action policy in the same environment; if random actions already capture most of the performance, the contribution is in the environment, not in the learning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This SoK paper identifies 11 methodological pitfalls in applying deep reinforcement learning to cybersecurity, organized across modeling, training, evaluation, and deployment. It reports a systematic review of 66 papers (2018–2025), claiming that every paper exhibits at least two pitfalls, with an average of 5.8 per paper, and that the most common pitfall (policy convergence) appears in 71.2% of papers. The paper further demonstrates the practical impact of the pitfalls through controlled experiments in autonomous cyber defense (MiniCAGE), adversarial malware creation (AutoRobust), and web security testing (Link and SQiRL), and provides recommendations for each pitfall. Appendices include search details, reviewer agreement, and MDP/hyperparameter specifications.

Significance. If the prevalence estimates are reliable, this is a valuable contribution to a growing and methodologically uneven literature. The qualitative findings — that convergence evidence is frequently missing, variance is often unreported, partial observability is widely unaddressed, and environments are often simplified — are credible and align with broader reproducibility concerns in DRL. The paper's strengths include a two-reviewer protocol with reported agreement, case studies using 20 runs with confidence intervals, and detailed appendices with MDP specifications and hyperparameters. However, the central quantitative claims rest on a taxonomy developed from a pilot subset of the same corpus and on coding categories with only moderate inter-rater reliability on several pitfalls. These issues need to be addressed before the headline prevalence numbers can be treated as established facts about the DRL4Sec literature.

major comments (4)
  1. [Section 3 (Review Methodology)] The 11-pitfall taxonomy was induced from a pilot review of the 37 April-collection papers and then applied to the full 66-paper corpus that includes those same 37 papers. This is a circular measurement design: the taxonomy is fit to part of the data it is used to describe. No external validation, pre-registration, or holdout analysis is reported. Since the December 2025 collection (29 papers) was not used in taxonomy development, the paper should report whether the 'every paper contains at least two pitfalls' and 'average 5.8' statistics replicate in that held-out subset as a validation check. Without such evidence, the abstract's prevalence claims are not yet established as stable, generalizable facts.
  2. [Section 3, Table 4] Several constituent pitfalls have only moderate inter-rater reliability (M-MS κ=0.579, M-PO κ=0.583, E-GA κ=0.594, D-UA κ=0.570). The headline counts combine 'present' and 'partially present' into a single 'pitfall' category — for example, 71.2% for policy convergence is 42.4% present plus 28.8% partially present, and 28.8% for underlying assumptions is 15.2% plus 13.6%. Moderate coding stability in these categories can materially shift per-paper totals and the reported average of 5.8. The paper should report prevalence under a stricter 'present-only' threshold, present the distribution of per-paper pitfall counts, and discuss how sensitive the average is to plausible coding error.
  3. [Section 3 / Corpus] A number of papers in the 66-paper corpus are authored or co-authored by members of the review team (e.g., refs. [7, 15, 48, 51, 88, 91, 131, 132, 133]). The manuscript does not disclose this or provide a stratified comparison of author-affiliated versus non-affiliated papers. This is a potential source of systematic leniency or severity bias in coding. The paper should disclose the overlap and, at minimum, report a sensitivity analysis with those papers excluded or a comparison of their prevalence scores against the rest of the corpus.
  4. [Section 5.4, Table 1] The 'Distinct States' versus 'Original' Link comparison is used to argue that removing modeling pitfalls 'significantly' improves performance. However, the 95% confidence intervals overlap substantially (Original 75.2 [59.4, 90.9]; Distinct 85.3 [76.7, 94.0]), and no significance test or paired comparison is reported. The 10.1 percentage-point increase is not statistically supported by the reported data. The authors should either provide a paired bootstrap or other appropriate test, or soften the claim to a suggestive result. The same issue appears in some MiniCAGE and SQiRL comparisons where confidence intervals overlap.
minor comments (7)
  1. [Section 3, Table 4] The text reports an overall Cohen's kappa of 0.712, but Table 4 does not include the overall kappa. Adding it would make the summary agreement easier to verify.
  2. [Throughout] The spelling of 'SQiRL' is inconsistent: both 'SQIRL' and 'SQiRL' appear. Also 'W A VSEP' should be 'WAVSEP'.
  3. [Figure 4] The stacked bar chart is dense; the three severity categories are shown only via a legend, and the individual pitfall labels are small. Consider adding the exact percentages and counts in a companion table, or using a labeled dot plot.
  4. [Section 5.4] The text states that removing 'unnecessary partial observability' avoids M-MC and M-PO. Since the original environment is described as having valid Markovian transitions, the issue is better characterized as a conflated observation function causing partial observability, not a direct Markov violation. Clarify this distinction.
  5. [Section 6.4] The recommendation to use N≥5 while also citing N≥20 and N≥50 for robust confidence intervals may appear contradictory. Clarify that N≥5 is a pragmatic minimum, not a substitute for a power analysis.
  6. [Section 8.3, Table 3] Several cross-condition comparisons in the MiniCAGE deployment table have overlapping confidence intervals (e.g., Mixed training vs B-line evaluation: -40.2 [-47.1, -33.4] vs -19.0 [-20.9, -17.2] actually do not overlap, but others do). Reporting paired significance tests would strengthen the claims about degradation under changed assumptions.
  7. [Sections 5–8] The 'Security Implications' paragraphs are somewhat repetitive across pitfalls; tightening them would improve readability and make the taxonomy easier to scan.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: prevalence figures are empirical review outputs, not fitted predictions; case studies are independent re-implementations; self-citations are non-load-bearing.

full rationale

The paper is a systematic review and case-study paper rather than a derivation chain. The central quantitative claims (e.g., 'an average of 5.8' pitfalls per paper, '71.2%' lacking convergence evidence) are measurements obtained by applying the authors' 11 pitfall definitions to a 66-paper corpus, not predictions derived from fitted parameters. The taxonomy was induced from a pilot review of 37 papers and then applied to the full corpus, which raises a legitimate external-validity and potential-overfitting concern for the prevalence statistics, but it does not make any reported quantity equivalent to an input by construction: the pitfall definitions are not defined in terms of the prevalence outcomes, and no equation or fitted parameter is renamed as a prediction. The case-study demonstrations in MiniCAGE, AutoRobust, Link, and SQIRL are re-implementations of existing environments, some authored by the review team, but they are compared against random/baseline policies and serve as independent evidence of the pitfalls' impact rather than as circular justifications. Self-citations such as [12] for the review framework are methodological precedents, not load-bearing support for the paper's conclusions. The limitations noted in Section 9 (leniency in ambiguity, two-reviewer process, discussion-based resolution) describe coding-reliability mitigations; the lack of external validation or pre-registration is a correctness/robustness risk about the prevalence numbers, not a circularity of the kind defined by the analysis rules.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim is an aggregate review statistic, not a fitted model, so there are no free parameters fitted to data. The axioms listed are the load-bearing premises behind the prevalence and case-study claims. The 11 pitfalls are analytical categories rather than invented physical or mathematical entities, so no invented entities are recorded.

assumptions (3)
  • domain assumption The 66-paper corpus selected via venue-tier and citation-count thresholds is representative enough to estimate pitfall prevalence in DRL4Sec literature.
    Prevalence percentages (e.g., 71.2% for policy convergence) are computed over this corpus; if the selection criteria skew toward higher- or lower-quality papers, the averages change. The authors acknowledge this in Section 9.
  • domain assumption The 11 pitfall definitions can be applied reliably by two reviewers, and the resulting present/partial/not coding reflects the papers' actual methodology.
    Section 3 reports Cohen's kappa 0.712 and 81.5% agreement, but the scheme was created by the authors from a pilot subset; no external audit or pre-registration is provided.
  • domain assumption The MDP/POMDP formalization is the correct lens for evaluating whether a cybersecurity DRL paper is methodologically sound.
    Used throughout Sections 5-8; a different evaluation lens (e.g., practical deployment metrics) might rank pitfalls differently.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity." pith.science (2026). https://pith.science/paper/2PLDSIEV

@misc{pith2026260208690,
  author       = {Pith},
  title        = {Pith review of: SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2PLDSIEV}},
  note         = {Machine review of arXiv:2602.08690}
}
read the original abstract

Deep Reinforcement Learning (DRL) has achieved remarkable success in domains requiring sequential decision-making, motivating its application to cybersecurity problems. However, transitioning DRL from laboratory simulations to bespoke cyber environments can introduce numerous issues. This is further exacerbated by the often adversarial, non-stationary, and partially-observable nature of most cybersecurity tasks. In this paper, we identify and systematize 11 methodological pitfalls that frequently occur in DRL for cybersecurity (DRL4Sec) literature across the stages of environment modeling, agent training, performance evaluation, and system deployment. By analyzing 66 significant DRL4Sec papers (2018-2025), we quantify the prevalence of each pitfall and find an average of over five pitfalls per paper. We demonstrate the practical impact of these pitfalls using controlled experiments in (i) autonomous cyber defense, (ii) adversarial malware creation, and (iii) web security testing environments. Finally, we provide actionable recommendations for each pitfall to support the development of more rigorous and deployable DRL-based security systems.

Figures

Figures reproduced from arXiv: 2602.08690 by the authors.

Figure 1
Figure 1. Common pitfalls of DRL when applied to cybersecurity, organized by development stage and relevant case studies. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Training performance of different hyperparameter [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Comparison of mean with 95/99% CI performance [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Complete breakdown of pitfall prevalence [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

154 extracted references · 17 linked inside Pith

  1. [1]

    Created by Maxwell Standen, David Bowman, Son Hoang, Toby Richer, Martin Lucas, Richard Van Tassel

    Cyber autonomy gym for experimentation chal- lenge 1.https://github.com/cage-challenge/ cage-challenge-1, 2021. Created by Maxwell Standen, David Bowman, Son Hoang, Toby Richer, Martin Lucas, Richard Van Tassel

  2. [2]

    Created by Maxwell Standen, David Bowman, Son Hoang, Toby Richer, Martin Lucas, Richard Van Tassel, Phillip Vu, Mitchell Kiely

    Cyber autonomy gym for experimentation chal- lenge 2.https://github.com/cage-challenge/ cage-challenge-2, 2022. Created by Maxwell Standen, David Bowman, Son Hoang, Toby Richer, Martin Lucas, Richard Van Tassel, Phillip Vu, Mitchell Kiely

  3. [3]

    Adkins, M

    J. Adkins, M. Bowling, and A. White. A method for evaluating hyperparameter sensitivity in reinforcement learning.Advances in Neural Information Processing Systems, 2024

  4. [4]

    Agarwal, M

    R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. G. Bellemare. Deep reinforcement learning at the edge of the statistical precipice.CoRR, 2021

  5. [5]

    Agarwal, M

    R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. G. Bellemare. Reincarnating rein- forcement learning: Reusing prior computation to ac- celerate progress.Advances in Neural Information Pro- cessing Systems, 2022

  6. [6]

    Al-Fawa’reh, J

    M. Al-Fawa’reh, J. Abu-Khalaf, P. Szewczyk, and James Jin Kang. Malbot-drl: Malware botnet detec- tion using deep reinforcement learning in iot networks. IEEE Internet of Things Journal, 2023

  7. [7]

    Al Wahaibi, M

    S. Al Wahaibi, M. Foley, and S. Maffeis. SQIRL: Grey- box detection of sql injection vulnerabilities using rein- forcement learning. In32nd USENIX Security Sympo- sium (USENIX Security 23), 2023

  8. [8]

    Kharkar, B

    Hyrum S Anderson, A. Kharkar, B. Filar, D. Evans, and P. Roth. Learning to evade static pe machine learn- ing malware models via reinforcement learning.arXiv preprint arXiv:1801.08917, 2018

Show all 154 references
  1. [9]

    Andrew, S

    A. Andrew, S. Spillard, J. Collyer, and N. Dhir. De- veloping optimal causal cyber-defence agents via cyber security simulation. InWorkshop on Machine Learning for Cybersecurity (ML4Cyber), 2022

  2. [10]

    Applebaum, C

    A. Applebaum, C. Dennler, P. Dwyer, M. Moskowitz, H. Nguyen, N. Nichols, N. Park, P. Rachwalski, F. Rau, A. Webster, and M. Wolk. Bridging Automated to Autonomous Cyber Defense: Foundational Analysis of Tabular Q-Learning. InProceedings of the 15th ACM Workshop on Artificial I...

  3. [11]

    Apruzzese, M

    G. Apruzzese, M. A.reolini, M. Marchetti, A.rea Ven- turi, and M. Colajanni. Deep reinforcement adversarial learning against botnet evasion attacks.IEEE Transac- tions on Network and Service Management, 2020

  4. [12]

    D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pier- azzi, C. Wressnegger, L. Cavallaro, and K. Rieck. Dos and Don’ts of Machine Learning in Computer Security. InProc. of the USENIX Security Symposium, 2022

  5. [13]

    Arulkumaran, M

    K. Arulkumaran, M. P. Deisenroth, M. Brundage, and Anil Anthony Bharath. Deep reinforcement learning: A brief survey.IEEE Signal Processing Magazine, 2017

  6. [14]

    A. S. Basnet, M. C. Ghanem, D. Dunsin, H. Khed- dar, and W. Sowinski-Mydlarz. Advanced persistent threats (apt) attribution using deep reinforcement learn- ing.Digital Threats: Research and Practice, 2025

  7. [15]

    Bates, V

    E. Bates, V . Mavroudis, and C. Hicks. Reward shap- ing for happier autonomous cyber security agents. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 2023

  8. [16]

    Beyond rewards in reinforcement learning for cyber de- fence, 2026

    Elizabeth Bates, Chris Hicks, and Vasilios Mavroudis. Beyond rewards in reinforcement learning for cyber de- fence, 2026

  9. [17]

    C. et al. Berner. Dota 2 with large scale deep reinforce- ment learning.arXiv preprint arXiv:1912.06680, 2019

  10. [18]

    B ¨ottinger, P

    K. B ¨ottinger, P. Godefroid, and R. Singh. Deep rein- forcement fuzzing. In2018 IEEE Security and Privacy Workshops (SPW). IEEE, 2018

  11. [19]

    Boutilier, T

    C. Boutilier, T. Dean, and S. Hanks. Decision-theoretic planning: Structural assumptions and computational leverage.Journal of Artificial Intelligence Research, 1999. 13

  12. [20]

    Burda, H

    Y . Burda, H. Edwards, A. Storkey, and O. Klimov. Ex- ploration by random network distillation.arXiv preprint arXiv:1810.12894, 2018

  13. [21]

    Caminero, M

    G. Caminero, M. Lopez-Martin, and B. Carro. Adver- sarial environment reinforcement learning algorithm for intrusion detection.Computer Networks, 2019

  14. [22]

    Cavenaghi, G

    E. Cavenaghi, G. Sottocornola, F. Stella, and M. Zanker. A systematic study on reproducibility of reinforcement learning in recommendation systems.ACM Transac- tions on Recommender Systems, 2023

  15. [23]

    S. C. Chan, S. Fishman, J. Canny, A. Korattikara, and S. Guadarrama. Measuring the reliability of reinforcement learning algorithms.arXiv preprint arXiv:1912.05663, 2019

  16. [24]

    Chatterjee and Akbar-Siami Namin

    M. Chatterjee and Akbar-Siami Namin. Detecting phishing websites through deep reinforcement learning. In2019 IEEE 43rd annual computer software and ap- plications conference (COMPSAC). IEEE, 2019

  17. [25]

    J. Chen, S. Hu, H. Zheng, C. Xing, and G. Zhang. Gail- pt: An intelligent penetration testing framework with generative adversarial imitation learning.Computers & Security, 2023

  18. [26]

    L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch. Decision transformer: Reinforcement learning via se- quence modeling.Advances in neural information pro- cessing systems, 2021

  19. [27]

    S. Chen, N. Carlini, and D. Wagner. Stateful detection of black-box adversarial attacks. InProceedings of the 1st ACM Workshop on Security and Privacy on Artifi- cial Intelligence, 2020

  20. [28]

    X. Chen, Y . Nie, W. Guo, and X. Zhang. When llm meets drl: Advancing jailbreaking efficiency via drl- guided search.Advances in Neural Information Pro- cessing Systems, 2024

  21. [29]

    K. L. Chung. Markov chains.Springer-Verlag, New York, 1967

  22. [30]

    Colas, O

    C. Colas, O. Sigaud, and P. Oudeyer. How many ran- dom seeds? statistical power analysis in deep reinforce- ment learning experiments.CoRR, 2018

  23. [31]

    Collyer, A

    J. Collyer, A. A.rew, and D. Hodges. ACD-G: En- hancing Autonomous Cyber Defense Agent General- ization Through Graph Embedded Network Represen- tation. InWorkshop on Machine Learning for Cyber- security (ML4Cyber) as part of the Proceedings of the 39 th International Conferen...

  24. [32]

    Corradini, Z

    D. Corradini, Z. Montolli, M. Pasqua, and M. Ceccato. Deeprest: Automated test case generation for rest apis exploiting deep reinforcement learning. InProceedings of the 39th IEEE/ACM International Conference on Au- tomated Software Engineering, 2024

  25. [33]

    Dabney, G

    W. Dabney, G. Ostrovski, D. Silver, and R. Munos. Im- plicit quantile networks for distributional reinforcement learning.CoRR, 2018

  26. [34]

    D. B. D’Ambrosio, S. Abeyruwan, L. Graesser, A. Is- cen, H. B. Amor, A. Bewley, B. J. Reed, K. Reymann, L. Takayama, Y . Tassa, K. Choromanski, E. Coumans, D. Jain, N. Jaitly, N. Jaques, S. Kataoka, Y . Kuang, N. Lazic, R. Mahjourian, S. Moore, K. Oslund, A. Shankar, V . Sindh...

  27. [35]

    De Silva, W

    R. De Silva, W. Guo, N. Ruaro, I. Grishchenko, C. Kruegel, and G. Vigna.{GuideEnricher}: Protecting the anonymity of ethereum mixing service users with deep reinforcement learning. In33rd USENIX Security Symposium (USENIX Security 24), 2024

  28. [36]

    X. Deng, M. Cen, M Jiang, and M. Lu. Ransomware early detection using deep reinforcement learning on portable executable header.Cluster Computing, 2024

  29. [37]

    R. R. Dos Santos, E. K. Viegas, A. O. Santin, and V . V . Cogo. Reinforcement learning for intrusion detection: More model longness and fewer updates.IEEE Trans- actions on Network and Service Management, 2022

  30. [38]

    Dulac-Arnold, D

    G. Dulac-Arnold, D. Mankowitz, and T. Hester. Chal- lenges of real-world reinforcement learning.arXiv preprint arXiv:1904.12901, 2019

  31. [39]

    Eimer, M

    T. Eimer, M. Lindauer, and R. Raileanu. Hyperparame- ters in reinforcement learning and how to tune them. In International conference on machine learning. PMLR, 2023

  32. [40]

    Emerson, L

    H. Emerson, L. Bates, C. Hicks, and V . Mavroudis. Cy- bORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents, 2024

  33. [41]

    J. Eom, S. Jeong, and T. Kwon. Fuzzing javascript interpreters with coverage-guided reinforcement learn- ing for llm-based mutation. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2024

  34. [42]

    Erd ˝odi, ˚A Sommervoll, and F

    L. Erd ˝odi, ˚A Sommervoll, and F. M. Zennaro. Sim- ulating sql injection vulnerability exploitation using q- learning reinforcement learning agents.Journal of In- formation Security and Applications, 2021

  35. [43]

    Evertz, N

    J. Evertz, N. Risse, N. Neuer, A. M ¨uller, P. Normann, G. Sapia, S. Gupta, D. Pape, S. Shaw, D. Srivastav, et al. Chasing shadows: Pitfalls in llm security research. arXiv preprint arXiv:2512.09549, 2025

  36. [44]

    Faillon, B

    M. Faillon, B. Bout, J. Francq, C. Neal, N. Boulahia- Cuppens, Fr´ed´eric Cuppens, and R. Yaich. How to bet- ter fit reinforcement learning for pentesting: A new hi- erarchical approach. InEuropean Symposium on Re- search in Computer Security. Springer, 2024

  37. [45]

    Z. Fang, J. Wang, B. Li, S. Wu, Y . Zhou, and H. Huang. Evading anti-malware engines with deep reinforcement learning.IEEE Access, 2019

  38. [46]

    Fawzi, M

    A. Fawzi, M. Balog, A. Huang, T. Hubert, B. Romera- Paredes, M. Barekatain, A. Novikov, Francisco J. R. Ruiz, J. Schrittwieser, G. Swirszcz, D. Silver, D. Hassabis, and P. Kohli. Discovering faster matrix 14 multiplication algorithms with reinforcement learning. Nature, 2022

  39. [47]

    R. Feng, A. Hooda, N. Mangaokar, K. Fawaz, S. Jha, and A. Prakash. Stateful defenses for machine learning models are not yet secure against black-box attacks. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, 2023

  40. [48]

    Foley, C

    M. Foley, C. Hicks, K. Highnam, and V . Mavroudis. Autonomous Network Defence Using Reinforcement Learning. InProceedings of the 2022 ACM on Asia Conference on Computer and Communications Secu- rity, ASIA CCS ’22, 2022

  41. [49]

    Foley and S

    M. Foley and S. Maffeis. HAXSS: Hierarchical Rein- forcement Learning for XSS Payload Generation. In 2022 IEEE International Conference on Trust, Security and Privacy in Computing and Communications (Trust- Com). IEEE, 2022

  42. [50]

    Foley and S

    M. Foley and S. Maffeis. Apirl: Deep reinforcement learning for rest api fuzzing. InProceedings of the AAAI Conference on Artificial Intelligence, 2025

  43. [51]

    Foley, M

    M. Foley, M. Wang, Z. M, C. Hicks, and V . Mavroudis. Inroads into Autonomous Network Defence using Ex- plained Reinforcement Learning. InConference on Ap- plied Machine Learning in Information Security (CAM- LIS), 2022

  44. [52]

    Gangupantulu, T

    R. Gangupantulu, T. Cody, P. Park, A. Rahman, L. Eisenbeiser, D. Radke, R. Clark, and C. Redino. Us- ing cyber terrain in reinforcement learning for penetra- tion testing. In2022 IEEE International Conference on Omni-layer Intelligent Systems (COINS). IEEE, 2022

  45. [53]

    Gleave, M

    A. Gleave, M. Dennis, C. Wild, N. Kant, S. Levine, and S. Russell. Adversarial policies: Attacking deep rein- forcement learning.arXiv preprint arXiv:1905.10615, 2019

  46. [54]

    D. Goel, K. Moore, M. Guo, D. Wang, M. Kim, and S. Camtepe. Optimizing cyber defense in dynamic ac- tive directories through reinforcement learning. InEu- ropean Symposium on Research in Computer Security. Springer, 2024

  47. [55]

    Gohil, H

    V . Gohil, H. Guo, S. Patnaik, and J. Rajendran. At- trition: Attacking static hardware trojan detection tech- niques using reinforcement learning. InProceedings of the 2022 ACM SIGSAC conference on computer and communications security, 2022

  48. [56]

    Gohil, S

    V . Gohil, S. Patnaik, D. Kalathil, and J. Rajendran. At- tackGNN: Red-Teaming GNNs in hardware security us- ing reinforcement learning. In33rd USENIX Security Symposium (USENIX Security 24), 2024

  49. [57]

    Ttcp cage chal- lenge 3.https://github.com/cage-challenge/ cage-challenge-3, 2022

    TTCP CAGE Working Group. Ttcp cage chal- lenge 3.https://github.com/cage-challenge/ cage-challenge-3, 2022

  50. [58]

    Ttcp cage chal- lenge 4.https://github.com/cage-challenge/ cage-challenge-4, 2023

    TTCP CAGE Working Group. Ttcp cage chal- lenge 4.https://github.com/cage-challenge/ cage-challenge-4, 2023

  51. [59]

    Y . Han, B. I. Rubinstein, T. Abraham, T. Alpcan, O. De Vel, S. Erfani, D. Hubczenko, C. Leckie, and P. Montague. Reinforcement learning for autonomous defence in software-defined networking. InInterna- tional conference on decision and game theory for se- curity. Springer, 2018

  52. [60]

    Hausknecht and P

    M. Hausknecht and P. Stone. Deep recurrent q-learning for partially observable mdps. In2015 aaai fall sympo- sium series, 2015

  53. [61]

    M. He, X. Wang, P. Wei, L. Yang, Y . Teng, and R. Lyu. Reinforcement learning meets network intrusion de- tection: A transferable and adaptable framework for anomaly behavior identification.IEEE Transactions on Network and Service Management, 2024

  54. [62]

    Henderson, R

    P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Pre- cup, and D. Meger. Deep reinforcement learning that matters. InProceedings of the AAAI conference on ar- tificial intelligence, 2018

  55. [63]

    Hessel, H

    M. Hessel, H. van Hasselt, J. Modayil, and D. Silver. On inductive biases in deep reinforcement learning.arXiv preprint arXiv:1907.02908, 2019

  56. [64]

    Hicks, V

    C. Hicks, V . Mavroudis, M. Foley, T. Davies, K. High- nam, and T. Watson. Canaries and Whistles: Re- silient Drone Communication Networks with (or with- out) Deep Reinforcement Learning. InProceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec ’23, 2023

  57. [65]

    S. Hore, J. Ghadermazi, D. Paudel, A. Shah, T. Das, and N. Bastian. Deep packgen: A deep reinforcement learning framework for adversarial network packet gen- eration.ACM Transactions on Privacy and Security, 2025

  58. [66]

    C. Hou, M. Zhou, Y . Ji, P. Daian, F. Tramer, G. Fanti, and A. Juels. Squirrl: Automating attack analysis on blockchain incentive mechanisms with deep reinforce- ment learning.arXiv preprint arXiv:1912.01798, 2019

  59. [67]

    Hsu and M

    Y . Hsu and M. Matsuoka. A deep reinforcement learn- ing approach for anomaly network intrusion detection system. In2020 IEEE 9th international conference on cloud networking (CloudNet). IEEE, 2020

  60. [68]

    Z. Hu, R. Beuran, and Y . Tan. Automated penetra- tion testing using deep reinforcement learning. In2020 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW). IEEE, 2020

  61. [69]

    Jumper, R

    J. Jumper, R. Evans, A. Pritzel, T. Green, M. Fig- urnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. ˇZ´ıdek, A. Potapenko, A. Bridgland, C. Meyer, Si- mon A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera- Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Pe- tersen, D. ...

  62. [70]

    L. P. Kaelbling, Michael L Littman, and Andrew W Moore. Reinforcement learning: A survey.Journal of artificial intelligence research, 1996

  63. [71]

    Kaiser, M

    L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H Campbell, K. Czechowski, D. Erhan, C. Finn, P. Koza- kowski, S. Levine, et al. Model-based reinforcement learning for atari.arXiv preprint arXiv:1903.00374, 2019

  64. [72]

    Z. Kan, S. McFadden, D. Arp, F. Pendlebury, R. Jor- daney, J. Kinder, F. Pierazzi, and L. Cavallaro. Tesser- act: Eliminating experimental bias in malware classifi- cation across space and time (extended version).arXiv preprint arXiv:2402.01359, 2024

  65. [73]

    S. Kim, S. Yoon, Jin-Hee Cho, D. S. Kim, Terrence J Moore, F. Free-Nelson, and H. Lim. Divergence: Deep reinforcement learning-based adaptive traffic inspection and moving target defense countermeasure framework. IEEE Transactions on Network and Service Manage- ment, 2022

  66. [74]

    Klees, A

    G. Klees, A. Ruef, B. Cooper, S. Wei, and M. Hicks. Evaluating Fuzz Testing. InProceedings of the 2018 ACM SIGSAC Conference on Computer and Communi- cations Security, CCS ’18, 2018

  67. [75]

    Kvasov, M

    A. Kvasov, M. Sahin, C. Hebert, and Anderson Santana De Oliveira. Simulating deception for web applications using reinforcement learning. InEuropean Symposium on Research in Computer Security. Springer, 2023

  68. [76]

    Landen, K

    M. Landen, K. Chung, M. Ike, S. Mackay, Jean-Paul Watson, and W. Lee. Dragon: Deep reinforcement learning for autonomous grid operation and attack de- tection. InProceedings of the 38th Annual Computer Security Applications Conference, 2022

  69. [77]

    Landis and G

    J R. Landis and G. G Koch. The measurement of ob- server agreement for categorical data.biometrics, 1977

  70. [78]

    Le Tolguenec, E

    P. Le Tolguenec, E. Rachelson, Y . Besse, F. Teichteil- Koenigsbuch, N. Schneider, H ´el`ene Waeselynck, and D. Wilson. Exploration-driven reinforcement learning for avionic system fault detection (experience paper). In Proceedings of the 33rd ACM SIGSOFT International Symposi...

  71. [79]

    S. Lee, S. Wi, and S. Son. Link: Black-box detection of cross-site scripting vulnerabilities using reinforcement learning. InProceedings of the ACM Web Conference 2022, 2022

  72. [80]

    Q. Li, M. Hu, H. H., M. Zhang, and Y . Li. Innes: An intelligent network penetration testing model based on deep reinforcement learning.Applied Intelligence, 2023

  73. [81]

    X. Li, X. Liu, L. Chen, R. Prajapati, and D. Wu. Al- phaprog: reinforcement generation of valid programs for compiler fuzzing. InProceedings of the AAAI Con- ference on Artificial Intelligence, 2022

  74. [82]

    Z. Li, C. Huang, S. Deng, W. Qiu, and X. Gao. A soft actor-critic reinforcement learning algorithm for net- work intrusion detection.Computers & Security, 2023

  75. [83]

    Lopez-Martin, B

    M. Lopez-Martin, B. Carro, and A. Sanchez- Esguevillas. Application of deep reinforcement learn- ing to intrusion detection for supervised problems.Ex- pert Systems with Applications, 2020

  76. [84]

    M. Luo, W. Xiong, G. Lee, Y . Li, X. Yang, A. Zhang, Y . Tian, Hsien-Hsin S Lee, and G Edward Suh. Auto- cat: Reinforcement learning for automated exploration of cache-timing attacks. In2023 IEEE International Symposium on High-Performance Computer Architec- ture (HPCA). IEEE, 2023

  77. [85]

    M. C. Machado, M. G. Bellemare, E. Talvitie, J. Ve- ness, M. Hausknecht, and M. Bowling. Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents.Journal of Artificial Intelligence Research, 2018

  78. [86]

    Maeda and M

    R. Maeda and M. Mimura. Automating post- exploitation with deep reinforcement learning.Com- puters & Security, 2021

  79. [87]

    Mavroudis, G

    V . Mavroudis, G. Palmer, S. Farmer, K. S. Whitehead, D. Foster, A. Price, I. Miles, A. Caron, and S. Pasteris. Guidelines for applying rl and marl in cybersecurity ap- plications.arXiv preprint arXiv:2503.04262, 2025

  80. [88]

    McFadden, M

    S. McFadden, M. Foley, M. D’Onghia, C. Hicks, V . Mavroudis, N. Paoletti, and F. Pierazzi. Drmd: Deep reinforcement learning for malware detection un- der concept drift. InProc. of the AAAI Conference on Artificial Intelligence, 2026

  81. [89]

    McFadden, M

    S. McFadden, M. Kan, L. Cavallaro, and F. Pierazzi. The impact of active learning on availability data poi- soning for android malware classifiers. InProceedings of the Annual Computer Security Applications Confer- ence Workshops (ACSAC Workshops). IEEE, 2024

  82. [90]

    McFadden, Z

    S. McFadden, Z. Kan, L. Cavallaro, and F. Pierazzi. Poster: Rpal-recovering malware classifiers from data poisoning using active learning. InProceedings of the 2023 ACM SIGSAC Conference on Computer and Com- munications Security, 2023

  83. [91]

    McFadden, M

    S. McFadden, M. Maugeri, C. Hicks, V . Mavroudis, and F. Pierazzi. Wendigo: Deep reinforcement learn- ing for denial-of-service query discovery in graphql. In IEEE Workshop on Deep Learning Security and Privacy (DLSP), 2024

  84. [92]

    L. Meng, R. Gorbet, and D. Kuli ´c. Memory-based deep reinforcement learning for pomdps. In2021 IEEE/RSJ international conference on intelligent robots and sys- tems (IROS). IEEE, 2021

  85. [93]

    Mirhoseini, A

    A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. Songhori, S. Wang, Young-Joon Lee, E. Johnson, O. Pathak, A. Nova, J. Pak, A. Tong, K. Srinivasa, W. Hang, E. Tuncer, Quoc V . Le, J. Laudon, R. Ho, R. Carpenter, and J. Dean. A graph placement method- ology for fast chipdes...

  86. [94]

    V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lilli- crap, T. Harley, D. Silver, and K. Kavukcuoglu. Asyn- chronous methods for deep reinforcement learning. In 16 International conference on machine learning. PmLR, 2016

  87. [95]

    V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller. Playing atari with deep reinforcement learning.arXiv preprint arXiv:1312.5602, 2013

  88. [96]

    Mohamed and R

    S. Mohamed and R. Ejbali. Deep sarsa-based reinforce- ment learning approach for anomaly network intrusion detection system.International Journal of Information Security, 2023

  89. [97]

    Moore, A

    A. Moore, A. Burke, M. Foley, A. Knack, C. Hicks, and V . Mavroudis. A fundamental research plan for au- tonomous cyber defence, 2025

  90. [98]

    T. T. Nguyen and Vijay Janapa Reddi. Deep reinforce- ment learning for cyber security.IEEE Transactions on Neural Networks and Learning Systems, 2021

  91. [99]

    L. Nie, W. Sun, S. Wang, Z. Ning, J. J. Rodrigues, Y . Wu, and S. Li. Intrusion detection in green internet of things: a deep deterministic policy gradient-based al- gorithm.IEEE Transactions on Green Communications and Networking, 2021

  92. [100]

    Nikishin, M

    E. Nikishin, M. Schwarzer, P. D’Oro, Pierre-Luc Bacon, and A. Courville. The Primacy Bias in Deep Reinforce- ment Learning. InProceedings of the 39th International Conference on Machine Learning, 2022

  93. [101]

    Nyberg and P

    J. Nyberg and P. Johnson. Training automated defense strategies using graph-based cyber attack simulations. arXiv preprint arXiv:2304.11084, 2023

  94. [102]

    Osband, C

    I. Osband, C. Blundell, A. Pritzel, and B. Van Roy. Deep exploration via bootstrapped DQN.CoRR, 2016

  95. [103]

    Patterson, S

    A. Patterson, S. Neumann, M. White, and A. White. Empirical design in reinforcement learning.Journal of Machine Learning Research, 2024

  96. [104]

    Paudel and G

    B. Paudel and G. Amariucai. Reinforcement learning approach to generate zero-dynamics attacks on control systems without state space models. InEuropean Sym- posium on Research in Computer Security. Springer, 2023

  97. [105]

    Pendlebury, F

    F. Pendlebury, F. Pierazzi, R. Jordaney, J. Kinder, and L. Cavallaro. TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time. In28th USENIX Security Symposium (USENIX Security 19), 2019

  98. [106]

    T. V . Phan and T. Bauschert. Deepair: Deep rein- forcement learning for adaptive intrusion response in software-defined networks.IEEE Transactions on Net- work and Service Management, 2022

  99. [107]

    Pierazzi, F

    F. Pierazzi, F. Pendlebury, J. Cortellazzi, and L. Caval- laro. Intriguing properties of adversarial ml attacks in the problem space. In2020 IEEE symposium on secu- rity and privacy (SP). IEEE, 2020

  100. [108]

    Praveena, A V ., P Chinnasamy, I

    V . Praveena, A V ., P Chinnasamy, I. Ali, R. Alroobaea, S. Y . Alyahyan, and Muhammad Ahsan Raza. Optimal deep reinforcement learning for intrusion detection in uavs.Computers, Materials & Continua, 2022

  101. [109]

    R. H. Randhawa, N. Aslam, M. Alauthman, M. Khalid, and H. Rafiq. Deep reinforcement learning based eva- sion generative adversarial network for botnet detec- tion.Future Generation Computer Systems, 2024

  102. [110]

    Rashid and J

    A. Rashid and J. Such. Malprotect: Stateful defense against adversarial query attacks in ml-based malware detection.IEEE Transactions on Information Forensics and Security, 2023

  103. [111]

    Rigaki and S

    M. Rigaki and S. Garcia. The power of meme: Ad- versarial malware creation with model-based reinforce- ment learning. InEuropean Symposium on Research in Computer Security. Springer, 2023

  104. [112]

    Romdhana, A

    A. Romdhana, A. Merlo, M. Ceccato, and P. Tonella. Deep reinforcement learning for black-box testing of android apps.ACM Transactions on Software Engineer- ing and Methodology (TOSEM), 2022

  105. [113]

    Schloegel, N

    M. Schloegel, N. Bars, N. Schiller, L. Bernhard, T. Scharnowski, A. Crump, A. Ale-Ebrahim, N.lai Bis- santz, M. Muench, and T. Holz. SoK: Prudent Evalu- ation Practices for Fuzzing. In2024 IEEE Symposium on Security and Privacy (SP), 2024

  106. [114]

    Schulman, S

    J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz. Trust region policy optimization. InInterna- tional conference on machine learning. PMLR, 2015

  107. [115]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017

  108. [116]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. CoRR, 2017

  109. [117]

    Sethi, R

    K. Sethi, R. Kumar, N. Prajapati, and P. Bera. Deep re- inforcement learning based intrusion detection system for cloud infrastructure. In2020 International Confer- ence on COMmunication Systems & NETworkS (COM- SNETS). IEEE, 2020

  110. [118]

    Sharma and M

    A. Sharma and M. Singh. Batch reinforcement learn- ing approach using recursive feature elimination for net- work intrusion detection.Engineering Applications of Artificial Intelligence, 2024

  111. [119]

    Shereen, D

    E. Shereen, D. Ristea, S. McFadden, B. Hasircioglu, V . Mavroudis, and C. Hicks. One pic is all it takes: Poi- soning visual document retrieval augmented generation with a single image.arXiv preprint arXiv:2504.02132, 2025

  112. [120]

    Simon, P

    R. Simon, P. Libin, and W. Mees. Learning Robust Penetration-Testing Policies under Partial Observabil- ity: A systematic evaluation, 2025. arXiv:2509.20008 [cs]

  113. [121]

    Skalse, N

    J. Skalse, N. Howe, D. Krasheninnikov, and D. Krueger. Defining and Characterizing Reward Gaming.Ad- vances in Neural Information Processing Systems, 2022

  114. [122]

    Sommer and V

    R. Sommer and V . Paxson. Outside the closed world: On using machine learning for network intrusion detec- 17 tion. In2010 IEEE symposium on security and privacy. IEEE, 2010

  115. [123]

    Standen, D

    M. Standen, D. Bowman, O. Naish, et al. Cy- ber operations research gym.https://github.com/ cage-challenge/CybORG, 2022

  116. [124]

    Su, Hong-Ning Dai, L

    J. Su, Hong-Ning Dai, L. Zhao, Z. Zheng, and X. Luo. Effectively generating vulnerable transaction sequences in smart contracts with reinforcement learning-guided fuzzing. InProceedings of the 37th IEEE/ACM Interna- tional Conference on Automated Software Engineering, 2022

  117. [125]

    J. Sun, T. Zhang, X. Xie, L. Ma, Y . Zheng, K. Chen, and Y .g Liu. Stealthy and efficient adversarial attacks against deep reinforcement learning. InProceedings of the AAAI conference on artificial intelligence, 2020

  118. [126]

    Richard S Sutton and Andrew G Barto.Reinforcement learning: An introduction. 2018

  119. [127]

    Cyberbat- tlesim.https://github.com/microsoft/ cyberbattlesim, 2021

    Microsoft Defender Research Team. Cyberbat- tlesim.https://github.com/microsoft/ cyberbattlesim, 2021. Created by Christian Seifert, Michael Betser, William Blum, James Bono, Kate Farris, Emily Goren, Justin Grana, Kristian Holsheimer, Brandon Marken, Joshua Neil, Nicole Nicho...

  120. [128]

    Terranova, A

    F. Terranova, A. Lahmadi, and I. Chrisment. Leverag- ing deep reinforcement learning for cyber-attack paths prediction: Formulation, generalization, and evaluation. InProceedings of the 27th International Symposium on Research in Attacks, Intrusions and Defenses, 2024

  121. [129]

    Tharewal, M

    S. Tharewal, M. W. Ashfaque, S. S. Banu, P. Uma, S. M. Hassen, and M. Shabaz. Intrusion detection system for industrial internet of things based on deep reinforce- ment learning.Wireless Communications and Mobile Computing, 2022

  122. [130]

    L. Tong, A. Laszka, C. Yan, N. Zhang, and Y . V orob- eychik. Finding needles in a moving haystack: Priori- tizing alerts with adversarial reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intel- ligence, 2020

  123. [131]

    Tsingenopoulos, J

    I. Tsingenopoulos, J. Cortellazzi, Branislav Bosansk `y, S. Aonzo, D. Preuveneers, W. Joosen, F. Pierazzi, and L. Cavallaro. How to train your antivirus: Rl-based hardening through the problem space. InProceedings of the 27th International Symposium on Research in At- tacks, I...

  124. [132]

    Tsingenopoulos, D

    I. Tsingenopoulos, D. Preuveneers, L. Desmet, and W. Joosen. Captcha me if you can: Imitation games with reinforcement learning. In2022 IEEE 7th Euro- pean Symposium on Security and Privacy (EuroS&P). IEEE, 2022

  125. [133]

    Tsingenopoulos, V

    I. Tsingenopoulos, V . Rimmer, D. Preuveneers, F. Pier- azzi, L. Cavallaro, and W. Joosen. The adaptive arms race: Redefining robustness in ai security. Research in Attacks, Intrusions and Defenses (RAID), 2025

  126. [134]

    Van Hasselt, A

    H. Van Hasselt, A. Guez, and D. Silver. Deep reinforce- ment learning with double q-learning. InProceedings of the AAAI conference on artificial intelligence, 2016

  127. [135]

    S. Vyas, V . Mavroudis, and P. Burnap. Towards the deployment of realistic autonomous cyber network de- fence: A systematic review.ACM Computing Surveys, 2025

  128. [136]

    J. Wang, L. Qixu, W. Di, Y . Dong, and X. Cui. Crafting adversarial example to bypass flow-&ml-based botnet detector via rl. InProceedings of the 24th international symposium on research in attacks, intrusions and de- fenses, 2021

  129. [137]

    Z. Wang, H. He, Z. Wan, and Y . Sun. Coordinated topology attacks in smart grid using deep reinforcement learning.IEEE Transactions on Industrial Informatics, 2020

  130. [138]

    Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas. Dueling network architectures for deep reinforcement learning. InInternational conference on machine learning. PMLR, 2016

  131. [139]

    C. J. Watkins and P. Dayan. Q-learning.Machine learn- ing, 1992

  132. [140]

    D. Wu, B. Fang, J. Wang, Q. Liu, and X. Cui. Evad- ing machine learning botnet detection models via deep reinforcement learning. InICC 2019-2019 IEEE Inter- national Conference on Communications (ICC). IEEE, 2019

  133. [141]

    X. Wu, W. Guo, H. Wei, and X. Xing. Adversarial policy training against deep reinforcement learning. In 30th USENIX Security Symposium (USENIX Security 21), 2021

  134. [142]

    P. R. Wurman, S. Barrett, K. Kawamoto, J. Mac- Glashan, K. Subramanian, Thomas J. Walsh, R. Capo- bianco, A. Devlic, F. Eckert, F. Fuchs, L. Gilpin, P. Khandelwal, V . Kompella, H. Lin, P. MacAlpine, D. Oller, T. Seno, C. Sherstan, M. D. Thomure, H. Aghabozorgi, L. Barrett, R....

  135. [143]

    X. Xu, Y . Sun, and Z. Huang. Defending ddos attacks using hidden markov models and cooperative reinforce- ment learning. InPacific-Asia Workshop on Intelligence and Security Informatics. Springer, 2007

  136. [144]

    C. Yang, A. Kortylewski, C. Xie, Y . Cao, and A. Yuille. Patchattack: A black-box texture-based attack with re- inforcement learning. InEuropean Conference on Com- puter Vision. Springer, 2020

  137. [145]

    Y . Yang, L. Chen, S. Liu, L. Wang, H. Fu, X. Liu, and Z. Chen. Behaviour-diverse automatic penetration test- ing: a coverage-based deep reinforcement learning ap- proach.Frontiers of Computer Science, 2025

  138. [146]

    S. Yu, R. Zhai, Y . Shen, G. Wu, H. Zhang, S. Yu, and S. Shen. Deep q-network-based open-set intrusion de- tection solution for industrial internet of things.IEEE Internet of Things Journal, 2023. 18

  139. [147]

    D. Zhan, W. Bai, X. Liu, Y . Hu, L. Zhang, S. Guo, and Z. Pan. Psp-mal: Evading malware detection via pri- oritized experience-based reinforcement learning with shapley prior. InProceedings of the 39th Annual Com- puter Security Applications Conference, 2023

  140. [148]

    Zhang, C

    T. Zhang, C. Xu, Y . Lian, H. Tian, J. Kang, X. Kuang, and D. Niyato. When moving target defense meets at- tack prediction in digital twins: A convolutional and hi- erarchical reinforcement learning approach.IEEE Jour- nal on Selected Areas in Communications, 2023

  141. [149]

    Zhang, D

    W. Zhang, D. Zhou, L. Li, and Q. Gu. Neural thompson sampling.arXiv preprint arXiv:2010.00827, 2020

  142. [150]

    K. Zhao, H. Zhou, Y . Zhu, X. Zhan, K. Zhou, J. Li, L. Yu, W. Yuan, and X. Luo. Structural attack against graph based android malware detection. InProceedings of the 2021 ACM SIGSAC conference on computer and communications security, 2021

  143. [151]

    W. Zhao, J. P. Queralta, and T. Westerlund. Sim-to-real transfer in deep reinforcement learning for robotics: a survey. In2020 IEEE Symposium Series on Computa- tional Intelligence (SSCI), 2020

  144. [152]

    D. Zhou, L. Li, and Q. Gu. Neural contextual bandits with ucb-based exploration. InInternational conference on machine learning. PMLR, 2020

  145. [153]

    A. Zou, Z. Wang, N. Carlini, M. Nasr, J Zico Kolter, and M. Fredrikson. Universal and transferable adversar- ial attacks on aligned language models.arXiv preprint arXiv:2307.15043, 2023. A Search & Review KeywordsWe defined the following keywords after reviewing the proceeding...

  146. [154]

    Reward FunctionSQIRLuses two rewards, an internal re- ward based on Random Network Distillation (RND), reward for finding new states [20]

    Sanitization Escape including obfuscation techniques such as capitalization, whitespace and SQL keyword encoding. Reward FunctionSQIRLuses two rewards, an internal re- ward based on Random Network Distillation (RND), reward for finding new states [20]. It also uses an external...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.