Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements

Isozaki, I · 2024 · arXiv 2410.17141

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

cs.CR · 2026-05-28 · unverdicted · novelty 7.0

Empirical study of 400 LLM attack runs finds exploitation success rates of 25-85% across four models against a fixed multi-service honeypot, with model-distinctive failure modes and p<0.001 differences.

citing papers explorer

Showing 1 of 1 citing paper.

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency cs.CR · 2026-05-28 · unverdicted · none · ref 9
Empirical study of 400 LLM attack runs finds exploitation success rates of 25-85% across four models against a fixed multi-service honeypot, with model-distinctive failure modes and p<0.001 differences.

Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements

fields

years

verdicts

representative citing papers

citing papers explorer