REVIEW 5 cited by
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Hacking poses a significant threat to cybersecurity, inflicting billions of dollars in damages annually. To mitigate these risks, ethical hacking, or penetration testing, is employed to identify vulnerabilities in systems and networks. Recent advancements in large language models (LLMs) have shown potential across various domains, including cybersecurity. However, there is currently no comprehensive, open, automated, end-to-end penetration testing benchmark to drive progress and evaluate the capabilities of these models in security contexts. This paper introduces a novel open benchmark for LLM-based automated penetration testing, addressing this critical gap. We first evaluate the performance of LLMs, including GPT-4o and LLama 3.1-405B, using the state-of-the-art PentestGPT tool. Our findings reveal that while LLama 3.1 demonstrates an edge over GPT-4o, both models currently fall short of performing end-to-end penetration testing even with some minimal human assistance. Next, we advance the state-of-the-art and present ablation studies that provide insights into improving the PentestGPT tool. Our research illuminates the challenges LLMs face in each aspect of Pentesting, e.g. enumeration, exploitation, and privilege escalation. This work contributes to the growing body of knowledge on AI-assisted cybersecurity and lays the foundation for future research in automated penetration testing using large language models.
Forward citations
Cited by 5 Pith papers
-
Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research
In offensive-LLM agent papers, dual-use risk is acknowledged in 39% of papers but concrete mitigations appear in only 7%.
-
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks
An autonomous LLM-driven agent can compromise accounts in a realistic Active Directory testbed, with reasoning models outperforming non-reasoning ones at competitive cost.
-
VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework
VulnBot, a three-role LLM agent team with a penetration task graph and summarizer, raises penetration testing completion rates over raw GPT-4o and Llama3.1 on public benchmarks, with one end-to-end real-machine succes...
-
Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks
A review of LLM-based agents as autonomous cyberattackers, arguing that they lower attack costs, scale up threats, and outpace existing defenses.
-
RedTeamLLM: an Agentic AI framework for offensive security
The paper reports that adding a separate reasoning step to a terminal-operating LLM agent reduces tool calls and improves completion on 4 of 5 entry-level CTF virtual machines, while the framework's memory and plan-co...
Discussion (0). Continue with ORCID to comment.