REVIEW 15 cited by
AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have demonstrated impressive results on natural language tasks, and security researchers are beginning to employ them in both offensive and defensive systems. In cyber-security, there have been multiple research efforts that utilize LLMs focusing on the pre-breach stage of attacks like phishing and malware generation. However, so far there lacks a comprehensive study regarding whether LLM-based systems can be leveraged to simulate the post-breach stage of attacks that are typically human-operated, or "hands-on-keyboard" attacks, under various attack techniques and environments. As LLMs inevitably advance, they may be able to automate both the pre- and post-breach attack stages. This shift may transform organizational attacks from rare, expert-led events to frequent, automated operations requiring no expertise and executed at automation speed and scale. This risks fundamentally changing global computer security and correspondingly causing substantial economic impacts, and a goal of this work is to better understand these risks now so we can better prepare for these inevitable ever-more-capable LLMs on the horizon. On the immediate impact side, this research serves three purposes. First, an automated LLM-based, post-breach exploitation framework can help analysts quickly test and continually improve their organization's network security posture against previously unseen attacks. Second, an LLM-based penetration test system can extend the effectiveness of red teams with a limited number of human analysts. Finally, this research can help defensive systems and teams learn to detect novel attack behaviors preemptively before their use in the wild....
Forward citations
Cited by 15 Pith papers
-
PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages
PoCGen combines LLM-based exploit generation with static taint analysis and dynamic validation to produce proof-of-concept exploits for 77% of 560 npm vulnerabilities in the SecBench.js dataset.
-
Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study
The Aw–Rascle–Zhang continuum model on directed lattice networks yields power-law spatiotemporal congestion clusters with finite-size scaling by linear system size.
-
PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts
An agentic LLM framework turns natural-language vulnerability descriptions into executable Foundry proof-of-concept exploits, beating prompting and workflow baselines on 23 real-world smart contract cases.
-
Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms
A closed-loop LLM simulation with genetic algorithms suggests that phishing strategies can evolve to bypass simulated victims' defenses, but the result has not been validated against real humans.
-
From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection
SHIELD, an LLM-aided pipeline combining a masked autoencoder, deterministic data augmentation, and multi-level prompting, detects host-based attacks with high precision on three public datasets.
-
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.
-
Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research
In offensive-LLM agent papers, dual-use risk is acknowledged in 39% of papers but concrete mitigations appear in only 7%.
-
Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub
A single vulnerable Node.js path traversal pattern was found in 1,756 GitHub projects, most rated critical, and the authors' automated pipeline produced patches, disclosures, and evidence that LLMs have learned the pattern.
-
Jailbreak Attack Initializations as Extractors of Compliance Directions
The CRI framework pre-trains several jailbreak attack suffixes, then picks the one with lowest first-step loss for each prompt, improving ASR and cutting steps to success.
-
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
A multi-agent LLM framework autonomously discovers and exploits OWASP-mapped IoT vulnerabilities with 95% success across 260 trials in IoTGoat and Metasploitable2.
-
A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges
LLM pentest agents co-evolved through four bottleneck-driven phases into RLVR systems, while CTF platforms became dual evaluation/training infrastructure and three linked reliability gaps remain.
-
Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models
A locally hosted 8B language model closed an autonomous observe-decide-act attack loop against a vulnerable target but completed only 10.9% of tasks, showing architectural feasibility without operational reliability.
-
Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing
An LLM agent with planner and summarizer modules solved roughly a third of PicoCTF and OverTheWire CTF challenges, and the authors release the agent and benchmarks.
-
Generative AI for Internet of Things Security: Challenges and Opportunities
A survey that catalogs 33 GenAI-for-IoT-security works through the MITRE ICS mitigations lens, with three small case studies on adapting LLMs to IoT incident response and security question answering.
-
On the Surprising Efficacy of LLMs for Penetration-Testing
A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.
Discussion (0). Continue with ORCID to comment.