Pith. sign in

REVIEW 15 cited by

AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01038 v1 pith:RMEJCKMK submitted 2024-03-02 cs.CR cs.AI

classification cs.CRcs.AI
keywords attacksllmsattacklanguagellm-basedpost-breachresearchsecurity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have demonstrated impressive results on natural language tasks, and security researchers are beginning to employ them in both offensive and defensive systems. In cyber-security, there have been multiple research efforts that utilize LLMs focusing on the pre-breach stage of attacks like phishing and malware generation. However, so far there lacks a comprehensive study regarding whether LLM-based systems can be leveraged to simulate the post-breach stage of attacks that are typically human-operated, or "hands-on-keyboard" attacks, under various attack techniques and environments. As LLMs inevitably advance, they may be able to automate both the pre- and post-breach attack stages. This shift may transform organizational attacks from rare, expert-led events to frequent, automated operations requiring no expertise and executed at automation speed and scale. This risks fundamentally changing global computer security and correspondingly causing substantial economic impacts, and a goal of this work is to better understand these risks now so we can better prepare for these inevitable ever-more-capable LLMs on the horizon. On the immediate impact side, this research serves three purposes. First, an automated LLM-based, post-breach exploitation framework can help analysts quickly test and continually improve their organization's network security posture against previously unseen attacks. Second, an LLM-based penetration test system can extend the effectiveness of red teams with a limited number of human analysts. Finally, this research can help defensive systems and teams learn to detect novel attack behaviors preemptively before their use in the wild....

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages

    cs.CR 2025-06 conditional novelty 7.0 of 10

    PoCGen combines LLM-based exploit generation with static taint analysis and dynamic validation to produce proof-of-concept exploits for 77% of 560 npm vulnerabilities in the SecBench.js dataset.

  2. Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study

    physics.soc-ph 2026-04 unverdicted novelty 6.0 of 10

    The Aw–Rascle–Zhang continuum model on directed lattice networks yields power-law spatiotemporal congestion clusters with finite-size scaling by linear system size.

  3. PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts

    cs.CR 2025-11 conditional novelty 6.0 of 10

    An agentic LLM framework turns natural-language vulnerability descriptions into executable Foundry proof-of-concept exploits, beating prompting and workflow baselines on 23 real-world smart contract cases.

  4. Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms

    cs.CR 2025-07 conditional novelty 6.0 of 10

    A closed-loop LLM simulation with genetic algorithms suggests that phishing strategies can evolve to bypass simulated victims' defenses, but the result has not been validated against real humans.

  5. From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection

    cs.CR 2025-07 conditional novelty 6.0 of 10

    SHIELD, an LLM-aided pipeline combining a masked autoencoder, deterministic data augmentation, and multi-level prompting, detects host-based attacks with high precision on three public datasets.

  6. MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

    cs.CR 2025-07 conditional novelty 6.0 of 10

    MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.

  7. Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research

    cs.CR 2025-06 reject novelty 6.0 of 10

    In offensive-LLM agent papers, dual-use risk is acknowledged in 39% of papers but concrete mitigations appear in only 7%.

  8. Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A single vulnerable Node.js path traversal pattern was found in 1,756 GitHub projects, most rated critical, and the authors' automated pipeline produced patches, disclosures, and evidence that LLMs have learned the pattern.

  9. Jailbreak Attack Initializations as Extractors of Compliance Directions

    cs.CR 2025-02 conditional novelty 6.0 of 10

    The CRI framework pre-trains several jailbreak attack suffixes, then picks the one with lowest first-step loss for each prompt, improving ASR and cutting steps to success.

  10. VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

    cs.CR 2026-07 conditional novelty 5.5 of 10

    A multi-agent LLM framework autonomously discovers and exploits OWASP-mapped IoT vulnerabilities with 95% success across 260 trials in IoTGoat and Metasploitable2.

  11. A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges

    cs.SE 2026-07 accept novelty 5.5 of 10

    LLM pentest agents co-evolved through four bottleneck-driven phases into RLVR systems, while CTF platforms became dual evaluation/training infrastructure and three linked reliability gaps remain.

  12. Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models

    cs.CR 2026-08 conditional novelty 5.0 of 10

    A locally hosted 8B language model closed an autonomous observe-decide-act attack loop against a vulnerable target but completed only 10.9% of tasks, showing architectural feasibility without operational reliability.

  13. Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing

    cs.CR 2025-07 conditional novelty 4.0 of 10

    An LLM agent with planner and summarizer modules solved roughly a third of PicoCTF and OverTheWire CTF challenges, and the authors release the agent and benchmarks.

  14. Generative AI for Internet of Things Security: Challenges and Opportunities

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey that catalogs 33 GenAI-for-IoT-security works through the MITRE ICS mitigations lens, with three small case studies on adapting LLMs to IoT incident response and security question answering.

  15. On the Surprising Efficacy of LLMs for Penetration-Testing

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.

Pith tools