Pith. sign in

REVIEW 19 cited by

LLM Agents can Autonomously Hack Websites

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06664 v3 pith:TVCQ4CW3 submitted 2024-02-06 cs.CR cs.AI

classification cs.CRcs.AI
keywords agentsautonomouslycapablellmsmodelswebsitescallcapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, large language models (LLMs) have become increasingly capable and can now interact with tools (i.e., call functions), read documents, and recursively call themselves. As a result, these LLMs can now function autonomously as agents. With the rise in capabilities of these agents, recent work has speculated on how LLM agents would affect cybersecurity. However, not much is known about the offensive capabilities of LLM agents. In this work, we show that LLM agents can autonomously hack websites, performing tasks as complex as blind database schema extraction and SQL injections without human feedback. Importantly, the agent does not need to know the vulnerability beforehand. This capability is uniquely enabled by frontier models that are highly capable of tool use and leveraging extended context. Namely, we show that GPT-4 is capable of such hacks, but existing open-source models are not. Finally, we show that GPT-4 is capable of autonomously finding vulnerabilities in websites in the wild. Our findings raise questions about the widespread deployment of LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery

    cs.CR 2026-07 conditional novelty 8.0 of 10

    A replay-based verifier with environment isolation, role separation, and a browser-execution sentinel lets white-box LLM agents report XSS exploits that are real, reproducible, attacker-to-victim vulnerabilities.

  2. "Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents

    cs.CR 2026-08 conditional novelty 7.0 of 10

    Mobile GUI agents systematically over-grant Android permissions, and their allow/deny choices depend on the visible app name and the active task even when the permission request is identical.

  3. Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

    cs.CR 2026-08 conditional novelty 7.0 of 10

    Training a 7B LLM planner with reinforcement learning on verifiable attack rewards lets it generate policies that reduce DRL defender scores by an average of 522% versus static red agents.

  4. AXE: Grey-Box Exploitability Confirmation for Localized Vulnerability Reports

    cs.CR 2026-02 conditional novelty 7.0 of 10

    Grey-box metadata (CWE + code location) plus a multi-agent LLM workflow raises automated web-exploit confirmation from ~10% to 30% on CVE-Bench, with actionable PoC output.

  5. Agent Security Needs Redefinition through a Holistic Framework

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Agent security should be redefined around four contextual authorization properties instead of the content of the action performed.

  6. Scale-free congestion clusters in large-scale traffic networks: a continuum modeling study

    physics.soc-ph 2026-04 unverdicted novelty 6.0 of 10

    The Aw–Rascle–Zhang continuum model on directed lattice networks yields power-law spatiotemporal congestion clusters with finite-size scaling by linear system size.

  7. Relativistic Quantum Thermal Machine: Harnessing Relativistic Effects to Surpass Carnot Efficiency

    quant-ph 2025-08 unverdicted novelty 6.0 of 10

    Relativistic motion of the reservoirs in a three-level maser is claimed to yield a generalized Carnot bound that allows efficiency above the ordinary Carnot limit.

  8. Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A tool-augmented 8B LLM fine-tuned with GRPO on a new procedurally generated crypto CTF dataset reaches 0.88 Pass@8 on unseen easy tasks, up from 0.10 in the body's tables.

  9. Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub

    cs.CR 2025-05 conditional novelty 6.0 of 10

    A single vulnerable Node.js path traversal pattern was found in 1,756 GitHub projects, most rated critical, and the authors' automated pipeline produced patches, disclosures, and evidence that LLMs have learned the pattern.

  10. MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem

    cs.AI 2025-05 conditional novelty 6.0 of 10

    MM-Agent, a multi-stage LLM pipeline with a hierarchical modeling method library, is claimed to outperform prior agents and award-winning human solutions on a new 111-problem MCM/ICM-based mathematical modeling benchmark.

  11. A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges

    cs.SE 2026-07 accept novelty 5.5 of 10

    LLM pentest agents co-evolved through four bottleneck-driven phases into RLVR systems, while CTF platforms became dual evaluation/training infrastructure and three linked reliability gaps remain.

  12. Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models

    cs.CR 2026-08 conditional novelty 5.0 of 10

    A locally hosted 8B language model closed an autonomous observe-decide-act attack loop against a vulnerable target but completed only 10.9% of tasks, showing architectural feasibility without operational reliability.

  13. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A structured review organizes cyber-capable-agent risks into five vulnerability classes and argues that evaluation environments must be treated as operational security systems rather than background.

  14. LLM Agents Should Employ Security Principles

    cs.CR 2025-05 conditional novelty 5.0 of 10

    A position paper proposing AgentSandbox, a framework that applies Saltzer-Schroeder security principles to LLM agents and reports large attack-success-rate reductions on AgentDojo.

  15. A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A website can identify AI agents by their digital fingerprints and serve them a poisoned hidden version of the page, hijacking their actions via indirect prompt injection.

  16. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  17. A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A synthesis of established risk management practices into a structured framework for frontier AI developers, centered on explicit risk tolerance, KRI/KCI thresholds, and governance.

  18. From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs

    cs.CR 2025-06 conditional novelty 3.0 of 10

    LLMs can assist both attackers and defenders in cybersecurity, but context limits, hallucinations, and weak reasoning make them unsafe to deploy without human oversight and real-world evaluation.

  19. Prompt Injection 2.0: Hybrid AI Threats

    cs.CR 2025-07 reject novelty 2.0 of 10

    A structured taxonomy of hybrid prompt injection attacks shows how XSS, CSRF, and SQL injection vectors converge with LLM manipulation to bypass traditional controls.

Pith tools