Pith. sign in

Paper Citation Record · LEDGER

AutoPenBench: Benchmarking Generative Agents for Penetration Testing

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2410.03225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03225 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:28:58.473598Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:38:33.520795Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ee89ef1e-f6a7-4954-b922-bc437c6b9978 · inbound

VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework cites this paper.

VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:35.751742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:13:35.751742Z digest=sha256:3820f0e22461e84a0b51ddb2c601efb5ebe2f70cdfcafa17878a943e227d4a6c

Observation de35d2ad-a344-45fa-8198-e4ca84011687 · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:10.953295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:10.953295Z digest=sha256:03813e20d1d527a9f75684a056bedc0bb12bb6a6795f444b2724532d4e390bd9

Observation 4abde3c0-a8d4-4c57-a7a8-53308416d707 · inbound

A Contemporary Survey of Large Language Model Assisted Program Analysis cites this paper.

A Contemporary Survey of Large Language Model Assisted Program Analysis AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-09T05:32:24.938897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:32:24.938897Z digest=sha256:42bdd122cc662ba4f087e4f03333a89a133f64d7a92295e9f61b7c619434285a

Observation 76a70774-8182-42d8-9bdf-c1a675c03734 · inbound

Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks cites this paper.

Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:58.473598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:28:58.473598Z digest=sha256:e354127eb8f8c3f4fe9f18bed40636a53dad6523e5892457825dbf20a55e9312

Observation ff71ded2-2ba4-4491-bc5e-b6dc976527d6 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:15.439899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:15.439899Z digest=sha256:d48b557b9d3cfcc1fc03366d0b0e7f3d4274347c7ba36bab596909c6acf340e9

Observation 4d7d9d3d-5ce4-418b-a742-22dd271fc41d · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:04.455902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:04.455902Z digest=sha256:e3b2186c0b80a6760c604cd6fcba886e666391aa56a41e22b7794a4db026948d

Observation 27bacd1b-08dd-410b-a165-3744a2b61ff2 · inbound

Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities cites this paper.

Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:06:42.909055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T18:04:09.528381Z digest=sha256:36b4f41da383709d937b815cee4cd052e347c8c5283b7b5e81a0590148bfa0b6

Observation d76c0297-a11c-4260-9d24-07155462fd0b · inbound

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts cites this paper.

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:09:33.496022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:09:33.496022Z digest=sha256:998281c42b5a976e287d6ff8edb8b75c92043553ab8adb7377982182c200f961

Observation ccc1f7da-0c8b-46ed-82fe-805fde92620c · inbound

From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software cites this paper.

From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:03:22.151429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T20:02:47.746439Z digest=sha256:6596a2c71c59b8e61688d6590003ece05752ac65e37f058524148cc8f35878c4

Observation a0829395-8fbf-40f8-b2ff-564d4ff23ae3 · inbound

Autonomous Adversary: Red-Teaming in the age of LLM cites this paper.

Autonomous Adversary: Red-Teaming in the age of LLM AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:26:10.541154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T09:09:58.566171Z digest=sha256:c686feab65dfbf24a1d237bd9fa86596759080d94e58e201307dd9bed3f26d36

Observation 7f0784d0-b838-488b-92e8-25ab5b550474 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.683047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:e776866b83f43c63a99b7b703d25b96c5e7c8a07124e429875fdf8fc6fa09aba

Observation 9710babd-2c8b-4b92-8776-7d7167ff2f8c · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:24.431295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:24.431295Z digest=sha256:df27c57dc090ec2402f8100a8d93bb1d0c081c1fb64f8d37c4e5a409e9cf1e97

Observation 6b1e73b4-ee16-48fd-86f7-6f2e67d87f44 · inbound

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly cites this paper.

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:59.018250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T21:19:59.005348Z digest=sha256:c208a829b8e084a5fc025b291dfee25de509da6a2e084114a9ec3b19dec27521

Observation ba9e09b6-a8cc-4093-ad41-4631de23ceca · inbound

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency cites this paper.

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:13:30.524834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T06:53:48.841077Z digest=sha256:95bd623c595c96946af264576ac351be5cb11d2c2c6a68ec4f1049763c68e7d6

Observation e90b2f68-dd9b-43ea-924c-728df9a05bfa · inbound

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems cites this paper.

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:38:33.522197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T06:25:17.298114Z digest=sha256:bdbbeb74322973ba58773d2f92260dffa82591475cc5c6e4593e18b8e4a76726

Observation f5fb2b92-04b1-4909-b1c4-7251da9723c6 · inbound

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems cites this paper.

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems AutoPenBench: Benchmarking Generative Agents for Penetration Testing

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:38.131562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T11:27:22.436981Z digest=sha256:13954c633d6c1c07007e902f86d812320b8bdb0a10442595ab2e6055ac28c578