Pith. sign in

Paper Citation Record · LEDGER

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges

As of 10 August 2026, this Paper Citation Record lists 100 of 145 outbound references and 1 inbound Pith citation observation for arXiv:2607.02605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02605 v1

Coverage vector

measured 100 of 145 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:10:11.585499Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:19:18.909239Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T04:19:19.261715Z

Reference resolution

100 of 145 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 210fd98f-ca49-46b3-b294-6e11877e3ae1 · outbound

This paper cites Service grid fed- eration architecture for heterogeneous domains,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Service grid fed- eration architecture for heterogeneous domains,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:1e09480c073dc10eaed2d1439a72e462d4192cde4564200f908c5968dd2a2ca3

Observation aa5a10c5-7a25-474a-8a5b-7689c1cd5256 · outbound

This paper cites Adaptive energy-aware computation offloading for cloud of things systems,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Adaptive energy-aware computation offloading for cloud of things systems,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ab670782e3000c10c5a554e3f257fc512326098dc78b116db13ff08d614e629e

Observation 14f8630f-b46a-400b-b917-1aee3ec82fd1 · outbound

This paper cites Dynamic service invocation control in service composition environments,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Dynamic service invocation control in service composition environments,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:88227c3b77375e33fb7d8925d466fa369ae8df46aa17cbb696cc46118f43cbd9

Observation 746a63a0-06d5-47fe-88ce-73976ae9637c · outbound

This paper cites Fine-grained two- factor access control for web-based cloud computing services,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Fine-grained two- factor access control for web-based cloud computing services,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c716dc35d0ae9f811d39e1488a793d59345f2ccd881f36d355200c7b6d3ef5bf

Observation 6a265dc0-411b-472f-9b50-cae8eec7bc2b · outbound

This paper cites Trust-based access control for secure cloud computing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Trust-based access control for secure cloud computing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:efcc8bf7f041510c5ba1d7836a4094ea0b7a0dc470cc6232f4fa791718c2e3e2

Observation bd9b68e5-cc83-4c24-b65c-d0f83be8eeb5 · outbound

This paper cites Weidman,Penetration Testing: A Hands-On Introduction to Hack- ing.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Weidman,Penetration Testing: A Hands-On Introduction to Hack- ing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:765f8e59a425cbfab1f38bde0abef91dbd2d714e1180b7ee527850e5c811d2a5

Observation 16791e6a-4846-4c2f-b339-04e653c29a3d · outbound

This paper cites Security and privacy challenges in cloud computing environments,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Security and privacy challenges in cloud computing environments,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ebf8dbe39530f8d12030429c5f6baa549a5f331e0c3f5f09fcc6100f984ba663

Observation 64f4989a-9296-4306-b6d5-5a9ce36d0edf · outbound

This paper cites Dynamic security risk manage- ment using bayesian attack graphs,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Dynamic security risk manage- ment using bayesian attack graphs,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d83b7f821c3e467274afbdabdb169bf5925ea7ad97ec6105dcee88a5b096d4c3

Observation 8bf6ee13-283f-459f-9fac-80d159266c96 · outbound

This paper cites Two-factor data security protection mechanism for cloud storage system,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Two-factor data security protection mechanism for cloud storage system,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c39e82e1d9e213fc04da4939eecf5b22e84a29d5c0f24d4d318289d2006cb658

Observation 3b3c516e-d9da-4166-b135-c23ac125a6db · outbound

This paper cites Engebretson,The basics of hacking and penetration testing: ethical hacking and penetration testing made easy.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Engebretson,The basics of hacking and penetration testing: ethical hacking and penetration testing made easy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:f6e74b32ee18ab3d4b671273d8963a51b91ff95fb23ca494d7cf7e8b03fa3401

Observation 75906ac1-11f8-48d5-8c00-97f805433ab6 · outbound

This paper cites Examining penetration tester behavior in the collegiate penetration testing competition,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Examining penetration tester behavior in the collegiate penetration testing competition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:1591c915be80701c7883c0fea9597bc8d65047dc5b879c710f9b48318354082c

Observation df531fa1-58ae-4872-b057-9d5fee0a8e34 · outbound

This paper cites Penetration testing–reconnaissance with nmap tool,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Penetration testing–reconnaissance with nmap tool,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:9f25e091753db913d7899cedb974b0d7ae1a0c5964931950d9b8aac859cddc9e

Observation bb35d593-0ab9-41c3-baab-0ef8822002f0 · outbound

This paper cites Toss a fault to your witcher: Applying grey-box coverage-guided mutational fuzzing to detect sql and command injection vulnerabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Toss a fault to your witcher: Applying grey-box coverage-guided mutational fuzzing to detect sql and command injection vulnerabilities,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c3d7be8b91faa80f7d67d689d37362fc88a25fbfa295ddaa9e08f41ddaed8180

Observation 8da05c22-b554-4a6f-b3e0-ef17449affc6 · outbound

This paper cites {FUGIO}: Automatic exploit generation for{PHP}object injection vulnerabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges {FUGIO}: Automatic exploit generation for{PHP}object injection vulnerabilities,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:fe33afb7c6d303c460ba2494640a407e226d2ae473cd3f8af121e738de20c66e

Observation 2ad030cb-5c8f-45df-ae90-fe6783ef50cf · outbound

This paper cites {ChainReactor}: Automated privilege es- calation chain discovery via{AI}planning,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges {ChainReactor}: Automated privilege es- calation chain discovery via{AI}planning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d2d979f7fdc597d880271e7aabf7d8e95135fbefea346f16118274caeec517db

Observation d748b416-21b0-4f2a-a84e-792d00440a53 · outbound

This paper cites A comprehensive detection method for the lateral movement stage of apt attacks,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges A comprehensive detection method for the lateral movement stage of apt attacks,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:e4ea56dd5c188beb2548c7978606e9268740b6a5221144dda5fd70428b2bbdd4

Observation 63d8d5b6-2503-425d-8047-0124bf642b1e · outbound

This paper cites A comprehensive overview of large language models,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges A comprehensive overview of large language models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d8f88bddcbf2ecadaa12ac9be068e8705861e2139e3606ce3dd4e69f53ebf6d0

Observation f667df0f-4b32-4971-97f7-54be8d515cba · outbound

This paper cites Large language models for cyber security: A systematic literature review,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Large language models for cyber security: A systematic literature review,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:503a22a7a05ac24c3df86b7b7098475099b8a0cb1009c890f5afad5c3bc11e53

Observation c743e59d-b0ec-478b-a0e3-156f12764254 · outbound

This paper cites {PentestGPT}: Evaluating and harnessing large language models for automated penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges {PentestGPT}: Evaluating and harnessing large language models for automated penetration testing,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:2d0f890140e590354076f05ac527c4e06ab14ba64b4982fb87bbc95f3d882339

Observation 05906e27-e947-4daf-a7a9-e86e6156c3fd · outbound

This paper cites Expel: Llm agents are experiential learners,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Expel: Llm agents are experiential learners,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:0c020bdfa151d4aa7106c84d0973ccfa940738758429e236b425b6383282300c

Observation c1c82a77-5e6e-4a14-96c4-9bf0c06497f2 · outbound

This paper cites Autotool: Efficient tool selection for large language model agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Autotool: Efficient tool selection for large language model agents,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:e63d256d4fc8a6cb8f4a3b329a0281a5c1b226d47bcfe3c43870edefa6ceb702

Observation 731d5387-d37a-41dc-a752-606544d8b4ae · outbound

This paper cites A survey on the feedback mechanism of llm- based ai agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges A survey on the feedback mechanism of llm- based ai agents,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b1eb06a99dc31bbfd52896212e282e292879411cac2584dd1222666f05127f56

Observation a7f54027-a8b6-4368-81df-6fe9368e9ef3 · outbound

This paper cites Spaiware: Uncovering a novel artificial intelligence attack vector through persistent memory in llm applications and agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Spaiware: Uncovering a novel artificial intelligence attack vector through persistent memory in llm applications and agents,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:570ba89d57619de3d975069022e00c9921ccbadba63be9e556a7c9898df7acef

Observation 85194bb2-bd3c-4ce1-89f6-ae447e0820fd · outbound

This paper cites Exe- cutable code actions elicit better llm agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Exe- cutable code actions elicit better llm agents,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:6009171f1ff459fc0db181d56ff3c4b2665485c9cab7c177a331acf09f6fc8ff

Observation 0d16d82f-9d02-4712-af8c-61a1619fc6e7 · outbound

This paper cites VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:45b0246c05ef151bf322e3533ff35d3a19b4e5964e74d4f54ed05ceb400de14c

Observation 2e3e7b8b-fb0b-4e6b-bf09-efde705691e2 · outbound

This paper cites Automated penetration testing with llm agents and classical planning,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Automated penetration testing with llm agents and classical planning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:546b233d429db3596162aea54dc35cbbb68e2ba95adc6c27549ce02ea0dcb39a

Observation af477695-0790-4068-bbbb-98c06d38af12 · outbound

This paper cites Pentest-r1: Towards autonomous penetration testing reasoning optimized via two-stage reinforcement learning,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pentest-r1: Towards autonomous penetration testing reasoning optimized via two-stage reinforcement learning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ec513a41e6015c3d927e13a77fa831202ea74ef1430e576efa79bfe876e5256f

Observation 02c34b39-37b5-4fb7-bfaf-ccfaf52fb370 · outbound

This paper cites Cyber- zero: Training cybersecurity agents without runtime,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Cyber- zero: Training cybersecurity agents without runtime,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:55b3c5385eb96c993176868312a48a78afd6f3d1643064f0aa8b6b91b46c0eab

Observation 5a93a239-8d0c-4381-ad78-5626f6cf1eed · outbound

This paper cites Cybench: A framework for evaluating cybersecurity capabilities and risks of language models,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Cybench: A framework for evaluating cybersecurity capabilities and risks of language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c42c553fcfa6f755580ac09729c4cc2de1c91e73ddf105a9c5f94fd8dc074302

Observation 5f3952f4-5d59-4ede-b3e7-05fa6bf75dd6 · outbound

This paper cites Autopenbench: A vulnerability testing benchmark for generative agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Autopenbench: A vulnerability testing benchmark for generative agents,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:29e390aa5ce0a95b710eb23681db80536ae5f43b65dff2bc357d35b3f36bf1ce

Observation c3992502-f160-4a23-9de0-28499f2165a9 · outbound

This paper cites Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Nyu ctf bench: A scalable open-source benchmark dataset for evaluating llms in offensive security,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:1238d7d5e2dbb1d5aa9ebc6133f7d434e0f1954cc5fb99a170e0a353189dfec7

Observation 080f2a37-45a0-44bf-860b-630447ce2d8d · outbound

This paper cites Towards effective offensive security llm agents: Hyperparameter tuning, llm as a judge, and a lightweight ctf benchmark,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Towards effective offensive security llm agents: Hyperparameter tuning, llm as a judge, and a lightweight ctf benchmark,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ec537c38086cd8908e727c697e79df1a743b52f192aad6efbeefbde07aeb174f

Observation 0d9976af-13a6-4820-ac4e-73807e93a6bc · outbound

This paper cites Enigma: Interactive tools substantially assist lm agents in finding security vulnerabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Enigma: Interactive tools substantially assist lm agents in finding security vulnerabilities,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:fa5b72f93588428b4179ec74b80a2aa33e883af0f23ca12c265374225c269e95

Observation 0e72e55a-846a-42f6-bff8-eb744e3cd7d1 · outbound

This paper cites Chimera: Harnessing multi- agent llms for automatic insider threat simulation,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Chimera: Harnessing multi- agent llms for automatic insider threat simulation,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d6784a727c9d81d6a923c7f416b530fec7e4b72e99941220286648260cf76394

Observation c351cd2a-c334-45e1-9ece-eb6d29f591d5 · outbound

This paper cites Awe: Adaptive agents for dynamic web penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Awe: Adaptive agents for dynamic web penetration testing,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:62c5bbc5cd1357e981e420ee0e27f713ffcea9da344b84ab3a100e2a7f8d3b1f

Observation d27b3552-3f0d-48be-862c-3c013b62a9e5 · outbound

This paper cites Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:80a66196747de947b8617d502ef11fb0d26238094a80432ba558df5e9b36e525

Observation 5371b2c9-2f19-4d1c-b521-5fb7f4011829 · outbound

This paper cites Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c8d3bf10f633a1361d3e4a223e773c78d943ec97ca9112a28cf1bcf161b22013

Observation a471ea7d-910a-43e2-ac91-5381519797e2 · outbound

This paper cites On the Surprising Efficacy of LLMs for Penetration-Testing.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges On the Surprising Efficacy of LLMs for Penetration-Testing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:fe87081baa089779eb9f11fb07cf0d8ee066a945f8e4a3340a9049364fff99b6

Observation 3100b41b-dfe7-489c-a06d-4ef7dcc5e6c7 · outbound

This paper cites A unified modeling framework for automated penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges A unified modeling framework for automated penetration testing,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:0e1b1c490f4aec1ae0fc1dc3040912276e2aa87a9b970713a9a39255c5343087

Observation 027175b6-5a64-49e2-8404-ea9e96fe8d24 · outbound

This paper cites Automated penetration testing: Formal- ization and realization,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Automated penetration testing: Formal- ization and realization,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d5119f13c644aaf38360082c1a9f6df43b87c935b772bc0cf99bc48b5007b76b

Observation db8f7259-7110-4b10-9818-1822490f4a19 · outbound

This paper cites {CTF}:{State-of-the-Art}and building the next generation,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges {CTF}:{State-of-the-Art}and building the next generation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:2edbe7343a7b5d7ae9d4e2e7fd735a6a78d6b2aefc28517363b7f28aaa2d70af

Observation b452e875-d806-452f-bec5-9f4c1e7a6a3e · outbound

This paper cites Cve-bench: A benchmark for ai agents’ ability to exploit real-worldweb application vulnerabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Cve-bench: A benchmark for ai agents’ ability to exploit real-worldweb application vulnerabilities,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d3f90396106fda8fab311eea76e34ee36263fac893c42fec27fe6060285d6edb

Observation c59add8c-1475-4f25-b80e-d9b977a154c6 · outbound

This paper cites Guidelines for performing systematic literature reviews in software engineering,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Guidelines for performing systematic literature reviews in software engineering,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:5dc1f20146bf2ea486f70381076e3151723cf9edcbb2cc5642956cd972c41a4b

Observation adfa7a0f-770b-4e09-af85-092056ae07bb · outbound

This paper cites Intelligent penetration testing through integrated knowledge graph and historical decision enhancement,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Intelligent penetration testing through integrated knowledge graph and historical decision enhancement,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:1c3f57040189c0379e2318724cd9994992a73cf5a271ca4240dc84e2e70e4ebe

Observation c2543969-2970-403d-b724-85015210bb5d · outbound

This paper cites Intercode: Stan- dardizing and benchmarking interactive coding with execution feed- back,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Intercode: Stan- dardizing and benchmarking interactive coding with execution feed- back,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:06201085cb5452c33dc2898bbdc184663ea8de6836258e1f4ba4cde75e73c65e

Observation 5e28fc9a-d73d-478b-a8d2-eb69709113e8 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:0aae9bb08f954d0c89c64ae3a57e88f5a8639a5b717ad043ad9bd0b893323602

Observation 8d49293a-1eb0-41b9-ada0-0ab3a11d48fb · outbound

This paper cites Got Root? A Linux Priv-Esc Benchmark.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Got Root? A Linux Priv-Esc Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d9bd71d64222179343e92973367dd282a65b124e0be544f6939aa8dc67dbcfcf

Observation 449cdf82-bf1b-40fa-b15d-776a14b71bb5 · outbound

This paper cites HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:73d6929e34c1241e30b3c452192ea7380a2d2b3ffe842623acc45c7a680c4385

Observation 6f3341c2-a8fd-46a0-b129-cf9e57cf8280 · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b4d07745b565f56bd9d6998480d7e310f0774fc086840838514c22c57d9dd32d

Observation 2b9ceb88-5530-4cbf-83ce-e80d7511db39 · outbound

This paper cites Catastrophic cyber capabilities benchmark (3cb): Robustly evaluating llm agent cyber offense capabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Catastrophic cyber capabilities benchmark (3cb): Robustly evaluating llm agent cyber offense capabilities,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:4146bb3ae0269923b3c5556aa96c31d9efb72c280877e1544fb9138221e2e3b5

Observation 4b997c96-2099-4abb-a3d1-0b4b3a37c1cf · outbound

This paper cites Towards automated penetration testing: Introducing llm benchmark, analysis, and improve- ments,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Towards automated penetration testing: Introducing llm benchmark, analysis, and improve- ments,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:4b76c8916b4119715c62b2a3f3c8b0cb645afc73cedb5ee6dc7b66888d9c3304

Observation 103f6b67-dd11-4fb7-9d37-93627dbd24c4 · outbound

This paper cites Pen- testeval: Benchmarking llm-based penetration testing with modular and stage-level design,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pen- testeval: Benchmarking llm-based penetration testing with modular and stage-level design,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:04abcbeebfb1dc356e00756365f40b06e54d64cf68bc4873719fd9740786991e

Observation 2be6268f-5191-41f5-8aa9-0bf34e110253 · outbound

This paper cites Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:90c7ab5fb44279e3628f81daa0c261e19ee056df29e568a96c1f85274865b627

Observation b7d26508-ee65-49ad-a541-effc16e84bde · outbound

This paper cites Hackworld: Evaluating computer-use agents on exploit- ing web application vulnerabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Hackworld: Evaluating computer-use agents on exploit- ing web application vulnerabilities,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:08a73009dce27d40c85ae1a3bc272a11f569cf5e09cc61055521a5bc0abf29ae

Observation 5a47ac48-02bf-4fbe-b5a7-465833cbf79e · outbound

This paper cites Pacebench: A framework for evaluating practical ai cyber-exploitation capabilities,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pacebench: A framework for evaluating practical ai cyber-exploitation capabilities,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:9824b3a4d14d9360df9dcd346d42a8b56d10b2ca42c0cd76ab1a56df3dc4bba8

Observation 29c5fa5b-1d1a-4ee2-a75e-c3f3368c8d93 · outbound

This paper cites CTFusion: A CTF-based Benchmark for LLM Agent Evaluation.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges CTFusion: A CTF-based Benchmark for LLM Agent Evaluation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:4b716901c916781a9e8b566a0f9184efc7376226f786037c6911f35966d88a5c

Observation b4732774-2ba4-49c4-8ca7-af1be4307ff3 · outbound

This paper cites How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:73594c701cc945f800891623b887e22ab6f19fc2c8de87ef3673a4aef7cae54d

Observation 2d5e2d96-e0eb-45df-9ae2-539a51d8200d · outbound

This paper cites CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:0ab61b05c1f1c21a845c873841de2104cce0b0c5810ffdf4a9288d2637358839

Observation 29e2d4b9-282d-48c0-bdcf-9bd3f2509b6f · outbound

This paper cites ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:634a7fe298dd3456ee86ede1a88a8cdff9ebf6f1cf5524c5f68de7293d6f203b

Observation 9a3eb5c2-626a-43b1-bcf9-1ea1b20f0200 · outbound

This paper cites Penheal: A two-stage llm framework for automated pentesting and optimal remediation,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Penheal: A two-stage llm framework for automated pentesting and optimal remediation,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:30be9a5074acb8b55b028e27a9630074e2779a63204372ae82220055602bfaa7

Observation 0f64a90a-baee-4c66-b62c-38515c759812 · outbound

This paper cites AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:86253a6213197678c6a5e8520f70b000bfff4433fd381dee911544e35e7bc04a

Observation fa369260-2f15-4ff5-a111-5fcccf3e9bec · outbound

This paper cites Pentest-ai, an llm-powered multi- agents framework for penetration testing automation leveraging mitre attack,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pentest-ai, an llm-powered multi- agents framework for penetration testing automation leveraging mitre attack,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:9151b64de6ba0adbbed8dd63bb49de9a7ca634bb47c630b061b819d010f63e5a

Observation f313c0ce-8f83-47a4-b297-9b721728f5cd · outbound

This paper cites Pentestagent: Incorporating llm agents to automated penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pentestagent: Incorporating llm agents to automated penetration testing,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:474138e5209acd0a8f52bee48aebbfc453b94b28660d7836ce867699d8e075ee

Observation f1d55eba-7b9d-4104-90df-da6d403b2c2f · outbound

This paper cites Autopentester: An llm agent-based framework for automated pentesting,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Autopentester: An llm agent-based framework for automated pentesting,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:0e83ff73b9ab4f656d7f89ac653f1b67c607935e4a556caea1d1b0a52cd0683c

Observation 0a3d033b-75d4-471a-87e1-827e61d1b54c · outbound

This paper cites PenTest++: Elevating Ethical Hacking with AI and Automation.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges PenTest++: Elevating Ethical Hacking with AI and Automation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ad2b361eb34cce93ece032edded63cdecbf850ca3f22eb4dfd9486158710b7c3

Observation fa266cf3-5f06-4f6a-ab9a-8ca169bd82c5 · outbound

This paper cites Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ae03bc8534279d87a4ba386066c7384ea9995d2699a646090f26952c5410a600

Observation f23085bf-13b3-4e39-bba4-d03b957eb832 · outbound

This paper cites Guided reasoning in llm-driven penetration testing using structured attack trees,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Guided reasoning in llm-driven penetration testing using structured attack trees,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:319c14b9fe47ba4bce3636e27c0161abd102b056dd994bbe8d793b9736dfe6e1

Observation 3635cbc8-0410-4ede-b126-1e066601d176 · outbound

This paper cites BreachSeek: A Multi-Agent Automated Penetration Tester.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges BreachSeek: A Multi-Agent Automated Penetration Tester

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:aaf8210f9c42a4d772030d83366e842b51bbe2b72fe3c76067d9ab3496963440

Observation 2848a064-1259-4149-a7da-c349c96efe7e · outbound

This paper cites Controller makes pentesting better: An improved multi-agent auto- mated penetration testing framework,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Controller makes pentesting better: An improved multi-agent auto- mated penetration testing framework,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b06c011b5654dc500969c6605c39114b20c54edcb268ec77ecf4d99b2cb3a19d

Observation eb4c7851-be73-4bfa-a7f7-a93bb8a540e5 · outbound

This paper cites Shell or nothing: Real-world benchmarks and memory- activated agents for automated penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Shell or nothing: Real-world benchmarks and memory- activated agents for automated penetration testing,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:6eaa2c52d304403584c8d9963829bf9b99cf2f70e1cbdae6cc71cb17396231c4

Observation a529eebf-bae3-4054-8b2d-1225ebe2467f · outbound

This paper cites Refpentester: A knowledge- informed self-reflective penetration testing framework based on large language models,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Refpentester: A knowledge- informed self-reflective penetration testing framework based on large language models,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:22bbc040ea0dcdccecf3925ff31abd405a901e7f4f1eba70c33d705fb2506c31

Observation 3d413c04-9332-4b6d-b902-ee2b6ccb284d · outbound

This paper cites Ptfusion: Llm- driven context-aware knowledge fusion for web penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Ptfusion: Llm- driven context-aware knowledge fusion for web penetration testing,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:198f29cda2f646e6ae493fea30035ca04f26b755a779596629650b49d6320e7b

Observation 98b41d4b-c539-4034-a614-b84d265cb790 · outbound

This paper cites Pentestmcp: Llm and mcp based multi-agent framework for automated penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pentestmcp: Llm and mcp based multi-agent framework for automated penetration testing,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:24a287e82e11805dcd3a5fe0df9afefda0a913045300c023dfaa7812c88f4029

Observation 5d800317-7377-490f-af9c-882f2e149a6d · outbound

This paper cites xoffense: An ai-driven autonomous penetration testing framework with offensive knowledge-enhanced llms and multi agent systems,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges xoffense: An ai-driven autonomous penetration testing framework with offensive knowledge-enhanced llms and multi agent systems,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b725275746c6f9bab13d387b7c52dbf304347029e6279d4702c17cd2dd50e606

Observation 78960ca3-6a94-40bb-b0aa-baf686c3e03b · outbound

This paper cites CAI: An Open, Bug Bounty-Ready Cybersecurity AI.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges CAI: An Open, Bug Bounty-Ready Cybersecurity AI

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:5145cb7cd05f53c0d71e5159de99cea227e66c3524e51ff48e71e162866323ca

Observation f0b5d879-7476-44fd-b3f0-434a6e37a026 · outbound

This paper cites Redteamllm: an agentic ai framework for offensive security,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Redteamllm: an agentic ai framework for offensive security,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:844bd61c2ceaeacdf796fd6f23cbc1f3680961a52ce91e99d70409e6c465ffe2

Observation 91414e80-c0a7-4176-b89a-748a6efd9007 · outbound

This paper cites Red-mirror: Agentic llm-based autonomous penetration testing with reflective verification and knowledge-augmented interac- tion,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Red-mirror: Agentic llm-based autonomous penetration testing with reflective verification and knowledge-augmented interac- tion,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:902cf4330d13bad93a2ad804e060755b3c77840e4d0b4105f8ecf1f7514a8c93

Observation 2c088a8b-63c1-4f2e-b69b-b2d3abe6d00a · outbound

This paper cites What makes a good llm agent for real-world penetration testing?.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges What makes a good llm agent for real-world penetration testing?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c4cfe1d007d1e7f97490375ae737b08b6bbfdb867b3c7e56328ad5ffc6659634

Observation 16f0dc1a-57cb-4c6e-8605-560147f5281e · outbound

This paper cites Incalmo: An autonomous llm-assisted system for red teaming multi- host networks,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Incalmo: An autonomous llm-assisted system for red teaming multi- host networks,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:7c31dc14c5bc37660cd9b6267d5fa39a3543b1c897868d3d456b8d0d06ee352a

Observation 2cf6539a-6cb5-4e4d-90f1-10965943e0b9 · outbound

This paper cites Cipher: Cybersecurity intelligent penetration- testing helper for ethical researcher,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Cipher: Cybersecurity intelligent penetration- testing helper for ethical researcher,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:820f3704fd24055960c163beebfbdf881e77a4b7edf7a1ce92864012718e9eae

Observation 31ecfd0c-d832-4929-8ca0-b4154b6c0e7d · outbound

This paper cites Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:44ad891e9b877088d5fd9c8781718b25f8f257197675afca9317a06ecbf3c1d6

Observation 10ff4b56-ac0d-4aa9-ab28-55a72fa44e00 · outbound

This paper cites From intent to invocation: A reasoning-first framework for natural language to pen- etration testing commands,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges From intent to invocation: A reasoning-first framework for natural language to pen- etration testing commands,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:f8928be379fad7a78862cd09d154bbcb32c991b301b16762fd77081301492377

Observation 70ff2b12-b7b1-4391-adcf-f750b30954a5 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges LLM Agents can Autonomously Hack Websites

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:55b35983b696b993c97e4aef975627871cbc8061b7e8f37be2fe78b930fdd547

Observation 2f0d08fa-93a9-4122-93f2-a8d10879c8ca · outbound

This paper cites AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:d8d8a1ddaf36f67b38e15984de5f988571b342d98c1a7854bbd433eee314c9c1

Observation 96aef585-e1f7-422e-b536-1e18a0396d6e · outbound

This paper cites Automated tactics planning for cyber attack and defense based on large language model agents,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Automated tactics planning for cyber attack and defense based on large language model agents,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:a0256b299b7fb77b3d109d0e3361ba6b1080c66a10ac64c88df7bf9fa8696440

Observation 94154dbd-a531-48dc-a0f4-4b7c04a314ab · outbound

This paper cites From capabilities to performance: Evaluating key functional properties of llm architectures in penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges From capabilities to performance: Evaluating key functional properties of llm architectures in penetration testing,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:14286630cfe919afae3cf1c40e89c8e588d9ecf4b750cc2a37172df4f426b0a0

Observation e9a648b7-6a46-4593-b160-55fa59a4ce80 · outbound

This paper cites Penforge: On-the-fly expert agent construction for automated penetration testing,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Penforge: On-the-fly expert agent construction for automated penetration testing,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:bb300917eb9f059cf250a2269e84b583cc6f3aa57379583e1ccb7c7feca33878

Observation 17035994-6e59-4683-a7e6-65929d33f127 · outbound

This paper cites Perses: Unlocking privilege escalation for small llms via extensible heterogeneity,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Perses: Unlocking privilege escalation for small llms via extensible heterogeneity,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:bc28b23077256fa17fdee35f31bd920097ea0df57f19c9925a77668fd1b11176

Observation 83ab6bc7-cefc-42bf-8263-7aab9d5039d4 · outbound

This paper cites Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:cd17c78e3461c37861dc17a747db86dcbc575f5a0461f8d3dd3b701407ec8f93

Observation f1ba5da7-160d-4928-85cb-99ecc5e676dc · outbound

This paper cites Can llms hack enterprise networks? autonomous assumed breach penetration-testing active directory networks,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Can llms hack enterprise networks? autonomous assumed breach penetration-testing active directory networks,

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:552a14f5fb1da014c8e56a757e4e483ea46ce90da535fadb94e54530007eb896

Observation 71464fe0-cc6d-41cf-9872-bf253de0b665 · outbound

This paper cites Multi-Agent Penetration Testing AI for the Web.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Multi-Agent Penetration Testing AI for the Web

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:673a43023c6b9e1c199aa5310d3bfe865aa1ac1d8ac68d79d8075d740e13dd41

Observation 5b87ce12-3fcd-4f1c-9c6a-d3c90ea08304 · outbound

This paper cites ARACNE: An LLM-Based Autonomous Shell Pentesting Agent.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges ARACNE: An LLM-Based Autonomous Shell Pentesting Agent

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b945362075cde08ecfc4e86ff853fa83454bd1cce2b1686a3ba06354622ab328

Observation 1bdbd712-9394-4429-aeb9-f7544ec9b9e3 · outbound

This paper cites Wifipentester: Towards governed genai-assisted wireless pentesting,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Wifipentester: Towards governed genai-assisted wireless pentesting,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:4f481701e02a6c2114ac561181f99baf6fa54e7ed930c4dde9280c92d02ef6c2

Observation c4d1d695-363c-439a-bd24-acedbda5d58b · outbound

This paper cites Ai-driven penetration testing for arm systems: Experimental evaluation and deployment framework across four paradigms,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Ai-driven penetration testing for arm systems: Experimental evaluation and deployment framework across four paradigms,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:8004f33b6366a7bcb95037728db1368c2a5caa409500e88f4079bd9a201588d0

Observation d9070392-1f83-4ca8-ba8d-897911781525 · outbound

This paper cites Getting pwn’d by ai: Penetration testing with large language models,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Getting pwn’d by ai: Penetration testing with large language models,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:8aa0c8196b861a4cc92a15e5947390097d53205a04a0a0e5dacd9181f094ae24

Observation d0a70912-4a6c-422a-8326-482f6fa986c6 · outbound

This paper cites AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:f1cbe180616ce62a1b57f82c23aa7d367714d15e6d2a31222a0d0902d71efa25

Observation 960b9023-4926-4141-9fbf-a86ff7787b6b · outbound

This paper cites Llms as hackers: Autonomous linux privilege escalation attacks,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Llms as hackers: Autonomous linux privilege escalation attacks,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:bfb8c9e5ac846c1395bc8bb4de0504ba83582d752dcb77b33f34f8cff95a1139

Observation 85a0061c-b2d1-4985-87a7-df0adff7c83d · outbound

This paper cites D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:7ab6445d13c6e0f2d4eea8fbb80ba35b3f349368354330a08929bc1dc6315715

Observation c949ec0a-e045-4586-928f-35b22e340274 · outbound

This paper cites Measuring and augmenting large language models for solving capture-the-flag challenges,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Measuring and augmenting large language models for solving capture-the-flag challenges,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b682907c9f5707cd217653017ea63df32f6b60c738bdfbd685aa28aec9ee9933

Observation 3602dbb9-74c9-42a7-85ee-00ed9d5398a4 · outbound

This paper cites Ctfagent: An llm-powered agent for ctf challenge solving,.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Ctfagent: An llm-powered agent for ctf challenge solving,

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:faeda425d6d454e5295a25a2ee4ad3a1ac95e297b96b1d156045a1d95ca245bd

Pith citing papers

Observation 4ae7e36f-590f-4925-8495-66c7521dbca7 · inbound

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models cites this paper.

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T04:19:19.265911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T04:19:18.909239Z digest=sha256:9c6a67206601f84e911f9452a9b0f383d2880fb9e0e7886df0df096b77196f8a