Pith. sign in

Paper Citation Record · LEDGER

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

As of 14 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 1 inbound Pith citation observation for arXiv:2604.05719.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.05719 v1

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:52:57.225878Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:10:11.585499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 138 outbound references displayed

  • verified exact21
  • verified fuzzy74
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b511081-5ba5-4848-a39b-b08ac41d67ef · outbound

This paper cites https://openstd.samr.gov.cn/bzgk/gb/newGbInfo?hcno=BAFB47E8874764186BD B7865E8344DAF.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing https://openstd.samr.gov.cn/bzgk/gb/newGbInfo?hcno=BAFB47E8874764186BD B7865E8344DAF

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.922338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d873c8306d798bd1ff956af00e6163b0c1ec650ef36b510bab1283f75ccad430

Observation ced6f320-887e-49de-88a4-d2b7c49063ba · outbound

This paper cites HexStrike AI MCP Agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing HexStrike AI MCP Agents

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.938328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:91ef6a76308ea5315e449d394550f21a47fe27761bfa2ad257586c6b0349211f

Observation 08a19b75-ef73-4418-9b25-e93e37403ef3 · outbound

This paper cites EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing EnIGMA: Interactive tools substantially assist LM agents in finding security vulnerabilities

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.909426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e5c66ab1423595ae73415fb20e037121c2db8db14f8fe85b6e1dc5671cc6cf71

Observation e54c609c-f1cd-4db5-b1e1-22c4981be066 · outbound

This paper cites Metasploit penetration testing cookbook.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Metasploit penetration testing cookbook

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.917910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d0dba72ec899d5e415fc6d071930e548e72209511cf31de915f139a7eab9df7d

Observation 7ad02f1d-8813-47c1-b803-218f81685de5 · outbound

This paper cites BreachSeek: A Multi-Agent Automated Penetration Tester.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing BreachSeek: A Multi-Agent Automated Penetration Tester

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.827605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b7961741cea4d6eaf17965ecb7f47c9c04f98b4bcae21f3a434027318c99a88a

Observation 1a4a38eb-8d14-4cf3-9467-0e9a30320173 · outbound

This paper cites Introducing the model context protocol.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Introducing the model context protocol

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.924718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:7fc7c7f70739c9af2ae8f237ab73eccb4032df132cdf989157d4e7d2ca3b9b5e

Observation e1eaed53-6bff-43da-8d0d-a69d00170f57 · outbound

This paper cites Agent skills.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Agent skills

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.913453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f6e8227af3896954287ee892d34d35f494265eb7bac17cd0e9eb663a20406031

Observation 4681805f-5208-430b-814e-4b77b21deb7c · outbound

This paper cites Claude code.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Claude code

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.907004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b8c1a04e004af96ba243f3beded5e4229302b0e23510ecf6de6e3889b191a023

Observation e2ce3f8d-c201-4d37-9e4a-8cd6f3af2364 · outbound

This paper cites Claude opus 4.6 system card.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Claude opus 4.6 system card

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.936155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:93b220a749dd01232dd19e15ee8241cb01833af70f2fa04cbbace70c0c88a186

Observation 4829b7bb-a499-47b1-95b3-d4329b765ac4 · outbound

This paper cites Introducing Claude Opus 4.6.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Introducing Claude Opus 4.6

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.919830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:dfade0010346da9f958e9758dc304e49861270d164df2cdfab4b1c9535680ca8

Observation df9bb2c7-d8f5-4857-b04a-cf3c90dc6b4d · outbound

This paper cites Software penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Software penetration testing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.911391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ba75758c2b6c5f3e9432af8e749afd0626b1455ac7a21be35ae2e5aca9b91cc4

Observation 9bfcb707-cc4c-42c6-9ec4-fa9ae02499ee · outbound

This paper cites Pentest-ai, an llm-powered multi-agents framework for penetration testing automation leveraging mitre attack.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Pentest-ai, an llm-powered multi-agents framework for penetration testing automation leveraging mitre attack

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.927615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2ed6adcba22e04ade89e37bf32137cd74590868b8607aed4e03ddc94426d5ada

Observation 4fe81752-659b-4afa-a2ec-b09ee134a740 · outbound

This paper cites About penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing About penetration testing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.933689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d658a22502dcf572bc3b052de0836cccb291f75a12f13f7450c966b4803bbbad

Observation 83c95459-01d8-42fd-87d1-08833cb7faba · outbound

This paper cites Coverage-based greybox fuzzing as markov chain.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Coverage-based greybox fuzzing as markov chain

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.915871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e15bf5cc394f10a5eba97d868799cd48c743397d56ea80e1a5281257af03974a

Observation 8b19fa8d-b63d-4098-80d7-651dfd3b9094 · outbound

This paper cites Language models are few-shot learners.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Language models are few-shot learners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.930889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:256d88b6a9678be6b722b5d7a79245587a235eeaea97665a1ee8c64f9225ae83

Observation 735af62d-a28b-4649-b203-b74839e71d3c · outbound

This paper cites The diamond model of intrusion analysis.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing The diamond model of intrusion analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.798392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:5c3a9acee7a4da916bc036ea427a99438f77e16ae7b006e506c179f4c7c71b22

Observation 25b37231-0e44-4535-95a4-1520382eaaf7 · outbound

This paper cites an unresolved cited work.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:22:55.796193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:aa67104aa3fafb276c7b04c91c8eea071f00393e02295f01dcc29233c8ab8005

Observation b3f955ad-ecab-4226-bf03-73196752c8a2 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Why Do Multi-Agent LLM Systems Fail?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:59.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d4b72c64a2d67bd0ca08f1f41fa8b22bea4d4ebc592eba78cc87b2d0be17b0a0

Observation 08536ac8-5ba6-4e02-b6d5-655d96d0404f · outbound

This paper cites tinyctfer.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing tinyctfer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.697301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a77bcd03f9f79386376fb86df86cf3b4a4ecb224960b75aadb9259964c28ea40

Observation 2ed5c236-14f4-491c-92de-16f4d088dbd8 · outbound

This paper cites RedTeamLLM: an Agentic AI framework for offensive security.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing RedTeamLLM: an Agentic AI framework for offensive security

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.800364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2c6d53cebd41f89cfa6c5bcfe4456bcbc1f27d2f20f7e33ff150dcf0b4b6169e

Observation 05c53256-eccc-4286-b99a-46af9ad72b53 · outbound

This paper cites Under the hoodie: Lessons from a season of penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Under the hoodie: Lessons from a season of penetration testing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.949648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ad8444fc12323848a83100bef200573659760812a4d2b095ee1b52f8f719d842

Observation d1856f37-ba71-46fa-905a-ba4c8f9b76db · outbound

This paper cites The growing importance of exposure management: Key insights from gartner hype cycle for security operations 2024.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing The growing importance of exposure management: Key insights from gartner hype cycle for security operations 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.826185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:45d5915c18331db14711567bb89939a8edf137b0b7f825e32295d4eaec3c6599

Observation 8d8c929c-c461-4124-ba8f-f18b68a554fa · outbound

This paper cites crewAI: Fast and Flexible Multi-Agent Automation Framework.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing crewAI: Fast and Flexible Multi-Agent Automation Framework

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.738107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2e272d068d730222b209ff845b6ed02ca679d2557f60b6a80ba1332b3f07e9a4

Observation 15a98dff-b576-4f2a-a0fc-27c0a41c9b0a · outbound

This paper cites CVE: Common Vulnerabilities and Exposures.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing CVE: Common Vulnerabilities and Exposures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.685544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:1c0231640d6b7ea8b404731ebf8682867b18a62936d21bd983ba5d99496c9544

Observation afc8c333-0720-4114-b781-cef46f630004 · outbound

This paper cites RefPentester: A Knowledge-Informed Self-Reflective Penetration Testing Framework Based on Large Language Models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing RefPentester: A Knowledge-Informed Self-Reflective Penetration Testing Framework Based on Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.952640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e48fd15b659cd8c3423d7bb9aa3b99748710a3329229012dd7aca6dc616eb191

Observation 74fb8b87-e886-448f-9553-ff1b42f0ea4b · outbound

This paper cites Multi-Agent Penetration Testing AI for the Web.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Multi-Agent Penetration Testing AI for the Web

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.773454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2656be8599d049b7eb086ca91ccca3b0730ddfc45e74d7f9fb677614e5c5819a

Observation 477558b5-eb75-406b-a2f9-3e4d0b4a1788 · outbound

This paper cites What makes a good llm agent for real-world penetration testing?.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing What makes a good llm agent for real-world penetration testing?

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.735714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:79e8f8721d1163024d50d282d5863d2ebc7989d4b9468ac227b0c203510301f3

Observation aa96cb8d-3d2e-49c9-a22f-3e5c82a909f5 · outbound

This paper cites {PentestGPT}: Evaluating and harnessing large language models for automated penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing {PentestGPT}: Evaluating and harnessing large language models for automated penetration testing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.972537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:18b92f1e53f72b51545005cac99c587ba64634b0be440376bd8602f6562021d2

Observation e551424a-534e-4971-8508-8b1dc06f0240 · outbound

This paper cites Cyberstrikeai.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cyberstrikeai

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.976586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e865336e37c6d25f2038b400696a8019aeb9916575a48d7b9603457095819f9c

Observation dd13d838-f950-491a-a37e-b9cbcda988bb · outbound

This paper cites Regulation (EU) 2022/2554 of the European Parliament and of the Council of 14 December 2022 on digital op- erational resilience for the financial sector (DORA).

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Regulation (EU) 2022/2554 of the European Parliament and of the Council of 14 December 2022 on digital op- erational resilience for the financial sector (DORA)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.987071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:89214959c72b9573891d866813967705bf98e17bbc20a84fa722d537cf43890d

Observation 73a5fa88-342c-4570-8797-d4143e7031c8 · outbound

This paper cites A survey on rag meeting llms: Towards retrieval-augmented large language models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing A survey on rag meeting llms: Towards retrieval-augmented large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.791451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:590bef6963bfdc42c98d63a2904c0a377eb4e343f5fbd9c973b8ef2a2bfb4507

Observation dfd86db7-4ebb-4620-bfac-b64a76345362 · outbound

This paper cites Llm agents can au- tonomously exploit one-day vulnerabilities.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Llm agents can au- tonomously exploit one-day vulnerabilities

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.768208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b05b19f47f961e2f86320bf412622e1145e1c53c936fb2900be5e648c9b138ef

Observation c72fffc0-1a6a-4728-99d2-a6fbfe733fe8 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing LLM Agents can Autonomously Hack Websites

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.756965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:559dbecc08068a9b5bf88d467b0649fd2183cf9696462d1ff20d1edadc343945

Observation 01e02962-7938-4873-94e6-5f261f9116f5 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.764061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:7b4a914bff0fd8eae3c219bfe372b786f6e2764005d67d8f3ce02814a01125b8

Observation aedab074-ba1f-4542-84ef-888942865865 · outbound

This paper cites PentestAgent.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing PentestAgent

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.802284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:419b5d6b4a16169b9fdd9d633838fd9011626a7a30f2a0b7e95285f0bdb23a57

Observation 1f6bd849-b4b0-455a-a271-d1221823af45 · outbound

This paper cites Automated Planning: theory and practice.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Automated Planning: theory and practice

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.954806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:3edf576bf32eebc699e6d67fff1817805518661df0a54f3c2573f4f224d63aba

Observation c080e9c1-1285-44e7-878a-27c9c2143f21 · outbound

This paper cites Autopenbench: A vulnerability testing benchmark for generative agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Autopenbench: A vulnerability testing benchmark for generative agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.844123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b38af8200a3d39d84d59f73b01045689c9f088e9538e66162434f1577eec9d27

Observation 1a128853-efde-4765-9d42-4daebe9fc288 · outbound

This paper cites Gemini 3.1 Pro.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Gemini 3.1 Pro

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.821382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fe11b0d64588d2e1ed757f1eec1ac14511aa1fa6645fcabf80769f4d23597054

Observation 3f683134-c97b-45d8-a750-4aea0f6df451 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.823878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:393c765cdd1f03681cab4eaba75a71ea9ef12e3bd19dbe75442419d60c93fffe

Observation 82cea609-b96b-43d8-90c0-f3f9a415a8f2 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.795747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:cfbfcdae0e547d002926e3719899e9604cbfb4f661727354e568d7783ea1f596

Observation 9ddd11c2-4c53-4ed2-bd86-41bbf13fb3e2 · outbound

This paper cites Hack The Box.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Hack The Box

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.773399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:10427085a3a0f9d4cec24baaaf5be29e9ae638f67f6f0cf515befd38cd8feaaa

Observation 307fbf41-9b28-488e-aa08-ce594834e3ab · outbound

This paper cites Hacking articles.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Hacking articles

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.833284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:cff73c3f898fdc35b4710b896ef305a2733581c218bb67b34126f1c4add13269

Observation a95bcb4f-082c-48fb-bdbd-0e7706c7a9dd · outbound

This paper cites Getting pwn nd by ai: Penetration testing with large language models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Getting pwn nd by ai: Penetration testing with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.730109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2deb3792b18dcb3d8ccfc975ca6ec5d2d4271cb4438906a37cdfba10bfc69f6a

Observation 92e71700-932e-425a-a430-4441a80ceeb9 · outbound

This paper cites Can llms hack enterprise networks? autonomous assumed breach penetration-testing active directory networks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Can llms hack enterprise networks? autonomous assumed breach penetration-testing active directory networks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.807046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:13dc6414946be44efbe1dab57d9ce67d3b9ecaa7b6ccaf4bf562dd55c810e384

Observation 5ff18c7b-91bd-477f-95d3-153df0404d8b · outbound

This paper cites On the Surprising Efficacy of LLMs for Penetration-Testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing On the Surprising Efficacy of LLMs for Penetration-Testing

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.913350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:83844cd2c887d07162aee0b80a5c291fb0166fbde4b69aff3b4943fe9e62a6fb

Observation 5b685132-b264-4697-aaca-25fb7aaf15d9 · outbound

This paper cites Got root? a linux priv-esc benchmark.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Got root? a linux priv-esc benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.707813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:90de64f547c0ef36dea090c8236fc3efed276c24ded4a137f57a1b7853c59b9a

Observation d5d0d05b-ae42-4395-a5f5-d7bb50cf264e · outbound

This paper cites Llms as hackers: Autonomous linux privilege escalation attacks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Llms as hackers: Autonomous linux privilege escalation attacks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.792002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2e4cf979ee8ce1e3b67e586df4e34e6e7837ca9a029815a3c3842ba76ae98b21

Observation 25d6001a-2db5-4982-b399-c6a12613c1d4 · outbound

This paper cites AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.782770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2809137a39e98b4f0378cf388ad031eb1a529f82a111d3d450e1553deebd6b4a

Observation 1b17aab1-a4cb-4474-9fc9-922a03da6a85 · outbound

This paper cites H-pentest.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing H-pentest

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.719105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:94068aca26b7565d9d09ca89e79d0abe97150efea658c5e699145d66dcade84d

Observation 4f22fead-67db-4381-816f-5244dd31e525 · outbound

This paper cites Context rot: How increasing input tokens impacts llm performance.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Context rot: How increasing input tokens impacts llm performance

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.761579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:c02ff4376e65dfc34a1d6a015710f6908b6999c9e1750830800543cc30f778c8

Observation c14b06d1-a65f-4136-865d-b0b58db96bdf · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Metagpt: Meta programming for a multi-agent collaborative framework

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.778255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:bad10af53367891cba2fc92fbd874b9730361be498320e7ebca57240b1cade76

Observation c1bb90ce-f5d7-4194-9440-4192aaf72ef6 · outbound

This paper cites Penheal: A two-stage llm framework for automated pentesting and optimal remediation.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Penheal: A two-stage llm framework for automated pentesting and optimal remediation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.754598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:5f8f0549de8d5429d6e2de9cd2c48cb4aff5fd0bf18829dae68f139b2497d4f2

Observation 9132d79d-3323-4586-8ef8-e99c740d6cda · outbound

This paper cites Qwen2.5-Coder Technical Report.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Qwen2.5-Coder Technical Report

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.748765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2411a7f633b87d19ece108bac3358abd1d89ea27950849faa5659372be482ce3

Observation 7d3879fa-7c87-4ac0-a4f4-cafb7eced539 · outbound

This paper cites newmapta.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing newmapta

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.745400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:32b3296a9052b193b56a84df0ac9b8c5d305be61b7ac5b2a0f97bb54c2270ead

Observation 7e2a4e60-cd4a-4499-a760-aba0b7f599d0 · outbound

This paper cites Intelligence-driven computer network defense informed by analysis of adversary campaigns and intrusion kill chains.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Intelligence-driven computer network defense informed by analysis of adversary campaigns and intrusion kill chains

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.964831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:071c39a4833439c4ffe6881f891e54ccac8f73f0dae3687b594845d51bb02a05

Observation 15635540-4f50-4d4b-925f-f73cdba19d3c · outbound

This paper cites Towards automated penetration testing: Introducing llm benchmark, analysis, and improvements.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Towards automated penetration testing: Introducing llm benchmark, analysis, and improvements

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.732953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:27b2623dfd1e7a4a60c74d452cf0957df51c97f3f06c06cb521c4d0e803d65f6

Observation 08b243d6-5102-4a75-983d-fb9500e9f5ba · outbound

This paper cites Measuring and augmenting large language models for solving capture-the-flag chal- lenges.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Measuring and augmenting large language models for solving capture-the-flag chal- lenges

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.961495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ee09ed0e830e3863919a773dae29d190f849cad6c46b51b6fe86565763c75d16

Observation 55e1c6d5-d05d-4cf4-97e7-4a77f4799ef2 · outbound

This paper cites Survey of hallucination in natural language generation.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Survey of hallucination in natural language generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.966988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a33e085654904a38059b3d0879a5850aba1a9059963563d7d3a7e12e4e930f36

Observation 313fe888-8bfe-48f9-8ec9-c9b48df9767b · outbound

This paper cites Sok: Agentic skills – beyond tool use in llm agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Sok: Agentic skills – beyond tool use in llm agents

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.978600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:43ae763ebb7b10d7427969b40bc9d99a81bcf4f91a0921a02b14c841c10d1fe6

Observation 90b6b10a-f210-45e5-9d39-4d0474bd60c3 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.695416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e997fd541d9d2254f2b7bae8c372b3eb1f2592cafa0165c774dcf094852a3fad

Observation 07ea5548-35e0-4710-a5d6-012c57fef10c · outbound

This paper cites Dense passage retrieval for open-domain question answering.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Dense passage retrieval for open-domain question answering

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.699805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:7de9ece3f13755dee7c156ef138cfe8de88db9abdc17519e554a061925507522

Observation e5edc5e0-73cd-408d-b6e8-1282c5c8cd02 · outbound

This paper cites Metasploit: the penetration tester’s guide.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Metasploit: the penetration tester’s guide

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.782399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a21f1b46560cb8b73841a24e7b264e0424b959e69ddf44538ca821216405e714

Observation 337c991f-c664-4ff4-9f66-92a4f5d068fe · outbound

This paper cites arXiv preprint arXiv:2508.07382 , year=.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing arXiv preprint arXiv:2508.07382 , year=

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.744985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:054ac1f25ce3df07860701b1bdab9f67f6ebb755f1578bb1275ce49638552ff7

Observation a600fa08-1553-40b0-99bb-001ea4d61aca · outbound

This paper cites VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.753000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e3c655f10db383901d906f6ab93651b08b700767f822352932a65d1be94185bc

Observation fd9323a4-041b-4fb8-aca6-5935541fc4be · outbound

This paper cites Se perspective on llms: Biases in code generation, code interpretability, and code security risks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Se perspective on llms: Biases in code generation, code interpretability, and code security risks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.693120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:16f67737fb9790daa9e58aab8425f983708bbcd101cabb953fdb0b87329ce2a9

Observation 8578de78-3a28-44e2-9518-44dedc88a023 · outbound

This paper cites LangGraph: Low-level orchestration framework for building stateful agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing LangGraph: Low-level orchestration framework for building stateful agents

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.688368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:0dcc5b7946d3af5e7c462902f1c3c11a70d9a334505ed051046cdf71a5a630aa

Observation 560a5d76-53a6-4569-8903-84934ca8e720 · outbound

This paper cites Lost in the middle: How language models use long contexts.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Lost in the middle: How language models use long contexts

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.752509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:0a883c045b7f1b3b4efeef19c4d37c0b91829ebe135a5b67bddf7150b9dc383c

Observation 808051ee-2187-47b0-8dc5-6209b888f61d · outbound

This paper cites Pacebench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities.ArXiv, abs/2510.11688, oct 2025.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Pacebench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities.ArXiv, abs/2510.11688, oct 2025

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.760810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:574b52c42ebf70cb3a5c1276e417440758613c12660fc9a6a6b277d0dd0aca37

Observation 4c7ec2b4-fed1-4d34-897f-83a7f644b1a0 · outbound

This paper cites Yu, and Ming Zhang.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Yu, and Ming Zhang

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.816490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f30e978f0452b2519e810337ad96c6ba73c487b3c8f20a24dcf444eec225455e

Observation 24ee55ed-9801-40dd-9d49-21e2b1e09b42 · outbound

This paper cites xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.866163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:95b3de6096f160ce40e605b82df94c122ed0662ecf42f37252bf8162b5d9f803

Observation 57a3d93d-51c1-4900-ae5c-a214e0bb8ce6 · outbound

This paper cites xbow-competition.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing xbow-competition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.980347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fca9d52cee5ae5e56f7d1511cbdfa48830676ccd07bfa59fd29bdbb033e8a4a8

Observation 81de04e4-3411-4450-a2aa-2967e39c4ee8 · outbound

This paper cites Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.816819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:90fb5cc8b09535802d092d4682b01846538bf6e3f08dc65e63162b7291ec9499

Observation 29fbf94f-0ec5-4fba-a4a5-c4cb4be3e5da · outbound

This paper cites Graphical user interfaces.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Graphical user interfaces

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.994363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:6fec31ec2e93c5865407cd42ef958d7a739c11b5015f6fffe0f93e12a7654bc7

Observation ae4ade74-0f14-461f-a995-cdf7a66b5d9b · outbound

This paper cites CAI: An Open, Bug Bounty-Ready Cybersecurity AI.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing CAI: An Open, Bug Bounty-Ready Cybersecurity AI

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.777811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f5996c4b6f557bf73ba9448be9aa3a1aef1bf8ab71722f90eecdfe0f029d6648

Observation f2586f11-f75b-468b-bedc-0b8860a1a230 · outbound

This paper cites an unresolved cited work.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:22:55.839497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:65b216d941beb086bf80931abaaf84550a8f9c72abcd2c41767f400d6d2de8b6

Observation 4f7150bb-3455-4d69-b156-1cf7f8d49865 · outbound

This paper cites CWE-Common Weakness Enumeration.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing CWE-Common Weakness Enumeration

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.800349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:68f8e5a13280d809f3ecc72c5a5df8342e05c240372d340463f8331d43b56fe5

Observation 2fa01074-e596-408d-ac40-5a8d047eb89f · outbound

This paper cites Kimi Code CLI.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Kimi Code CLI

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.784554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fb6b3016f63f8b91647e4ca16ed21ff18f062f5602ff0d3924460ded008593e1

Observation 11e04272-a99f-4219-9bce-73a6cd349f42 · outbound

This paper cites Penetration testing and ethical hacking services market size & share analysis - growth trends and forecast (2025 - 2030).

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Penetration testing and ethical hacking services market size & share analysis - growth trends and forecast (2025 - 2030)

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.992249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:875efcf9edc6cb5b579821b4ab540c9a4c6af9ebf9f35c453049985054c885ba

Observation ac72b7d5-c36c-4076-9c24-d44a01e37ff5 · outbound

This paper cites HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.946097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d8b3fafeee15b11d4951a0af58090c4f11626d8986c2eb5530bf6c810136893a

Observation 43bd97bf-c2b2-4b8b-b7aa-1cef2b5ea14e · outbound

This paper cites Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.889783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:e5042345e93214ae7b75da572580f8e8f39e54e3aafcaf40a16ffbeea642ba8d

Observation 3437b781-ae95-4d11-8d11-4210f02e680c · outbound

This paper cites ARACNE: An LLM-Based Autonomous Shell Pentesting Agent.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing ARACNE: An LLM-Based Autonomous Shell Pentesting Agent

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.854441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:78372b142f0e7b35dbbdbc6d76768d4c6b934c64b192b11e46155585a7ee0160

Observation e80334b4-8736-4575-9b26-0192498cdf6b · outbound

This paper cites Passage Re-ranking with BERT.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Passage Re-ranking with BERT

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:38:57.456949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b3223cdf7d8ba360d093ea94be1e3b01e7dd3c234f560852de4aca6d28d1e923

Observation 56d2614a-dcc8-407e-ab80-2f59b5e508ec · outbound

This paper cites Introducing GPT-5.2.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Introducing GPT-5.2

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.787016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:152aac23ce5e36b08468b3c0101019266ecb042dc0b4d768e64691078e98c442

Observation 1524d535-a30b-447a-87d7-98b7d19fda3e · outbound

This paper cites GOAD (Game Of Active Directory).

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing GOAD (Game Of Active Directory)

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.957288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:eb2e75476d61878cd3afa29642617bd7f480a352b34959a44818c867b4ded5ca

Observation c986c91f-130c-4387-a2a0-cfe3665b7409 · outbound

This paper cites OverTheWire: Wargames.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing OverTheWire: Wargames

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.709904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:2ebda2dd31a01b81439eac9e449dc1fc41e9025e5de5db58a4a3c1e9c38f8f4d

Observation 3efd8319-81cb-43ad-a734-2f54b633d769 · outbound

This paper cites OW ASP Top Ten Web Application Security Risks.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing OW ASP Top Ten Web Application Security Risks

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.742928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fda178ccaa4043f83d7b6d90dc94c937cedefd0365b0a71b5fb0e729452eb793

Observation ba9215ea-bc59-48bc-8229-268b351e6e72 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Generative agents: Interactive simulacra of human behavior

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.702770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:cd090b29a3b69d31c2244ed543dbbb84f67d61383a73c9b7348344f8a26c9f1a

Observation 0505bcc1-83c5-4228-a86b-9c338242f49a · outbound

This paper cites ctfsolver.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing ctfsolver

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.722597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:63acbb5118ce1301aa856877b0e6317834e8f184975477dcbfb39b06d731de09

Observation 2f7d2d1d-0e8f-42f9-868f-edfa3db28ddd · outbound

This paper cites PCI Security Standards Council.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing PCI Security Standards Council

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.793706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:f03da3c7d481593c4f3c002aaa2c3a5def57f9ca791d431f0eeaba4a8c0d7005

Observation 249cb4d2-f8cc-48ec-8b86-ecc9739939ce · outbound

This paper cites The penetration testing execution standard.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing The penetration testing execution standard

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.947122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:c42b5a9427dcd8e7978696cd4afe288f83d8f6b2b4e3a8a792528825b85dfc05

Observation be7e0ebc-081a-40d5-baa1-b2d94f33344c · outbound

This paper cites Chatdev: Communicative agents for software development.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Chatdev: Communicative agents for software development

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.819287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:620d279a42a31ffc35a0b1c03db2a7cf1ec1321934ac636250c559724e93197e

Observation 019c8e3d-e4ca-4352-a12e-3f939e6593a5 · outbound

This paper cites Hackworld: Evaluating computer- use agents on exploiting web application vulnerabilities.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Hackworld: Evaluating computer- use agents on exploiting web application vulnerabilities

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.982490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:acff3702ecbdca0dfb63e3ccf933daafcce54567ab62b7915cc3448fdf7cf98c

Observation 517fae7d-fe0c-448b-88eb-8562dc1316c5 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Code Llama: Open Foundation Models for Code

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:52.921947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:88088aaf582d1dd6e042344ef3b69e028bcb64025fe21f40b3f2960bc3157c10

Observation 289c2aae-390f-4b8c-8afc-ac6b2e6201ca · outbound

This paper cites Luan1aoagent.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Luan1aoagent

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.846624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:098aa829c2b0ff1d240ad9c38f9ee805232ef7c46916c746332a8d6dad0f8f51

Observation 075b3a53-3cb6-4ed9-903b-08d7d659b540 · outbound

This paper cites Techni- cal guide to information security testing and assessment.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Techni- cal guide to information security testing and assessment

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.812007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:747614ea634020ae81e9772892e9ee57752cb7ff102d1e42d2498bdc5b7625b9

Observation 3005aae5-69d8-4203-82f4-f5c97b708553 · outbound

This paper cites Toolformer: Lan- guage models can teach themselves to use tools.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Toolformer: Lan- guage models can teach themselves to use tools

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.809835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ce06c69a0f05874b18649bb6724f6285937bfc2f78f82f18b167df80ed5b8dd5

Observation 63ea598c-b462-46c4-bd7a-a290cee2202a · outbound

This paper cites An Empirical Evaluation of LLMs for Solving Offensive Security Challenges.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.935601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:0a901dee089a136cd5f629868ce046b14e32daa41de3529b76a14855c4c74f3c

Observation 36921e58-8a6b-468f-b8ba-acbcc7d8984c · outbound

This paper cites Nyu ctf bench: A scalable open-source benchmark dataset for evalu- ating llms in offensive security.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Nyu ctf bench: A scalable open-source benchmark dataset for evalu- ating llms in offensive security

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:17:55.968879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:a3386cc7fe1903c0866ad1cf306273a07e7dbdd9d550ef4f9482b2ccc6cb545a

Observation edde6d80-bd73-46d5-9e26-f33de730e60f · outbound

This paper cites Pentestagent: Incorporating llm agents to automated penetration testing.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Pentestagent: Incorporating llm agents to automated penetration testing

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.747708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:22f031a133576ae310fa3851fbe3689036a4f67195bc826d2d5490c34cc483dc

Observation 839dae70-06ec-4174-9fdb-6537cdba8dfc · outbound

This paper cites Llms in software security: A survey of vulnerability detection techniques and insights.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Llms in software security: A survey of vulnerability detection techniques and insights

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:22:55.716041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:ae9629973be21d50b2769a7468020f09be1abe13b956971861047b233c481ffc

Pith citing papers

Observation d27b3552-3f0d-48be-862c-3c013b62a9e5 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:c5dd042ea47d41d7d01c1b8305ffa2052c726c6706da596ca2b45ac385d6d8b0