Pith. sign in

Paper Citation Record · LEDGER

HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2412.01778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01778 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:11.031542Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4cc83794-0f5c-44dd-90b2-ac04f2fcd385 · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.031542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.031542Z digest=sha256:9228360f7cf124884b49ec5b132d991f2f81c192c661794c2faba1e139bf13ac

Observation b8433a1f-db6f-44a2-973e-4a657cf4068e · inbound

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges cites this paper.

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:35.401593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:35.401593Z digest=sha256:35367cdd22b71fec37b3465b50b54cbebdbd0d2d511567af06498b513a632b3a

Observation 82562bdd-9d90-4023-96d6-b295d55b26e5 · inbound

Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities cites this paper.

Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 152

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:06:43.054517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:04:09.528381Z digest=sha256:8b68f5c81912845653d31550d5666c5132c235c7d23892430ff661ebcffd8ffd

Observation 21886b58-514e-4675-abba-88549a69c205 · inbound

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models cites this paper.

xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T16:41:38.195679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T16:38:45.171505Z digest=sha256:5a4d1653be8da50825400a5ef7f747dc565000f710e49b67ba15d549bdc1fb61

Observation 56861ea8-a6ca-447d-b970-86941d6a8b75 · inbound

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting cites this paper.

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:58.660386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:58.660386Z digest=sha256:24c6fc4e30e3260e312591cab3eebacba95766a537abc8d634644d316496988e

Observation ac72b7d5-c36c-4076-9c24-d44a01e37ff5 · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.946097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:fea86b686c04fb1eec4b8422de0a8d8dd78e24932a6d8d8555b0d2934841a179

Observation 57ed953c-4fb8-4b5e-a148-623c862c0eb3 · inbound

Challenges and Future Directions in Agentic Reverse Engineering Systems cites this paper.

Challenges and Future Directions in Agentic Reverse Engineering Systems HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:50:24.760427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:49:31.238578Z digest=sha256:cde6097577559775ad13fa9c295ab427efbd839bb59d791a8133e4af8cb7393b

Observation bd5fc79b-618b-4684-a6db-7ec21eaf40d0 · inbound

Dynamic Cyber Ranges cites this paper.

Dynamic Cyber Ranges HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:16:53.786335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T03:04:03.611481Z digest=sha256:2628c94a69b24a7deed9e82b43779c5a1960978209ec69a772a61b52d558da53

Observation 530fb5bd-c552-4ab8-a1b2-8554e677c23d · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.854253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:de879060fe8737c7248398e8eb47f17c6301e398a767d796e7b0addfcb0cad39

Observation 36436658-d8c2-4171-9c7a-4c4e88422b23 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:25.329346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:25.329346Z digest=sha256:8ace5a3569c5ec8ad8d8ce82449792fac4cfc80dc8fa564ed366e2869194b92f

Observation f31ef441-6cb3-43d6-9282-f1debc5704f4 · inbound

uGen: An Agentic Framework for Generating Microarchitectural Attack PoCs cites this paper.

uGen: An Agentic Framework for Generating Microarchitectural Attack PoCs HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T15:47:38.578630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T15:46:52.746700Z digest=sha256:b0db6ed3aadc05aa1a074534fd5ec57072c321d1fee29c06a98d8934a7cc3935

Observation c8a6b39a-832b-4fcb-8d5c-0f1a40ae56f3 · inbound

A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems cites this paper.

A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:28:25.609556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:23:47.507490Z digest=sha256:28cd82a51bd936ee3dc5ecf36b9bce5050e4c3de4dd2ffd5b271331621b41b7e

Observation 44723b6e-f54f-4eee-b76d-253797ee3450 · inbound

Benchmarking Mythos-Linked Bug Rediscovery cites this paper.

Benchmarking Mythos-Linked Bug Rediscovery HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:17:57.557350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T23:14:34.164854Z digest=sha256:4b1e49831443c629e1e55c726d487fcd9444cfa92c6c0d5c60ba2f32d5162e8e

Observation 9f7241b1-0b0c-4fd1-b6fa-3b38cff13600 · inbound

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks cites this paper.

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:30:20.160101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T04:27:57.881555Z digest=sha256:6ae6bcfeab6b4781c867ba8578f42d0bf06de22451ddf07a243013bf29c05bdb

Observation f9d42804-c1c6-428b-97c0-2e7bf82976ea · inbound

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks cites this paper.

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:35:12.940275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T16:25:47.480189Z digest=sha256:61a75a679e8ac26b67b567a5f1f862efe69849d1f2aaaa00a72a680fdccd93a9

Observation 9b10175f-e341-42eb-80a9-80c48ae0abeb · inbound

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency cites this paper.

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:13:30.546421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T06:53:48.841077Z digest=sha256:6a7fa3eb1516f2f9f4c4f51f70c6f9771120826c8b29bbf67facfb487ada46a2

Observation 5f5bcf9d-bbb0-4aad-815d-702d144743fb · inbound

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense cites this paper.

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.773069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:26:22.434043Z digest=sha256:de48cc22ad6a40534615d8491c7a9152873a6b7b7131b92a64c51782e1662854

Observation 2e29f34e-fe48-4e76-913e-532b949c1c5b · inbound

Synthetic APTs: the Collapse of TTP-Based Attribution cites this paper.

Synthetic APTs: the Collapse of TTP-Based Attribution HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.317063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:02:35.881369Z digest=sha256:2fffa627d03e7b230c344c3bd1e2188a28fb94145b1acc8b5fc69fd0265e9697

Observation 097a62f0-a635-456e-a5f3-91e6553d8c96 · inbound

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation cites this paper.

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:10:57.027838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T12:55:22.831264Z digest=sha256:3edb40f2be918e63a8e6ceb9e641ffa77cc8c67769afe6b07d4795642c19aa4e

Observation 449cdf82-bf1b-40fa-b15d-776a14b71bb5 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:9953476d694db5d35f10cdb8b78f2250cbe91c56397dd3deef9550b4eea22207

Observation 24b59bcc-8ffd-4dfc-b0c8-06c750c80ac2 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T02:40:29.385913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:40:29.385913Z digest=sha256:04638f9c4a8f86d9c13c0eed49f40b3f46bc40d978ac4911277545a1402339a5

Observation 5d41388c-3db1-4e58-bb07-d8b86d680402 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:20.837262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:20.837262Z digest=sha256:c496a20e033083286a92f4975f5f4a3ee32617d808d24b9c5b22fa86ebd5d4e1

Observation 9198dcf8-bd12-4b6a-a7ec-0677a79acbe1 · inbound

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents cites this paper.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:36.891218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:36.891218Z digest=sha256:55919f9bf9a2496ec5f7953e474ab880dda9f4a98409ca27324bc17b22f5d4bc