Pith. sign in

Paper Citation Record · LEDGER

An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.11814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11814 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:11.088564Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 666b8d88-e5e3-490d-a843-7faf46e44eff · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.088564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.088564Z digest=sha256:9430436a1c0824d4349ce3ef483877b66fc0ae58e6384b6ba1bd20eedce90e30

Observation 62362588-4ee9-4dd3-87b6-468428c4147d · inbound

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution cites this paper.

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:51.648562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:51.648562Z digest=sha256:5608a21d8e890b1008dd6df1c0638e5dce2e86e04edd5f0564b22148bd030693

Observation 531bb9b8-a3ff-4145-b1e8-b789c3bd6c80 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.704955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.704955Z digest=sha256:c47e02855b0f3dbeb598a7c7e03aafad6ac39daeccd350fdcc0d7f364ae48660

Observation a0e3fa75-76a2-41c3-a90e-714a8f0405ac · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.656665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.656665Z digest=sha256:ee7c628c9db511121272648a655deea080e1ee7b9dba183571cc8cae5d18b0df

Observation 63ea598c-b462-46c4-bd7a-a290cee2202a · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.935601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:d7357c875ef4ec35d3e534d348f305c77d2c7a52384e096a67c973bfc50b8878

Observation e6a33eb4-e0dc-45c9-b173-a5f249f0ccd2 · inbound

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments cites this paper.

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:35:44.724362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:32:57.161371Z digest=sha256:f47ac150df81ace069bfd7374376eba7de1b7a1e3f4e09e8fdf4af25b2fafaa3

Observation 4bb081e1-779b-4042-a3d7-129c3a192a6f · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.100715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:708ff370b1c9056ba9954da62f520e79a71e5dd25a894d17e546e6ba2413c851

Observation 6de744e9-7328-4157-b677-b69d9a320542 · inbound

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs cites this paper.

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:55:21.293998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:43:10.944954Z digest=sha256:923439b9ecf11518d99ab090e610cff19f0ae71402286c725ad38701f84cc15a

Observation 5a02c696-be0b-43d4-851f-134fe6f7b267 · inbound

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge cites this paper.

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:34:47.041453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:32:59.120345Z digest=sha256:99c304099e8123cb0c52c0339ebf609e9051dc6e8c0fcd1ea2df4e2f2a8defc6

Observation 11fde743-fb16-43d0-8e06-4a41b9ff1d5e · inbound

Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption cites this paper.

Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.078458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:01:15.586948Z digest=sha256:8649d631d4b07109a0730a83ce1ea2bb8a3683dad0d3c4a1b29ee5d50a8abca4

Observation 6f3341c2-a8fd-46a0-b129-cf9e57cf8280 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:ad7920fb4257c51fa6997aaf89128f62378a47755f0ef22f3dc4536b8f39d8c0