Pith. sign in

Paper Citation Record · LEDGER

An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2402.11814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11814 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:36:09.632242Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 50f8e8cb-e477-46cc-b372-20167b0f48ea · inbound

Psychometric-Based Evaluation for Theorem Proving with Large Language Models cites this paper.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.632242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.632242Z digest=sha256:061b93664273d618d9a57667d28030a2bbcd6fabde0f2c0709db678b53e5f243

Observation 666b8d88-e5e3-490d-a843-7faf46e44eff · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.088564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.088564Z digest=sha256:b37cb4d02a77148deb085566f6087e55c2a3757ac49de652301bfb35cb9467d5

Observation 62362588-4ee9-4dd3-87b6-468428c4147d · inbound

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution cites this paper.

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:51.648562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:51.648562Z digest=sha256:e223ea54c43bf1c1d7c742df14dab095b804c618c793f331f2573b9a9304649a

Observation 531bb9b8-a3ff-4145-b1e8-b789c3bd6c80 · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:05.704955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:05.704955Z digest=sha256:b59effd73551f8b186c09e8c2b0cd9784309405e0b86a10a1dec60fe4e147887

Observation a0e3fa75-76a2-41c3-a90e-714a8f0405ac · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.656665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.656665Z digest=sha256:778084d7be446972ab97a565e3dcc46944de5a8f7bcc76e9aa93ace1b9103bc5

Observation 63ea598c-b462-46c4-bd7a-a290cee2202a · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:52.935601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b975c274f3eed2d7bed307d5f4c4eb1df059cb3e4e57532f9eda38272095344a

Observation e6a33eb4-e0dc-45c9-b173-a5f249f0ccd2 · inbound

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments cites this paper.

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:35:44.724362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:32:57.161371Z digest=sha256:a4c60928c79bf83f1ae2d9b97a9d34693fd65dd767418fc1d3b213d6af8578dd

Observation 4bb081e1-779b-4042-a3d7-129c3a192a6f · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.100715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:0b98f7905f877f2894da15cab2df48913b280f0a32042adb4edcf1677782cc43

Observation 6de744e9-7328-4157-b677-b69d9a320542 · inbound

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs cites this paper.

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:55:21.293998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:43:10.944954Z digest=sha256:04eea49e64689d70d3c1b65eec057a8938bdc192c416dd6565b2cbc733fe7e5d

Observation 5a02c696-be0b-43d4-851f-134fe6f7b267 · inbound

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge cites this paper.

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:34:47.041453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:32:59.120345Z digest=sha256:4820be869205e4b22071a8184c49447731dfef48889d3939359fd0d8832b2481

Observation 11fde743-fb16-43d0-8e06-4a41b9ff1d5e · inbound

Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption cites this paper.

Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.078458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:01:15.586948Z digest=sha256:263e1862a90d5f2d887baea8f6d59a8b945c8b2ebab98954a67b2ce50ec6edbc

Observation 6f3341c2-a8fd-46a0-b129-cf9e57cf8280 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:b4d07745b565f56bd9d6998480d7e310f0774fc086840838514c22c57d9dd32d