Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2504.10112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10112 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:10.965917Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:07.161491Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 36a33f6c-8109-4d23-bf1d-1af92c0df62d · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:10.965917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:10.965917Z digest=sha256:aca9f737eb1d02580850011cac44a73f575704797532e56cc56c9799619258be

Observation 39cd706b-239e-4186-b0b0-bbd991284f8a · inbound

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs cites this paper.

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:59.924739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:59.924739Z digest=sha256:29998f71f470063a7f17e02b036287244e022c387034de1f170691c2fc3e1aee

Observation bed81608-05d1-4c89-a81a-27894b52f624 · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:04.491227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:04.491227Z digest=sha256:84595745d5323a6e9d76ae138f3d6860dc18f8cf4cb1be0a3bd00597bcf53dee

Observation a7a0094c-8ea6-4482-9e2b-b5cff3cee040 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:27.628828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:9a8f54718128bf90c6973a9e7f5ec24a6efb59bef0422fa5cc3d6b663236d8bf

Observation f27747f2-bda1-48df-a2ac-0167ef37d704 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:24.494937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:24.494937Z digest=sha256:f660ead6c0a0223a00bcf953e1bf75c9739e01fd6648251e689524e6112a6436

Observation a7931abb-6c44-44a5-b9f6-26956f8d76bf · inbound

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency cites this paper.

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:13:30.540842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T06:53:48.841077Z digest=sha256:a73e52f62e19cfea37b4d019cd0eecc8a48ec2274b96320e7cb1e90c070d8e76

Observation 965f875d-4656-4dcb-ba14-a9069b9c74c0 · inbound

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense cites this paper.

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.787116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:26:22.434043Z digest=sha256:4063d9b5cce5188d52dedafe1b62ed43417b0c3517a64d41cd7667882916592a

Observation cd011455-5dfd-41fa-a915-3f2ceec8281d · inbound

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing cites this paper.

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:07.162950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T21:23:32.255067Z digest=sha256:fd5f6d256628e9c59e15db3b8d4c4cfa14bf234c29f330911adcd8cf2ab3d0f8

Observation f4717228-edf6-44d4-814c-bec27062b365 · inbound

Hephaestus: Toward a Cybersecurity AI Scientist cites this paper.

Hephaestus: Toward a Cybersecurity AI Scientist Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:54:44.479468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T05:42:32.460183Z digest=sha256:ab01c1d75f8a250db5eb998c590755819d7b315ddb191de0b25f1dddf8a29a57

Observation 5371b2c9-2f19-4d1c-b521-5fb7f4011829 · inbound

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges cites this paper.

A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:11.585499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:11.585499Z digest=sha256:10b261c305dde2dda0e1943835c0d65d90e6c07c344fc9f719decf1c6be19a4c

Observation 770e4217-b261-4293-ba14-82a3cefdce29 · inbound

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI cites this paper.

Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:28:52.906906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:28:52.906906Z digest=sha256:8d2f204a0322bc7446160e874667f5b387026272c22b06b35ab1c094353e440d