Pith. sign in

Paper Citation Record · LEDGER

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

As of 12 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.15263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15263 v3

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:44:37.460029Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f9b988f-390e-488a-9a40-9ffbc46839e1 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.305552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.305552Z digest=sha256:dcc83b4f605ec36eacd35977bb1a41c22864d1c07f4f1933406223f1917ab091

Observation bb4687d3-3666-4250-ba48-4b349a2448f5 · outbound

This paper cites Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.414433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.414433Z digest=sha256:3353b22223b974554c011e939b9c4affaa7a361cabf6d39d4df8411b43b13d12

Observation 6f15dcdf-788d-4a8e-916d-c88dd10634d8 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Benchmark Data Contamination of Large Language Models: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.694280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.694280Z digest=sha256:ad2ec000e36ed0c914a9498fd06948e472a7137cdaa5c3d4b93c729fb54ea785

Observation 633f3a86-5e26-45a0-b68f-5367c8e1bbfb · outbound

This paper cites Benchmarking LLMs in an Embodied Environment for Blue Team Threat Hunting.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Benchmarking LLMs in an Embodied Environment for Blue Team Threat Hunting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.754442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.754442Z digest=sha256:f685175fcbbeef1415a345a042c33696f952da65fa45a11586863df381e912b8

Observation 11c482a4-d7cf-4797-b409-3391f45c000e · outbound

This paper cites Paul Mockapetris.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Paul Mockapetris

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.833647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.833647Z digest=sha256:26d1783449fd695310cba20119c69cf58663be8a67b79a738d99c0c9e3e38a05

Observation 8d20a8c4-390a-4650-a315-0bcba00456e7 · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.919216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.919216Z digest=sha256:687684edcdb395ae9af30b0f475f0bffff20fb5256b946d7be0bb26b69233b3d

Observation 46ea5393-03e2-4810-abf8-8a73dd5fbedc · outbound

This paper cites Hacking CTFs with Plain Agents.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Hacking CTFs with Plain Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.985500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.985500Z digest=sha256:82068eaf493e294720c8b9aa42db80c345fea0fbf151f901c1aa3ba25fda71d6

Observation a28e5a39-caff-46fe-bdd2-8c5025751ff8 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.044681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.044681Z digest=sha256:09a3240552c9ebddb04f76ec4e9bec5516ab9874ab2444d9932a95eb2bb8f961

Observation f42e648e-7f65-4ff9-8770-9e6a013df65d · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Benchmarking Benchmark Leakage in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.126253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.126253Z digest=sha256:f8a73b2368777723d378c71b9bf362f4709d7c1a8ba296ee7f7bfc964ca0f544

Observation be40b139-ad95-49c7-ad4b-9990fcec6388 · outbound

This paper cites ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.200360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.200360Z digest=sha256:abc5dc968f07586da131f32b9e88a7c6431c478ebe89e16477b61f9df26f500b

Observation 873331ca-5b6c-450b-b430-feab0bd94684 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.248320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.248320Z digest=sha256:f7130f6ee1a083a61850dcda4f73d0565d0deda5c76387b59a41f6ba3fd4f572

Observation e5ec7e9b-7c8e-48ae-9d35-a983f0fcfb48 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.323064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.323064Z digest=sha256:0d47d3ad9c1c90d97701e770f959e554b70f95223e8da4d9b99f8db6a8127205

Observation 17b90d1e-5d53-4080-9e58-8f1b0e0cae06 · outbound

This paper cites Yuxuan Zhu, Antony Kellermann, Dylan Bowman, et al.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Yuxuan Zhu, Antony Kellermann, Dylan Bowman, et al

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.402394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.402394Z digest=sha256:36cd615b0193e3d6707d3a3f80884561ec9c976c0ed538edc617efb08f53be59

Observation 00a49258-7ad4-4e0c-8a36-163edbb7f068 · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.460029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.460029Z digest=sha256:51267a8c319dedbcaaee82448d5432839d04e61e15e2100f75bafd66e71d596a

Observation a183dcca-28ed-4f5b-9a6e-bebc856ba4c5 · outbound

This paper cites rfc-editor.org/rfc/rfc3912.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents rfc-editor.org/rfc/rfc3912

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.359957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.359957Z digest=sha256:964821f88379acea0acf56d210a9b266f4eab52b0a505cd210fbc5f4ff408fec

Observation f6781fd7-ad21-4b2c-9022-04703038ec8a · outbound

This paper cites URL https://aclanthology.org/2023.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents URL https://aclanthology.org/2023

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.487351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.487351Z digest=sha256:adc2b90fc3af794aa3204583a96a304c84f6fce18dc99052eedffb819652bb1e

Observation 9ef3975f-3be5-422f-bde5-452fd7a31623 · outbound

This paper cites URLhttps://aclanthology.org/2024.eacl-long.5/.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents URLhttps://aclanthology.org/2024.eacl-long.5/

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.222177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.222177Z digest=sha256:fef36aad77c5660004d126ee0987c4542c68c1c605416e88441c061e18d503b9

Observation bb12d759-4240-4c45-9150-1a714929e72c · outbound

This paper cites SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.537877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.537877Z digest=sha256:56d48719e4895ed5d96f4909ae9d753cbb1dcb4cabc88f8815aa25dd8861cc7e

Observation 1066791f-c241-420e-8b08-88f33fc59dee · outbound

This paper cites Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:36.622992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:36.622992Z digest=sha256:38f513f8fc4468f7c07d802dfdf0c1d725eb6402f7bbd8dfd87135ca90e0e701

Pith citing papers

No inbound Pith citation observations are available.