Pith. sign in

Paper Citation Record · LEDGER

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2408.08926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08926 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:04:33.587373Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T18:37:31.168670Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0feb9130-5989-413f-818a-86b3d31b83ba · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:35:51.135418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:e79b8e72d660206f859c5b0fe081f38fccc96e33837e55b68b9ab0fc1845f288

Observation 936f24ff-bbee-4b68-afa7-b98308ff62ee · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:22:01.635191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:000242398128b6dc425e2fb2b27305842eb6002fa87fd1d8660407e0c24428f0

Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.350304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:593a9c8e04be8cc47fbcfa3d7cbf3643ec1f06c17494e6497488b09148b077e1

Observation 405ac15e-b490-4ee3-996c-fbf0126befe2 · inbound

LLM Cyber Evaluations Don't Capture Real-World Risk cites this paper.

LLM Cyber Evaluations Don't Capture Real-World Risk Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.587373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.587373Z digest=sha256:2c0a7eb87748d913033416ee44f8b641b50a0fc270338db0c1bf884beb4b1695

Observation fcd94e89-ae74-4576-a430-c68e8cb1942c · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.149366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.149366Z digest=sha256:065a988b844d41049a6925f86fea2bad74811b548c937b95491ffe7f5e1851bd

Observation 4dd76008-2c7c-4b8e-a7d3-32227712fbce · inbound

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management cites this paper.

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:44.828234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:48:44.828234Z digest=sha256:4437262d46f313d810252851ef003c90e997700eb1754c168c07ac8610f2582f

Observation 9c21991c-eadb-487d-9a76-7927a1c1cf7b · inbound

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution cites this paper.

CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:52.914666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:52.914666Z digest=sha256:9c257f0f230ee21afc7bd3f795aabdb7253f2877c71b504beebdd58fe6448d2b

Observation bc7718ec-e1a9-48e5-92f1-5b9dca55f357 · inbound

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges cites this paper.

Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:35.406709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:35.406709Z digest=sha256:12f4cbb6de2f238df4fe681895618e8679fdfc2aad8212e13163e9306e7fedcf

Observation 7794c2eb-86ab-446d-9979-d9af2c171628 · inbound

Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research cites this paper.

Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:20.166338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:20.166338Z digest=sha256:722160170684c30d69890073bcb2fb713c781c1b6c09fd7ebdd479017493db49

Observation 71a0b358-a58c-4eae-a135-ff47b069821d · inbound

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs cites this paper.

From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:35:03.624759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:35:03.624759Z digest=sha256:4f59bcaaaba674d6d3953faffd2d4331ba65cc0ba6f15b7fcfee714dd68d20d0

Observation cb576ed1-53fe-4da0-ab9c-9c13e0aa1f7e · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.287235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.287235Z digest=sha256:b911e828d3e0a7ae6c370522142dff9e63ab3be36fd8cb3d29e517eb938e9e3c

Observation 290af98a-b2cf-432e-bfdb-cf3e875068bd · inbound

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework cites this paper.

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:59.356503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:59.356503Z digest=sha256:ec43d9ae5e0f3b548d9250e36b523df2539f510305d2b6e9c8ef2189b2bae4fa

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · inbound

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights cites this paper.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:4c3a1b1d916ffc92267a504d5e0c1f7e332df1b05ec048cc03f34069f18d2837

Observation 00418eaf-ba6c-4c13-b005-b4182b51cdd1 · inbound

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation cites this paper.

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T04:42:04.771929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T04:37:33.942379Z digest=sha256:84f465ec533e0916dab3d0f4cd3fc78f090ef0116f3c8a25c2218a89ec822c9e

Observation ca2ca46d-83e1-4bf8-b634-cbf1ce3a42f3 · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:23.006196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:23.006196Z digest=sha256:60826a401aae7b34d0ef2938a75348586b40a856dd44cae4e179ab69dd9e64a0

Observation a2e33d96-9c84-4793-a4d4-08cedb088a61 · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.491612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.491612Z digest=sha256:5cd3538a2fe1e79a3fffd251f516284c2097026a61c157c68954f3963c549589

Observation 5207f360-b370-4439-a14b-e9273761fc7f · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.858357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.858357Z digest=sha256:9fdb42194a83987dd316d3d2c2d7e5ae49dc355710d931395a044840282dbea4

Observation 063639a0-1e72-4497-a8a7-3174e901d29a · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.522003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.522003Z digest=sha256:9a1a0e421364d1cbe3ef898d81cbeaa5f0c2848ea75cbcb0c2787112b45ab86c

Observation 99d6da7f-8a57-4eda-a3e2-58af7100db2c · inbound

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts cites this paper.

PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:09:36.085164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:09:36.085164Z digest=sha256:e3e0b740c1a25e91d603702649d754eeaf08449d80f21ff51d19f54b7517c6bf

Observation d5932e34-fc2d-4149-bfda-417ff6e36dd9 · inbound

Quantifying Frontier LLM Capabilities for Container Sandbox Escape cites this paper.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T19:43:40.853518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:43:40.853518Z digest=sha256:2cc31b4434afa2da3e01889efd621bc63ec6f5e9d076353330390e7b4df73a17

Observation 8ffd15fe-f1cb-47df-870a-e1f7d5365307 · inbound

Quantifying Frontier LLM Capabilities for Container Sandbox Escape cites this paper.

Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T06:01:01.313606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:01:01.313606Z digest=sha256:9503f2742c1c24d89314eb834e6e0304b65bda67fcb1a1225d8ec1f5df7399e8

Observation d0189277-b3d4-40eb-b3dc-9def2a74622a · inbound

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing cites this paper.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:52.939979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:52:57.225878Z digest=sha256:b94a5ed29f3b61100c0979a20bb4244b2686399af4adbcc296225c9330c51381

Observation 3c7fb686-9676-4ddb-b92a-76b5e3a3191a · inbound

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills cites this paper.

SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:51.515451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:51.515451Z digest=sha256:9807027089c121570920c53ac9bf5795fea2c72f989c5e32be4c79e571d17e74

Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · inbound

AlphaEval: Evaluating Agents in Production cites this paper.

AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:45:58.817569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:30:51.886471Z digest=sha256:4189ce01f0b61ab0e3e145b8f797b9d43fbd391806e34ec3dbb65d1ae60ea24b

Observation 670e3a77-f49b-4e6a-b904-6acda6aaf242 · inbound

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks cites this paper.

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.087529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T06:02:32.399075Z digest=sha256:9c599f7eee8a2dd8884eebb3d771eb77050765f45e2e36f12c359705d7c6a064

Observation 98a4d09b-b3c1-483c-8fd2-b68f5d949d6f · inbound

Can LLMs be Effective Code Contributors? A Study on Open-source Projects cites this paper.

Can LLMs be Effective Code Contributors? A Study on Open-source Projects Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:10.852344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T08:09:33.211692Z digest=sha256:5f78fa778f01f6e00103fdc0bb1bb641fb770c4915133e090c5fc35649aaf8d5

Observation af10d1ef-b6f1-4213-8cdb-52e0b21494d8 · inbound

Dynamic Cyber Ranges cites this paper.

Dynamic Cyber Ranges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:16:53.348658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T03:04:03.611481Z digest=sha256:6ab21020d60bcb1e938fc11bdbc6dd6b2b05b11e7ff4b0adcc1257b0365ce9aa

Observation 901fe0bd-f957-4028-ad60-c19de5e945c4 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.399595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:b359fca6118ea4b1d7f4724a88db622ab506a4f402ef615e4cc957da1419b2fc

Observation f3479441-714c-411a-b0ed-5fb2f8e5e57e · inbound

XekRung Technical Report cites this paper.

XekRung Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:51:14.588143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T20:55:10.400291Z digest=sha256:aee668465899ca0d7db6165ff360c79fa1d785bd43d8b5e3d4753ea885886e49

Observation 6f2e4b70-ed75-4997-8ad4-4ded7d60ee4d · inbound

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting cites this paper.

Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:18.994757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T15:18:43.426681Z digest=sha256:e1c1bef813972858904d24224adff6fd9c8ccb405ab7c6bc65b2d9bcc9ea7b98

Observation 8c487224-e811-4da9-89bd-d92904e4ba7f · inbound

Autonomous Adversary: Red-Teaming in the age of LLM cites this paper.

Autonomous Adversary: Red-Teaming in the age of LLM Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:26:10.508553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T09:09:58.566171Z digest=sha256:331bf148c32eeb0de63babfcf9470acd217cd7bedcdb9612673a9cec1d13eae3

Observation fb629ce5-673b-4c01-a4d6-b413708b2c71 · inbound

Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches cites this paper.

Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:31:11.685007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T08:51:33.112454Z digest=sha256:6afebe323c5ea685b73806f4f3f0d9d184f477c550ed66e80d0adce8d92d36ec

Observation 056324e2-fb99-4104-b553-afcf4b86cac3 · inbound

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios cites this paper.

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:10:57.348437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:54:46.391345Z digest=sha256:9b332044059445ee38a482932c6d16106d573380519cb5acbe0045da9ee4ddef

Observation f204818c-98b3-493f-ab84-07171811d024 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.972213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:57f15e22bf5d14866fa2fa3cea340a74e51f7b1cd285faae7f696c8b1e53d837

Observation 24ac8b9d-0e14-4b24-9694-01757c5bc2be · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:27.655588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:27.655588Z digest=sha256:51b222ca0eec6222c274acc758321b5d329d0c0bd516c93cbab59bfa134c4e35

Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.563721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:504eefa1a879f90bc9458f717df480694c7a37e30552b96b7330051cbb7766ea

Observation 33d9d0cc-241a-4c13-8df6-342a25d52c98 · inbound

Benchmarking Mythos-Linked Bug Rediscovery cites this paper.

Benchmarking Mythos-Linked Bug Rediscovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:17:57.552731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T23:14:34.164854Z digest=sha256:6764e4f28c82a62b86f8e6149cfaa190c3a2fb457c6c7495d4d702f5eee5f291

Observation 8584ade9-95b6-446c-bf9e-18902bde487f · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:18:11.749362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:a95e7e9651c54d35e00a18e3cea7c712c48b728a931e2013271d540e8607788a

Observation 305c9254-29ad-43f3-8569-8b25cfa8cddf · inbound

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection cites this paper.

HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:54:45.823514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T08:52:34.079804Z digest=sha256:0fae05da5fc72f0915f19a3d6d595e231fb1b3b84820a926541c992e320cc08a

Observation 7e042dcc-a9f5-4faf-8701-a82bb982bb3f · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.929199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:a66dd6d311e3f6ba0b1adcff5abc5b45a002c42d3c09758558f905279d903340

Observation 001041d1-9921-45e9-9866-f3db5b16d6ed · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.185520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:a6dcd2c0a67b5cacdd7be7d13ef5705079080c15dbad3e9f85b99c90f2d23a43

Observation 0e1e3c4c-a3a3-4fc0-85fc-a7b7e84f2e31 · inbound

Cybersecurity AI (CAI) Dataset cites this paper.

Cybersecurity AI (CAI) Dataset Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:13:26.856759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T12:07:49.656453Z digest=sha256:c195f6f112748a86c6f27b2a681345051b56c9938eeb40f6090b7697d38c5580

Observation ab106e19-42cb-47ed-b501-4f3b97b30ebf · inbound

Stateful Online Monitoring Catches Distributed Agent Attacks cites this paper.

Stateful Online Monitoring Catches Distributed Agent Attacks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:11.064697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T21:54:44.072929Z digest=sha256:6de282a031bacc3a15cbd4a069c5bc872516cd60781c5cd3423537768d3440ea

Observation d56a6cf9-90bf-410a-b13a-6f3bd59e972b · inbound

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios cites this paper.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.344299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T03:39:36.657903Z digest=sha256:0be47a7171d37210a3cd888fc7e7c7963182429e85635c9bbd6f6209760126b6

Observation 99c16f9a-1b02-49bb-b4e3-14ce6bf99e4a · inbound

Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents cites this paper.

Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:40:01.102907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-25T23:27:14.610097Z digest=sha256:df2d1f668556296de5e323de835bf0a9308afcb3c7785b98fda2564034a82080

Observation f1e2d4fa-4aef-4e74-9e82-9d19555f84d9 · inbound

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations cites this paper.

Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.125038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:49:28.819402Z digest=sha256:2b16431b614331d81b6989f72ef52e2caa4baf8a108e60d7fe9de9be462b0598

Observation b35eadfd-1caa-44b8-a2e7-2eb5d55a9303 · inbound

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction cites this paper.

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.502047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T14:02:06.173081Z digest=sha256:2ed475c034a0941108eaf0c51e287da28ba3790453ccc35404e4906cdf600756

Observation ef5f186d-bd5e-4b54-b7d1-5f3115183312 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:37:31.169937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T18:29:50.731038Z digest=sha256:9fe883060e359c48ee23c2dc65aa61046c24e38842fe2de26cde6802b07d3c3d

Observation 65c92e11-d7bc-443a-8ac8-4bfc8b6376c6 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:47.662275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:49:47.662275Z digest=sha256:3d6caa63a4680e8a707bb7f63ad36bf3fac7492e61edc89339c748889bc9f40a

Observation e5ec7e9b-7c8e-48ae-9d35-a983f0fcfb48 · inbound

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents cites this paper.

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:44:37.323064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:44:37.323064Z digest=sha256:0d47d3ad9c1c90d97701e770f959e554b70f95223e8da4d9b99f8db6a8127205

Observation ceff4a75-aac0-40df-b438-27b6d540f7d0 · inbound

Harmonizing AI Safety Thresholds cites this paper.

Harmonizing AI Safety Thresholds Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:24:15.880141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:24:15.880141Z digest=sha256:6d40c0ee5ae7bd727125d8059783050c524a7b89745f0c77ffcca63f5216dbe0

Observation 23646153-646d-483d-acfc-45130e22460a · inbound

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery cites this paper.

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:09.259606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:09.259606Z digest=sha256:0b73ea125e97f6d45d896200349b4baf3c6011d5e093e17de85128ccb5e0e201

Observation 911f64eb-a401-452b-86ca-d0fc8217fff4 · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:26.953693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:26.953693Z digest=sha256:a55783e7b61428c18a20c72faaf1088a5e729752cb98f54a481812a2b85ed880

Observation 579da816-ab2d-4ffb-98f7-0939b6df73a4 · inbound

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play cites this paper.

The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:31:26.073853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:31:26.073853Z digest=sha256:202887337f73fde34d5de9435e8ebedd8d27ee513d7fbfe615eae851a1e1403a

Observation 8c75942d-277b-4329-a925-2a8d7343f2de · inbound

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents cites this paper.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.291705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.291705Z digest=sha256:bb2b5f3583eb634821362374c1278b50f43c40f6109b3a8d546d259a9d46c758

Observation ccc8c6be-b185-4ddf-95ee-b317d5299044 · inbound

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models cites this paper.

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:19:18.902915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:19:18.902915Z digest=sha256:1873ff63e74ecade72810c226d63b95e0181f48a54de8c53806ddeade5f4a8bf