Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2408.08926.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:04:33.587373Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T18:37:31.168670Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0feb9130-5989-413f-818a-86b3d31b83ba · inbound
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 936f24ff-bbee-4b68-afa7-b98308ff62ee · inbound
Frontier Models are Capable of In-context Scheming Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation be75ed09-53c6-4815-a5d7-86f23bc03ac8 · inbound
Humanity's Last Exam Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 405ac15e-b490-4ee3-996c-fbf0126befe2 · inbound
LLM Cyber Evaluations Don't Capture Real-World Risk Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd94e89-ae74-4576-a430-c68e8cb1942c · inbound
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dd76008-2c7c-4b8e-a7d3-32227712fbce · inbound
A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c21991c-eadb-487d-9a76-7927a1c1cf7b · inbound
CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7718ec-e1a9-48e5-92f1-5b9dca55f357 · inbound
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7794c2eb-86ab-446d-9979-d9af2c171628 · inbound
Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71a0b358-a58c-4eae-a135-ff47b069821d · inbound
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb576ed1-53fe-4da0-ab9c-9c13e0aa1f7e · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 290af98a-b2cf-432e-bfdb-cf3e875068bd · inbound
Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · inbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00418eaf-ba6c-4c13-b005-b4182b51cdd1 · inbound
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca2ca46d-83e1-4bf8-b634-cbf1ce3a42f3 · inbound
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e33d96-9c84-4793-a4d4-08cedb088a61 · inbound
Agent Identity Evals: Measuring Agentic Identity Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5207f360-b370-4439-a14b-e9273761fc7f · inbound
Evaluation and Benchmarking of LLM Agents: A Survey Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 063639a0-1e72-4497-a8a7-3174e901d29a · inbound
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d6da7f-8a57-4eda-a3e2-58af7100db2c · inbound
PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5932e34-fc2d-4149-bfda-417ff6e36dd9 · inbound
Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffd15fe-f1cb-47df-870a-e1f7d5365307 · inbound
Quantifying Frontier LLM Capabilities for Container Sandbox Escape Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0189277-b3d4-40eb-b3dc-9def2a74622a · inbound
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c7fb686-9676-4ddb-b92a-76b5e3a3191a · inbound
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f31584-7ea9-41de-ab10-7890120d3c43 · inbound
AlphaEval: Evaluating Agents in Production Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 670e3a77-f49b-4e6a-b904-6acda6aaf242 · inbound
Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98a4d09b-b3c1-483c-8fd2-b68f5d949d6f · inbound
Can LLMs be Effective Code Contributors? A Study on Open-source Projects Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af10d1ef-b6f1-4213-8cdb-52e0b21494d8 · inbound
Dynamic Cyber Ranges Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 901fe0bd-f957-4028-ad60-c19de5e945c4 · inbound
Risk Reporting for Developers' Internal AI Model Use Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f3479441-714c-411a-b0ed-5fb2f8e5e57e · inbound
XekRung Technical Report Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f2e4b70-ed75-4997-8ad4-4ded7d60ee4d · inbound
Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c487224-e811-4da9-89bd-d92904e4ba7f · inbound
Autonomous Adversary: Red-Teaming in the age of LLM Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb629ce5-673b-4c01-a4d6-b413708b2c71 · inbound
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 056324e2-fb99-4104-b553-afcf4b86cac3 · inbound
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f204818c-98b3-493f-ab84-07171811d024 · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 24ac8b9d-0e14-4b24-9694-01757c5bc2be · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · inbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 33d9d0cc-241a-4c13-8df6-342a25d52c98 · inbound
Benchmarking Mythos-Linked Bug Rediscovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8584ade9-95b6-446c-bf9e-18902bde487f · inbound
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 305c9254-29ad-43f3-8569-8b25cfa8cddf · inbound
HIDBench: Benchmarking Large Language Models for Host-Based Intrusion Detection Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e042dcc-a9f5-4faf-8701-a82bb982bb3f · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 001041d1-9921-45e9-9866-f3db5b16d6ed · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e1e3c4c-a3a3-4fc0-85fc-a7b7e84f2e31 · inbound
Cybersecurity AI (CAI) Dataset Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab106e19-42cb-47ed-b501-4f3b97b30ebf · inbound
Stateful Online Monitoring Catches Distributed Agent Attacks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d56a6cf9-90bf-410a-b13a-6f3bd59e972b · inbound
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 99c16f9a-1b02-49bb-b4e3-14ce6bf99e4a · inbound
Poisoned Playbooks: Demystifying Knowledge Poisoning Effects on AI Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f1e2d4fa-4aef-4e74-9e82-9d19555f84d9 · inbound
Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b35eadfd-1caa-44b8-a2e7-2eb5d55a9303 · inbound
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ef5f186d-bd5e-4b54-b7d1-5f3115183312 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 65c92e11-d7bc-443a-8ac8-4bfc8b6376c6 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ec7e9b-7c8e-48ae-9d35-a983f0fcfb48 · inbound
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceff4a75-aac0-40df-b438-27b6d540f7d0 · inbound
Harmonizing AI Safety Thresholds Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23646153-646d-483d-acfc-45130e22460a · inbound
RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 911f64eb-a401-452b-86ca-d0fc8217fff4 · inbound
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579da816-ab2d-4ffb-98f7-0939b6df73a4 · inbound
The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c75942d-277b-4329-a925-2a8d7343f2de · inbound
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc8c6be-b185-4ddf-95ee-b317d5299044 · inbound
Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.