Pith. sign in

Paper Citation Record · LEDGER

LLM Cyber Evaluations Don't Capture Real-World Risk

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2502.00072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00072 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:04:33.593990Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:10:11.025662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T02:07:33.281632Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b184fcb-075b-470a-a0cf-0d22dbc76a76 · outbound

This paper cites Introducing computer use, a new Claude 3.5 Sonnet , and Claude 3.5 Haiku , 10 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Introducing computer use, a new Claude 3.5 Sonnet , and Claude 3.5 Haiku , 10 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.663333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.328815Z digest=sha256:7e210039ec1b30d2a7094366e9f25083372ca6c5deeab3512985aee96b73028f

Observation fc76035d-be20-42f0-87af-4cb0ed18e822 · outbound

This paper cites Phishing Activity Trends Report , 4th Quarter 2023.

LLM Cyber Evaluations Don't Capture Real-World Risk Phishing Activity Trends Report , 4th Quarter 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.646352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.335713Z digest=sha256:3968ffcc25ed5ebd200b9e15ade6e4168e4883c5d53d6eab14229df281796cce

Observation 08c4b518-f91c-4c21-9ebd-b50c53109244 · outbound

This paper cites Phishing Activity Trends Report , 3rd Quarter 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Phishing Activity Trends Report , 3rd Quarter 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.628895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.341958Z digest=sha256:fd04e5053589da15f57dbe103746235ae9492c964ad17cabe250b67864ce6682

Observation fb3ca9dc-e754-4cb1-aafd-d21afdbda140 · outbound

This paper cites and Kogtenkov, A.

LLM Cyber Evaluations Don't Capture Real-World Risk and Kogtenkov, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.611472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.347979Z digest=sha256:b61b386b5c101a4ff6420436dd19250c26247c25db3ffcd1809011b9ae4fa03c

Observation ed79ba16-725a-4a3d-8408-82246b218a83 · outbound

This paper cites Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities.

LLM Cyber Evaluations Don't Capture Real-World Risk Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.354128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.354128Z digest=sha256:5e94762bfca43fe7e7733946450ddc51d6d2962b682196d9f5dca31334966809

Observation 1d434117-3d9b-49e2-921a-e9a825aa9ec7 · outbound

This paper cites "Real Attackers Don't Compute Gradients": Bridging the Gap Between Adversarial ML Research and Practice.

LLM Cyber Evaluations Don't Capture Real-World Risk "Real Attackers Don't Compute Gradients": Bridging the Gap Between Adversarial ML Research and Practice

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.361159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.361159Z digest=sha256:454639f6917bdd9613e07a11b47ce96bc34996c1dad5c9eedacc2c229e5b5c72

Observation a859d208-3314-48fe-8ea4-cb7d3c5f5364 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.368368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.368368Z digest=sha256:75715819c2b4a0c3832363cd96a58dabfb13a5421e08335a82a13704fb7f4c6a

Observation 95c0f915-db67-47e4-ae77-5fb642fa5e28 · outbound

This paper cites From naptime to big sleep: Using large language models to catch vulnerabilities in real-world code, 11 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk From naptime to big sleep: Using large language models to catch vulnerabilities in real-world code, 11 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.594239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.374456Z digest=sha256:d1ea9158182d29c3192c4aa57328a46fc43a4893a902fd30b85f173f8043e575

Observation 8485747e-d971-41e3-9105-3c754be4f8ed · outbound

This paper cites The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation.

LLM Cyber Evaluations Don't Capture Real-World Risk The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.379294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.379294Z digest=sha256:e501c07fecba5e714dfb2938964cbd77ee646ba5d5d117bf9e2a3384bead9736

Observation 43d800fb-3ce9-4eb8-b9c0-5a955b284f57 · outbound

This paper cites FunkSec – alleged top ransomware group powered by ai, January 2025.

LLM Cyber Evaluations Don't Capture Real-World Risk FunkSec – alleged top ransomware group powered by ai, January 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.577335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.385053Z digest=sha256:e9a208d6c7156fbf8ad3d92e1855d63d1ffac57e3e678a58f160447ef9d65d1f

Observation bc70c310-3b1d-48ab-9d0a-263ab40871b0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LLM Cyber Evaluations Don't Capture Real-World Risk Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.390435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.390435Z digest=sha256:d75c23e49e4f3a8c2217d2c166387235de3cc59aa65fbbdb23a1ac00aa82f503

Observation 453d3976-525f-473d-8a7e-c3502375ead2 · outbound

This paper cites PentestGPT: An LLM-empowered Automatic Penetration Testing Tool.

LLM Cyber Evaluations Don't Capture Real-World Risk PentestGPT: An LLM-empowered Automatic Penetration Testing Tool

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.395712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.395712Z digest=sha256:392e0fa4523c49cfe02dc1acb2913eaae60d1c609eee0e951ba90999e97913ed

Observation 210dca36-a026-428d-8629-726b72f7e63f · outbound

This paper cites Robust physical-world attacks on deep learning visual classification.

LLM Cyber Evaluations Don't Capture Real-World Risk Robust physical-world attacks on deep learning visual classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.560656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.401150Z digest=sha256:9889e0b8070394fd1a64f6b175a7eeaaf0dc7c1dd5163d05377d74b4fce827d5

Observation 63bcd5fb-1dcb-4b55-bca7-d0b4571c825c · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

LLM Cyber Evaluations Don't Capture Real-World Risk LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.405915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.405915Z digest=sha256:62579f63c479bac37343ffa5ec439919d4e96941d20dac339e47e87063a6f390

Observation c15c0154-c66e-497c-a7a7-88b441c96339 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

LLM Cyber Evaluations Don't Capture Real-World Risk LLM Agents can Autonomously Hack Websites

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.411892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.411892Z digest=sha256:8e45cdad11d20b28b65eec1222829ba5c73fc8f683290097164c0a3a34c0c290

Observation fe24e7d6-45c6-4908-9b35-38e2300b9234 · outbound

This paper cites Teams of LLM Agents can Exploit Zero-Day Vulnerabilities.

LLM Cyber Evaluations Don't Capture Real-World Risk Teams of LLM Agents can Exploit Zero-Day Vulnerabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.418057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.418057Z digest=sha256:f68d56d4ee33385c5291183d403931f57102fa4007bfa1dd678c954d51e501d6

Observation cb522df9-8e68-4256-a28d-4503783d83f1 · outbound

This paper cites 2023 Internet Crime Report.

LLM Cyber Evaluations Don't Capture Real-World Risk 2023 Internet Crime Report

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.542751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.424270Z digest=sha256:62a5d877e10950ca67f72b25436a9f53bc33e40d077b157b5095b8f63d1ce454

Observation e22701fb-69b4-4d82-b6fd-02173e688d60 · outbound

This paper cites BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B.

LLM Cyber Evaluations Don't Capture Real-World Risk BadLlama: cheaply removing safety fine-tuning from Llama 2-Chat 13B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.429874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.429874Z digest=sha256:0f317c43d15b8ea1beb6802f316d61ca1e436ebb8321d82ec54a3c55649d8723

Observation 5c626d25-7818-471a-b928-31f4b8ba6b14 · outbound

This paper cites Sour grapes: stomping on a Cambodia -based ``pig butchering'' scam.

LLM Cyber Evaluations Don't Capture Real-World Risk Sour grapes: stomping on a Cambodia -based ``pig butchering'' scam

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.524676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.436838Z digest=sha256:087ee131025f41349a54a49d8a442c8cc09c6dc68aa0529bc66ae34dd86007f4

Observation 5e4213b4-87fb-45f4-8f6b-6a4dcc5b77cb · outbound

This paper cites Safety case template for frontier AI: A cyber inability argument.

LLM Cyber Evaluations Don't Capture Real-World Risk Safety case template for frontier AI: A cyber inability argument

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.442414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.442414Z digest=sha256:5dea50a3237259259c672ac48e1ba94a966badd37f33ac299364d3b395cfc6d1

Observation b51324d8-4b85-41b8-900d-458036ad85ff · outbound

This paper cites Adversarial misuse of generative AI , 2025.

LLM Cyber Evaluations Don't Capture Real-World Risk Adversarial misuse of generative AI , 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.505804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.448387Z digest=sha256:2742d03aecd2fa0dbecff4440c4f11e088310587b20b2ee6cfa0ef947f7dfde7

Observation d423440c-c0c8-4d27-882c-bca395718d24 · outbound

This paper cites and Cito, J.

LLM Cyber Evaluations Don't Capture Real-World Risk and Cito, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.452963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.452963Z digest=sha256:38be06ad1460e1451a83814a4e8c3a14ec81815b5886dd06a21f8b56b984db92

Observation ba7e8388-6430-4110-b26f-cedd2ab27edb · outbound

This paper cites Spear Phishing With Large Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk Spear Phishing With Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.457578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.457578Z digest=sha256:f588bc7e3c7814eba0b296718c00aae4adef00255314c6196d0f5c8662ba51fb

Observation 5c2bca68-90e1-41da-8702-9d652720202e · outbound

This paper cites Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects.

LLM Cyber Evaluations Don't Capture Real-World Risk Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.462290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.462290Z digest=sha256:2907f5f5791259638b3da59276c21dd3902de26c463a118e8ba45d3671e8c904

Observation d5bdeb83-b35b-4899-90dd-6e33de2dbc6d · outbound

This paper cites Devising and detecting phishing emails using large language models.

LLM Cyber Evaluations Don't Capture Real-World Risk Devising and detecting phishing emails using large language models

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-09T22:04:48.912079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.467563Z digest=sha256:58ae93e3dd86a939764c4623732ca4508876881be4b718de77e6e335d46fd3e6

Observation db3a8d13-1121-4919-8dfb-e6678f9de43d · outbound

This paper cites An Overview of Catastrophic AI Risks.

LLM Cyber Evaluations Don't Capture Real-World Risk An Overview of Catastrophic AI Risks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.472242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.472242Z digest=sha256:d81c7a3db6c437c06612a72eb06628c79c2342104d8d93ab981f3554fc6159cc

Observation 7c6542e2-94c5-4c4d-9565-6125617200a6 · outbound

This paper cites Why do nigerian scammers say they are from nigeria? Proceedings of the Workshop on the Economics of Information Security, 01 2012.

LLM Cyber Evaluations Don't Capture Real-World Risk Why do nigerian scammers say they are from nigeria? Proceedings of the Workshop on the Economics of Information Security, 01 2012

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.486659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.477435Z digest=sha256:00a2c195d46f7e49ca0593080a95b812a8a958f305b273a9b15304e61cee039f

Observation 0352e941-d60b-4224-b8ea-e544af213495 · outbound

This paper cites Now you see me, now you don't: Using LLMs to obfuscate malicious JavaScript.

LLM Cyber Evaluations Don't Capture Real-World Risk Now you see me, now you don't: Using LLMs to obfuscate malicious JavaScript

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.466991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.483374Z digest=sha256:9f4a26ec0447b5cbaf508d15e99e763c957812bf300bfe1e5680744bb50d0d57

Observation 204d8de5-1490-4f04-a694-d6614981b0e3 · outbound

This paper cites Stored cross-site scripting (xss) in 2FAuth : How XBOW found a stored XSS in 2FAuth ( CVE -2024-52597), 12 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Stored cross-site scripting (xss) in 2FAuth : How XBOW found a stored XSS in 2FAuth ( CVE -2024-52597), 12 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.447646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.488615Z digest=sha256:12770332c58b32748f75c537e2c261c017769c2d189d007663238c79c5a998fb

Observation 27b08ecc-d8b3-4fe2-960e-f12d2f66672e · outbound

This paper cites Translated: Talos ' insights from the recently leaked Conti ransomware playbook, September 2021.

LLM Cyber Evaluations Don't Capture Real-World Risk Translated: Talos ' insights from the recently leaked Conti ransomware playbook, September 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.424914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.494029Z digest=sha256:7c4e78b3763e12565c8e99097f293cd1b63cf4aac9fe51cd1ba806937c73cdcd

Observation a4a79575-6512-4eeb-a6e1-73bcf464c3a5 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

LLM Cyber Evaluations Don't Capture Real-World Risk HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.499853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.499853Z digest=sha256:568693090620864043bab8ac785aaf04d14e0416f361ffe3a8d123d83fb11166

Observation 12c1686f-abf5-498f-93b0-62c2c9ca7a6a · outbound

This paper cites An update on our general capability evaluations.

LLM Cyber Evaluations Don't Capture Real-World Risk An update on our general capability evaluations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.405115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.505725Z digest=sha256:b4baca0b98c65b6328bd9aa7ced51ef3327491c0a794898a81025dcd64e43e7e

Observation ec9198f9-877b-451e-b4d4-c1d99862bdcd · outbound

This paper cites The Threat of Offensive AI to Organizations.

LLM Cyber Evaluations Don't Capture Real-World Risk The Threat of Offensive AI to Organizations

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-09T22:04:33.793851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.511132Z digest=sha256:733963611bc9cdbc870a4066f63019fd90a6388a84c436769797f44e4927b616

Observation c34d9cf7-baea-4763-8bc3-6a0ec7a7c574 · outbound

This paper cites MITRE ATT&CK.

LLM Cyber Evaluations Don't Capture Real-World Risk MITRE ATT&CK

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.385448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.516850Z digest=sha256:953998db07c5dab873b4678cb41a8b443f77fd81dd0961689f149112f02b67cc

Observation 8b6af47e-590e-4d83-aa3a-ef7020aaee5c · outbound

This paper cites and Flossman, M.

LLM Cyber Evaluations Don't Capture Real-World Risk and Flossman, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.364755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.522871Z digest=sha256:820edde335c9426ba3795e17f4a9456b82be8d2557b70c0848800f7e5b357ec8

Observation 42633144-4e19-4a8d-a742-cd32bce9bb40 · outbound

This paper cites Introducing Operator , 1 2024 a.

LLM Cyber Evaluations Don't Capture Real-World Risk Introducing Operator , 1 2024 a

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.344228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.528353Z digest=sha256:6ed6b6e68e6eb040df79055860ed5af4b12749d63c01cbc9fe3080278d0b8a8d

Observation 584f2ce6-6679-4b7b-a517-7f9fd7efe21e · outbound

This paper cites Disrupting malicious uses of AI by state-affiliated threat actors, February 2024 b.

LLM Cyber Evaluations Don't Capture Real-World Risk Disrupting malicious uses of AI by state-affiliated threat actors, February 2024 b

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.325794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.533729Z digest=sha256:6612a896b0305ad95afcffe4562a58ba25b4270ea5e080b5eb137fc3c7b0471d

Observation d37c8ff2-dc4e-4943-acbb-6e223e991712 · outbound

This paper cites and Haugen, S.

LLM Cyber Evaluations Don't Capture Real-World Risk and Haugen, S

Reference 38

Resolution
verified exact
doi, observed 2026-08-09T22:04:33.640065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.539035Z digest=sha256:eb51e1a2f9bc1f4b5b1e8bc81860a9e563e7ae6f9d770e09be7949c0000681db

Observation efc7f0e1-9d42-424e-89e6-55b39d2c85f9 · outbound

This paper cites No, LLM agents cannot autonomously "hack" websites, feb 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk No, LLM agents cannot autonomously "hack" websites, feb 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.306353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.543901Z digest=sha256:f15d3db5b14a0594184b27ff84586671eb12d3b1d62d4717144d661790277924

Observation 22c56035-901b-4691-9c58-eb886fe33a14 · outbound

This paper cites SoK: On the Offensive Potential of AI.

LLM Cyber Evaluations Don't Capture Real-World Risk SoK: On the Offensive Potential of AI

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T22:04:33.767100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.548908Z digest=sha256:2b5fd1a63d70e72cbd4127fbd6aa9a4f3a2ca2d430e148727f366b9eaff20fc7

Observation 04ff266c-6a39-49da-ba99-b6aab206b6b1 · outbound

This paper cites NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security.

LLM Cyber Evaluations Don't Capture Real-World Risk NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.554414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.554414Z digest=sha256:8d88054d5fb6e883c731d6e9eef1e757565e6889c31ad14125eec8357c61aa7b

Observation 465a581a-dcfb-4462-9f82-f3ccd21ef63e · outbound

This paper cites Model evaluation for extreme risks.

LLM Cyber Evaluations Don't Capture Real-World Risk Model evaluation for extreme risks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.559543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.559543Z digest=sha256:082c1b5e1f0a90c83796e822348f870f777c2d39d2eba608ccb2e76984065b5e

Observation 5329ec5e-6b9a-45b2-8b17-6366992c2269 · outbound

This paper cites Hacking CTFs with Plain Agents.

LLM Cyber Evaluations Don't Capture Real-World Risk Hacking CTFs with Plain Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.564625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.564625Z digest=sha256:e59c962d675746d741f43a5cd2fe5de56f9582acb7e43cb6ff5892aa711f25d4

Observation a33ab2e4-6974-4332-9a83-d6bd7a00f22f · outbound

This paper cites Advanced AI evaluations may update, 2024.

LLM Cyber Evaluations Don't Capture Real-World Risk Advanced AI evaluations may update, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.286541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.570205Z digest=sha256:dec2d46da13cf2c1f8fa158219d679b43bc5ef73aad64d472b1abda4dd859bd8

Observation 7fe49fbc-4484-401c-a2a3-7794bab2e757 · outbound

This paper cites The near-term impact of AI on the cyber threat.

LLM Cyber Evaluations Don't Capture Real-World Risk The near-term impact of AI on the cyber threat

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:05:04.267306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:04:33.575782Z digest=sha256:12592a3e4fcc700a3665d95275c66622fc75800a10dd1eae786d3ede943a7cf1

Observation 5ef8df1f-46c0-45de-876c-9a1ce7f36373 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.581400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.581400Z digest=sha256:b27055cdcbd1e48dd614ce2fdd0016203eb9b14db9cb0b82064d447d5670e718

Observation 405ac15e-b490-4ee3-996c-fbf0126befe2 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

LLM Cyber Evaluations Don't Capture Real-World Risk Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.587373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.587373Z digest=sha256:2c0a7eb87748d913033416ee44f8b641b50a0fc270338db0c1bf884beb4b1695

Observation a7b32975-a92c-48b2-a216-75f3a2c4a383 · outbound

This paper cites write newline.

LLM Cyber Evaluations Don't Capture Real-World Risk write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.593990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.593990Z digest=sha256:6f9577f54c836e07c4eb130a7a2eaeece39b8bc827ebbc06fdbad70ee7eecec3

Pith citing papers

Observation cb5c6c48-7cbe-4570-96c3-612dbd8bfe28 · inbound

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks cites this paper.

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T23:10:11.025662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:10:11.025662Z digest=sha256:2e7e90a83c88e4eaa7bfaebf54ac7f0e96f0269eac18fa2d49bd7772128721ee

Observation 2064aa91-2129-48b3-b60f-39c01393918b · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:06.577088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:06.577088Z digest=sha256:3d3aff86c1e3a0267d14f75566930f46f4cc61fbd52f69bc2abd9db02268f0b3

Observation 6374f621-c707-4497-b01a-f9b1c90a2c30 · inbound

The coordination gap in frontier AI safety policies cites this paper.

The coordination gap in frontier AI safety policies LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:24:10.760233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T12:21:06.900491Z digest=sha256:f7acab28d14c05ff5ded7a37807670e3798294fd502ceadd56688d814bd3f2f6

Observation dd764713-2be5-4852-b2a9-32a041fabb52 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting LLM Cyber Evaluations Don't Capture Real-World Risk

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.283723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:5930c02d8a562a6298b37e82af2e8a41e174b15ea7c95973f5a5d6c960f62f4b