Pith. sign in

Paper Citation Record · LEDGER

Evaluating Frontier Models for Dangerous Capabilities

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2403.13793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.13793 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:34:05.033901Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

11
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5397a145-bf4e-4251-a510-aa73e680b479 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Evaluating Frontier Models for Dangerous Capabilities

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T04:18:27.721335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:1b2d45f74dacba98ba103eac402ec771b9bcd73a838d4a106e3050ba87c88371

Observation 5e4744a9-9afb-4475-8606-520bfe93d33a · inbound

Gemma 2: Improving Open Language Models at a Practical Size cites this paper.

Gemma 2: Improving Open Language Models at a Practical Size Evaluating Frontier Models for Dangerous Capabilities

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:11:16.476742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T12:11:16.326752Z digest=sha256:3d86cd9bd2951a090eae7edcfbc2ae2f8aa57ba3c075dffc6d3b4e509f226156

Observation 8ebdf006-f847-4ed3-a907-983e3caeb6ac · inbound

Safety case template for frontier AI: A cyber inability argument cites this paper.

Safety case template for frontier AI: A cyber inability argument Evaluating Frontier Models for Dangerous Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T22:01:11.507302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T22:01:11.507302Z digest=sha256:18ddf7b9a7e60d6bb7dca6b88c00a46d982e039774a4e269822a7b85a29fcd45

Observation 15f0547b-94c7-4ce5-b869-0556aad8a54d · inbound

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment cites this paper.

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment Evaluating Frontier Models for Dangerous Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:02.620872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:15:02.620872Z digest=sha256:d81701cfea8ad041049e8505d721866fafbc68eaab1f096e93c7be0e24b80d78

Observation 0353eefc-190f-4539-bbd6-3a7d36c013b0 · inbound

Dimensions of Generative AI Evaluation Design cites this paper.

Dimensions of Generative AI Evaluation Design Evaluating Frontier Models for Dangerous Capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:21.945350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:15:21.945350Z digest=sha256:f8153b854e3ac7885f8e126cb96d353830b396350a38cf239c0c02a2195659e5

Observation 2f64c4e1-828a-4706-9b42-47ed2dace3a8 · inbound

Best-of-N Jailbreaking cites this paper.

Best-of-N Jailbreaking Evaluating Frontier Models for Dangerous Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T22:21:46.938158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:21:46.938158Z digest=sha256:6c4877ca405bf1ce172489d77f61a3656fe0d426e0d0d16773505dff8849a472

Observation 84930135-d3d2-41d7-a0ea-024504b98d5c · inbound

Towards Data Governance of Frontier AI Models cites this paper.

Towards Data Governance of Frontier AI Models Evaluating Frontier Models for Dangerous Capabilities

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:54.204725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:54.204725Z digest=sha256:567b15c72107c2507b9c9f40c6dddef6fe29c4bae5d990c2a91ab38854bf75f8

Observation fa14d73f-3b98-467b-98df-ed70fa9be0ee · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Evaluating Frontier Models for Dangerous Capabilities

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:22:01.671094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:1521e1dfcd46acd3bc395e299253727029042d112cf19df21cfe802721f421de

Observation d7d3fc15-ed60-4b0e-984f-dca0c53661e7 · inbound

What AI evaluations for preventing catastrophic risks can and cannot do cites this paper.

What AI evaluations for preventing catastrophic risks can and cannot do Evaluating Frontier Models for Dangerous Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:20.991463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:56:20.991463Z digest=sha256:96bdd9f5b168c169f764af6d793df04cf8f12aaf0f021dc99145ce7d43fe0828

Observation d4832d0c-7de9-4495-92f3-9564d4a34d9b · inbound

Frontier AI systems have surpassed the self-replicating red line cites this paper.

Frontier AI systems have surpassed the self-replicating red line Evaluating Frontier Models for Dangerous Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:37:44.732532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:37:44.732532Z digest=sha256:d840f0ac60dc2f5d756236c676c75fa45ab071f8c490535717edab40e8c5fa26

Observation b08d6fd2-6db4-4753-ab21-4885354ae96e · inbound

Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations cites this paper.

Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations Evaluating Frontier Models for Dangerous Capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:43.131236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:43.131236Z digest=sha256:57e549ff314e17d4f3c36d42a02533c280b474f2c5d9d8133853b4920fd15af0

Observation f85bd960-a8cd-4ffc-8e96-fbaf221f5797 · inbound

OpenAI o1 System Card cites this paper.

OpenAI o1 System Card Evaluating Frontier Models for Dangerous Capabilities

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:42:39.642455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-23T06:39:44.542350Z digest=sha256:ddf507cdd0c9c87955f498dfdf57ac9e43fbcdcd7ecd035194f763321a53948a

Observation 28512f1c-7477-4b92-a651-cde7deb30b94 · inbound

Authenticated Delegation and Authorized AI Agents cites this paper.

Authenticated Delegation and Authorized AI Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T19:50:42.581504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:50:42.581504Z digest=sha256:c55c3e633d60a221a8b9d8b8c8b0ad27bfa35afd5bf43a2c0619dcd18fb8abd9

Observation 54ab1573-da71-4eb2-a70f-7d5bdffbc599 · inbound

The AI Agent Index cites this paper.

The AI Agent Index Evaluating Frontier Models for Dangerous Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T14:48:35.492465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:48:35.492465Z digest=sha256:eff09ff3ada6cb258d6884e75ecc394f9bb9dfd1ac3ef63db1f754bd5fa319a8

Observation aa0b066f-1fcc-4b5a-812a-f607318d8510 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Evaluating Frontier Models for Dangerous Capabilities

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.576679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:82ecd7233f872857156dffd98f9395ad5cefac8f1d07fff40962b8213223c70d

Observation c7c49b1b-fc83-4f17-bc1f-82192cf068a2 · inbound

Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures cites this paper.

Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures Evaluating Frontier Models for Dangerous Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:34:05.033901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:34:05.033901Z digest=sha256:124736295101faa21efd3fb07295fbc2554354cbc8ef850208adb49c09bcdb64

Observation 60966de0-cb01-47a8-9c2a-e3f296d50abb · inbound

JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift cites this paper.

JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift Evaluating Frontier Models for Dangerous Capabilities

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:57:58.287237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:57:58.287237Z digest=sha256:5c6eebd77bf0ea55547281064972d5dcbc57f646f20e10200b17d44d21386ad3

Observation 49cc80c4-77d1-44db-b5ae-09685e1e801c · inbound

Real-World Gaps in AI Governance Research cites this paper.

Real-World Gaps in AI Governance Research Evaluating Frontier Models for Dangerous Capabilities

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:53:11.969029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:53:11.969029Z digest=sha256:9231b1dd9a16269640a941b78dcc1853c37ae37bef104072a74fc85fcefcf772

Observation fff4a8ad-7dad-4e6c-9731-1915a374fa2c · inbound

Evaluating Frontier Models for Stealth and Situational Awareness cites this paper.

Evaluating Frontier Models for Stealth and Situational Awareness Evaluating Frontier Models for Dangerous Capabilities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:22:42.676750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:22:42.676750Z digest=sha256:ff691a811db31c7aec09d2dfb5e19baa44cbaf1b11743d3f4d60d35a31076068

Observation 4badcd9d-26c5-4958-8f5e-052cd73a5167 · inbound

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions cites this paper.

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions Evaluating Frontier Models for Dangerous Capabilities

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:30.430189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:27:30.430189Z digest=sha256:21d43f07a11394c88017eff411951956e5f82026238ef689b8c4b0e1af3a3cd0

Observation 1f46d63e-ab30-4b48-b0fb-6d5d6d486985 · inbound

ACSE-Eval: Can LLMs threat model real-world cloud infrastructure? cites this paper.

ACSE-Eval: Can LLMs threat model real-world cloud infrastructure? Evaluating Frontier Models for Dangerous Capabilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:00.711731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:00.711731Z digest=sha256:52559f0fcf10b95365051dfc8bce5f8e859c107104ec5433a29b61d0910b9646

Observation 291bbf65-44a7-4f41-afc8-e085cebc48d4 · inbound

Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks cites this paper.

Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks Evaluating Frontier Models for Dangerous Capabilities

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:59.431537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:28:59.431537Z digest=sha256:e0583357cc25b46998f003c445819b808f5bcf0721cbfcaa677411d2193161e6

Observation 316b0a8f-dd34-417e-8046-86a1e7ec13d6 · inbound

Benchmarking Misuse Mitigation Against Covert Adversaries cites this paper.

Benchmarking Misuse Mitigation Against Covert Adversaries Evaluating Frontier Models for Dangerous Capabilities

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:32:14.642218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T10:29:05.104520Z digest=sha256:b7cad6ac5dce9cb934700cf356a47048257f96214f29343b75186238c3990853

Observation 40c2ff1d-0346-4269-9b7a-7b608aa67487 · inbound

Data-Centric Safety and Ethical Measures for Data and AI Governance cites this paper.

Data-Centric Safety and Ethical Measures for Data and AI Governance Evaluating Frontier Models for Dangerous Capabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.884266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.884266Z digest=sha256:1368ef1882a20fde654c058afbbed59278d2dee14dbca4a3db4da9ff13038a75

Observation 64cb3526-2168-4d5e-96d6-e947c741a752 · inbound

Contemporary AI foundation models increase biological weapons risk cites this paper.

Contemporary AI foundation models increase biological weapons risk Evaluating Frontier Models for Dangerous Capabilities

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:20.085006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:20.085006Z digest=sha256:2ae46853f910e51f50fc1a4e7176eab49b6026ec8dfb7d1494ebbac73d684c33

Observation ff62de0c-4161-493e-8c1b-09bf990d1e6f · inbound

International Security Applications of Flexible Hardware-Enabled Guarantees cites this paper.

International Security Applications of Flexible Hardware-Enabled Guarantees Evaluating Frontier Models for Dangerous Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:47:59.062215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:47:59.062215Z digest=sha256:f73db28c02339ebc927813a3ad45ff15d5514899f445f129951fec343b2fe386

Observation 4a56befd-c9be-42e5-9209-0b30a053139b · inbound

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents cites this paper.

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:52.954857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:52.954857Z digest=sha256:5aa43caaafcc0b60ab9777ce68d13c946f562534d5b1b5ffa04ee08cd007cb9f

Observation be9f9fa3-1e59-4bc6-b345-e655913e4d13 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Evaluating Frontier Models for Dangerous Capabilities

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.931876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:d225da8dc52599ff3de34ce44486fdc9aa0e8e0f6f16efbbb22a1dcc9dd12aaf

Observation e107757b-5e07-42fd-a604-df71d28c0cf0 · inbound

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework cites this paper.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Evaluating Frontier Models for Dangerous Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:37.129859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:37.129859Z digest=sha256:9b15d2e5a1bf6aa89f6ffc95c673f196de073fa9a68c4232f18a389e77c35551

Observation ded1746d-a757-43f2-ae1c-421e4d68f658 · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Evaluating Frontier Models for Dangerous Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:22.926909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:22.926909Z digest=sha256:d481e10d5faa3c06ed7a3b1e3e3f375b5d2a71cd04af787fbd89d2f0ea4accea

Observation cdd23794-a56e-411d-9767-1858edc9d8c5 · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:53.446112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:53.446112Z digest=sha256:707932b57a04ad68525a5fa0429ec10c1a6c4894f970f53825084a483f412f78

Observation 7c944593-4e77-41ff-8120-20da83a075f2 · inbound

Private, Verifiable, and Auditable AI Systems cites this paper.

Private, Verifiable, and Auditable AI Systems Evaluating Frontier Models for Dangerous Capabilities

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:59.371101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:59.371101Z digest=sha256:94c9e28d4705a36950fbbfe7f24799d82eeada19c114062af69741e11ce97412

Observation 93071eb1-00da-4785-8cae-d7d45886470a · inbound

Estimating the Empowerment of Language Model Agents cites this paper.

Estimating the Empowerment of Language Model Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T14:53:17.688392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:53:17.688392Z digest=sha256:248cd55b86ed5d8f948a90e3349349e9615afed6d9dea1861625d32af9ad4d18

Observation 98004cc5-5adf-4f1f-b56a-8b3576daf855 · inbound

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users cites this paper.

Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users Evaluating Frontier Models for Dangerous Capabilities

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:23:39.955248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T23:22:42.431997Z digest=sha256:8318fe56843627e14fcd3c6edd7fd36c24527a505bec32e2f4ce0fae670ceab0

Observation 9fa259d1-9099-4208-a1de-04ac5e511324 · inbound

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems cites this paper.

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems Evaluating Frontier Models for Dangerous Capabilities

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:46:36.241926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T20:41:49.137745Z digest=sha256:3ae03123ba4960d0764c693f6332900820c7bdcdbe9e08d49642c2096def5f51

Observation bf4cfde7-10f5-433a-9099-a2b0e77d19be · inbound

From Disclosure to Self-Referential Opacity: Six Dimensions of Strain in Current AI Governance cites this paper.

From Disclosure to Self-Referential Opacity: Six Dimensions of Strain in Current AI Governance Evaluating Frontier Models for Dangerous Capabilities

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:00:22.250927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T11:55:50.822764Z digest=sha256:63ae94ff7c6b036cfea2210971f7dc5a9b119583710e7b60c55188c64eef2c67

Observation 3fb4d26d-f3b0-4c3e-ad66-36d994d3cc4a · inbound

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge cites this paper.

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge Evaluating Frontier Models for Dangerous Capabilities

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:34:47.049775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:32:59.120345Z digest=sha256:4f34f4c79bf85a0e514d3f53338582347228fa096025ad51c494e4471ee2ebef

Observation 8a653bfd-6ffc-4603-9442-4775de29256c · inbound

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? cites this paper.

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? Evaluating Frontier Models for Dangerous Capabilities

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:09.009924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T03:04:14.130717Z digest=sha256:49a333da3d273920da379e62de80e3b56492b99e75c48476eb784086f0d613aa

Observation d378f5ed-b3e3-4af3-ab17-32288ec8738b · inbound

Rollout Cards: A Reproducibility Standard for Agent Research cites this paper.

Rollout Cards: A Reproducibility Standard for Agent Research Evaluating Frontier Models for Dangerous Capabilities

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:02:23.587743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:57:48.614881Z digest=sha256:f847209140f1504e1af047b1dc6a26fae5a23eec89a2bc0b57e5296e352358da

Observation 0672b9c5-6fb0-439c-9e09-65e98af93209 · inbound

Measuring Safety Alignment Effects in Autonomous Security Agents cites this paper.

Measuring Safety Alignment Effects in Autonomous Security Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.916500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:cec94823f6dca6356b7f8633bf3aa1d4f482c6c49523e75a0abd00930902316a

Observation 2e85b7ac-258d-4977-8ac5-73376046bb6f · inbound

Comprehensive AI governance requires addressing non-model gains cites this paper.

Comprehensive AI governance requires addressing non-model gains Evaluating Frontier Models for Dangerous Capabilities

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:31.663670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T08:11:23.860723Z digest=sha256:bb558e27a3b3f3f09f782b56ed52dcc21765537dc0fcc266822d67183de947df

Observation b787e572-707f-449d-9223-02a0fb9e8fcb · inbound

Solipsistic Superintelligence is Unlikely to be Cooperative cites this paper.

Solipsistic Superintelligence is Unlikely to be Cooperative Evaluating Frontier Models for Dangerous Capabilities

Reference 199

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.746920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T10:00:41.374121Z digest=sha256:2c0504876f572feeb65e19ad902bcc7927bbb1c2948f4db0e8a22749590f41d9

Observation 80ecbb19-e52f-46ae-9021-c419ae918ce2 · inbound

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing cites this paper.

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing Evaluating Frontier Models for Dangerous Capabilities

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.572672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:17:01.173495Z digest=sha256:656900519493dc07abbf5794cacb5322d8b9a8cdfb47f0dc0c3cc2349edd33db

Observation 238c4677-edad-4e28-805c-67e7e959149c · inbound

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems cites this paper.

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems Evaluating Frontier Models for Dangerous Capabilities

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:38:33.531225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:25:17.298114Z digest=sha256:5f2e11b01f10b4bb8d0a1437400a5922da5c07ccf7435d4084b030d39adbbd83

Observation 498984e8-a9ec-4251-80d7-a6089f9abafb · inbound

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems cites this paper.

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems Evaluating Frontier Models for Dangerous Capabilities

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:38.126424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T11:27:22.436981Z digest=sha256:fced67b45578848500a5b66fa0653ad08f35b142618db76012ea373ebcfddfaf

Observation f4441bf2-4b62-4cd4-9a17-81337fcf7b1b · inbound

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework cites this paper.

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework Evaluating Frontier Models for Dangerous Capabilities

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:59.996827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T23:42:20.304205Z digest=sha256:6abe03a5b9fdf320d2802a0eff2aa55f850f6d071f7d86d837a3e3e0977a1a12

Observation 2943cd42-a5e2-4f2c-b94b-4cd3e36984e8 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Evaluating Frontier Models for Dangerous Capabilities

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.646309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:70abc6ed0b7a151d745ac51eb005d4d2a61489108293febfb0f2205bfbe6ef9c

Observation c0a59a1d-43da-4459-8037-1aed539ea542 · inbound

The Oracle's Gambit: A Game-Theoretic Framework for Responsible AI Release cites this paper.

The Oracle's Gambit: A Game-Theoretic Framework for Responsible AI Release Evaluating Frontier Models for Dangerous Capabilities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T01:32:59.751620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:32:59.751620Z digest=sha256:fffccd44bf68de6b2e42fe1151611aa4daf37758de813d2a4413aed2e702dcb9

Observation 0746f56a-9ea6-4c5c-9ca1-9d2e0f48bd6a · inbound

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring cites this paper.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring Evaluating Frontier Models for Dangerous Capabilities

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:40.994995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-10T00:52:47.537142Z digest=sha256:95dac0f5d35029fa8f995c28b53cf068acd0fbed51a112664bd676cf88150138

Observation 1cc1b510-a081-4d37-a465-c8fa80c98ce4 · inbound

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI cites this paper.

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI Evaluating Frontier Models for Dangerous Capabilities

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T04:12:58.764877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:12:58.764877Z digest=sha256:6794623b2773ebade48350707dbed5e4aa4b7d8f0e39b962a68cd2eb64318737

Observation 4b1eb21c-0605-4abb-a4cf-76023772b512 · inbound

Accountability Asymmetry and Structural Trust in Autonomous AI Systems cites this paper.

Accountability Asymmetry and Structural Trust in Autonomous AI Systems Evaluating Frontier Models for Dangerous Capabilities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:51:31.797330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:51:31.797330Z digest=sha256:0ce4fa6981ddfad47a60ed8c5fc211bdbf56f1e232ad93b28524908d51ad2bb9