Pith. sign in

Paper Citation Record · LEDGER

AI Security Leaderboard: Methodology, Results and Minimal Standard

As of 21 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.03070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03070 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:04:40.936182Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 486f349b-b9d2-42d9-8d97-0c6bd8f05cbd · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

AI Security Leaderboard: Methodology, Results and Minimal Standard gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.800351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.800351Z digest=sha256:2ed650bf57641a7ab1aefa8b7a3f40361b9c2368635b8f0d76b95522d0dc1c4d

Observation 00822f6b-9cd3-486a-a667-1ad9952ecd6b · outbound

This paper cites Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.717747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.804895Z digest=sha256:0f9244b198cc2f090f328821a2a0199f1a44843dbc0b1059970c9e22bddea408

Observation 921e6ce0-e3b2-4dbe-9683-f8857cdb3622 · outbound

This paper cites Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.708262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.808590Z digest=sha256:762d371204a11440472f32547a18576ac4b6679d44266b834e8fe58cbd60cb5a

Observation 68433ac3-357a-459b-b927-30ab800cbf9e · outbound

This paper cites Constitutional Classifiers: Defending against universal jailbreaks, February 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against universal jailbreaks, February 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.699247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.812066Z digest=sha256:d471a97dbd760d19aae1ad7aa8621747f633e3722ebec3901dc957ff48ecaf78

Observation 5e9af83b-b20c-4971-8b95-8ac4c98bfb4e · outbound

This paper cites Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.689793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.815425Z digest=sha256:6a788130e0dc0a9f586031cc8e7d6f94addd58639f981dadd401d468e6d7b798

Observation 8ebe1614-0dac-47a8-b3d5-f23ef00d43fd · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

AI Security Leaderboard: Methodology, Results and Minimal Standard Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.818991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.818991Z digest=sha256:aaaccd16c4cdc744c332e854ee654710fc150a37fbd0d002e2c89a090d8fa0a3

Observation dbe28a52-b414-451e-999a-f397cc2723b0 · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.822825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.822825Z digest=sha256:170d6c9c5038df7f437d471f96fcecab2c53de4237a500eb71ed742afa5974c6

Observation bd399f56-0609-4170-915b-3919fcb5cd6c · outbound

This paper cites Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.825990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.825990Z digest=sha256:0d0ed3cecda3286bd95b356cc5d2189ab8f97d7f540fbd9166efb294b4959d85

Observation d4c42b4f-2a40-4877-8248-ec8866a67713 · outbound

This paper cites Boundary point jailbreaking of black-box llms, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Boundary point jailbreaking of black-box llms, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.829056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.829056Z digest=sha256:e2d00bf6c914ea6111a11d51bdcbd55827cb5bc58496e468c9cee2243e31b86f

Observation 8e302ba0-1a01-4499-8b35-d7db3113287f · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.679664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.832135Z digest=sha256:f723e4cfc820d2a159db2f044caaa347fad04db5d4a760bcfdf4da511cf838f1

Observation 974e515b-6c36-4a76-b531-59b2b6b748e1 · outbound

This paper cites The safety gap toolkit: Evaluating hidden dangers of open-source models.

AI Security Leaderboard: Methodology, Results and Minimal Standard The safety gap toolkit: Evaluating hidden dangers of open-source models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.669906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.835276Z digest=sha256:3fd7317b0f3a78f718bd4004baa0ae590d510940269baf83c44761eded69a54b

Observation c2cc9033-fcb5-404e-b5bc-6a51103e706d · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.660329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.838399Z digest=sha256:541475d6a66df2a7bced6c58cd793f50cacb7f8ba47d715442c75f84e1cbe222

Observation 50f5ab3b-2cb7-4d9f-b6dd-cb0e3db16493 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.841373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.841373Z digest=sha256:d088aff1f5d1da4150dfd1dafccf476e84da5df479b7176849d3f2390f60726b

Observation 5dac27c1-ce00-407f-b11b-92a491cf6dca · outbound

This paper cites Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark.

AI Security Leaderboard: Methodology, Results and Minimal Standard Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.844948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.844948Z digest=sha256:0797b5f57f4377035d62e2ec0da2642ea0d88592de1f1f56d4e3e48a91b7c5f6

Observation 774cb533-c83e-4232-9a8b-840c9f9bfdc6 · outbound

This paper cites Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.650799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.848365Z digest=sha256:245a87aa1245d7d77331d25f09f3896812c15a5b4029e7ae0f3ec033fb3050cf

Observation ac7b7765-dcc5-4c5c-a68d-0f42c6539bd1 · outbound

This paper cites God Has Helped Us, and So Will AI.

AI Security Leaderboard: Methodology, Results and Minimal Standard God Has Helped Us, and So Will AI

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.641029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.851461Z digest=sha256:6c17b9df5ac1253f863d5e190db6423ed691a59e21a818558d190b5607c780ed

Observation 5c60ab4c-7a46-4812-8b03-d95718c6c93e · outbound

This paper cites McKenzie, Oskar J.

AI Security Leaderboard: Methodology, Results and Minimal Standard McKenzie, Oskar J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.854468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.854468Z digest=sha256:e06178ac3886578f8b0cf1e4f41d49442c4167d82915ddd452ef5b343915e121

Observation 5352fe49-0ae7-4ba0-be5b-86492d0ddb08 · outbound

This paper cites Common elements of frontier AI safety policies, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Common elements of frontier AI safety policies, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.631341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.857439Z digest=sha256:35ddce1874cee0f7cc3d52fc9e1b5f1f53107469f053e5846ad999eabd3613cc

Observation b1277843-1da5-4094-9db4-f6d75e1bfed1 · outbound

This paper cites The Jailbreak Tax: How Useful are Your Jailbreak Outputs?.

AI Security Leaderboard: Methodology, Results and Minimal Standard The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.860705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.860705Z digest=sha256:530d397310216aaf751c9e9e81b793328c70b4fcc9b8a4bb8f41a21982e70ced

Observation 46d6afd3-527e-4ea0-bb38-de08a31a3502 · outbound

This paper cites ChatGPT Agent System Card, July 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard ChatGPT Agent System Card, July 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.621501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.864122Z digest=sha256:f91b4ce9362a2b1e2774c0c1e6649e138693ca072de55956e6d9936a1d62b6e4

Observation 67ef4c4d-23e3-4bdf-9a3c-4d8d35ffc6a5 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.867237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.867237Z digest=sha256:657cd963120dd29ee05f3891833eaa80f1b4f095167844525142a1aa5f7b7e11

Observation 94c2008a-fb9e-4f1a-a5e5-861605ed6319 · outbound

This paper cites Exposing the systematic vulnerability of open-weight models to prefill attacks.

AI Security Leaderboard: Methodology, Results and Minimal Standard Exposing the systematic vulnerability of open-weight models to prefill attacks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.871031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.871031Z digest=sha256:c1593d8491371cb7567649b63f27480d2a6220749a6fc2dc222db5f1988228ef

Observation b8ef9201-8fba-4294-bb23-4a0ac23f70e8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Qwen3.5: Towards native multimodal agents, 2026

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.612012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.874332Z digest=sha256:38bfc6a0b13e56a13e9c52ddfa936a0af2507924d0cf8724f6eaf1ee27e43b00

Observation ffed2393-7a4e-41ac-aa69-45d8c2381186 · outbound

This paper cites Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.602321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.877689Z digest=sha256:4ef5f5cdb9dc071a1095158c4bc7be4e7680d8479d5cbdb11cdbe87edf79a0ec

Observation 8fe0585a-8534-4ddd-82f2-fbb94be7e2b4 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

AI Security Leaderboard: Methodology, Results and Minimal Standard The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.880760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.880760Z digest=sha256:9850563975529a0ff47616c2e256128b21c8a6e91b1e4e2d9ee9553c1c8ad72c

Observation d3f94878-4a23-4e07-b4ef-da0b4d1def24 · outbound

This paper cites Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.884319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.884319Z digest=sha256:3aa184f986a345947da422678d930e10474beb52b10766dd854bc7f5fc5642fc

Observation 89692e7a-d5e4-4df7-8875-d5e7e26c8edc · outbound

This paper cites Jailbroken Frontier Models Retain Their Capabilities.

AI Security Leaderboard: Methodology, Results and Minimal Standard Jailbroken Frontier Models Retain Their Capabilities

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:04:40.979412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.887782Z digest=sha256:4ff1db520697cc26b19cbf70dedbc0fe2e5e75a59f3e6e7bb4365544d39e3909

Observation 677b28d0-f0f8-4248-88f9-255d12d54f4b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.891362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.891362Z digest=sha256:843e83cf43c707ff4c620a1e8efb330695b449f17378b89619a6e5578a5665f4

Observation 407c913d-b43a-4987-b53d-247c77fac7eb · outbound

This paper cites Sorry,” “I can’t help with that, but….

AI Security Leaderboard: Methodology, Results and Minimal Standard Sorry,” “I can’t help with that, but…

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.591935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.895326Z digest=sha256:a91031a7eb97bd6012d40e906d79029d27c7c132d329052c55665c58bf684870

Observation 273bbf32-5b57-4029-9f26-a0ffb085e131 · outbound

This paper cites not jailbroken.

AI Security Leaderboard: Methodology, Results and Minimal Standard not jailbroken

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.582787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.898975Z digest=sha256:42e62a7f5bbf0142e59b27bcd3cb760bff319f1d1286ac0843926aee17e317d1

Observation 01daf259-7b40-443d-a04f-4bc89046a70b · outbound

This paper cites jailbroken (1).

AI Security Leaderboard: Methodology, Results and Minimal Standard jailbroken (1)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.574059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.902515Z digest=sha256:5989d6917c68613c883f557fe6c48bdfb780db45c61400deefa935f47fa1f2fd

Observation b3de8b6a-0ba0-4112-a24b-acb5836644cc · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.565345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.906031Z digest=sha256:c68915859e70a57822b601ae5129cedb78fd73873dcd0f35d730b92738b19a73

Observation e870fbac-2b7d-433c-99c5-f8fe535c39f4 · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.555538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.909296Z digest=sha256:09acd64e69f6dde2ab7e30ccec80c290850089de826a227e9c735e6f5dee9b1d

Observation ce5d3ada-f4f9-4bef-a33a-18668093c0ff · outbound

This paper cites the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0).

AI Security Leaderboard: Methodology, Results and Minimal Standard the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0)

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T01:04:41.545319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.912861Z digest=sha256:467f881cf7fefbcc3cb07fdb8ed009fd500025d6269abdafd1798550aae6d569

Observation 051e85b6-83a9-447b-a859-cf1d9c650a1f · outbound

This paper cites The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price.

AI Security Leaderboard: Methodology, Results and Minimal Standard The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.535433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.917008Z digest=sha256:3997a340d9397c07765f2fbbea8e3b3238ad957946844decf98630fb640128b8

Observation f4cf8811-900f-4081-9503-7d1530ef92f7 · outbound

This paper cites If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops.

AI Security Leaderboard: Methodology, Results and Minimal Standard If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.525030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.920005Z digest=sha256:1460c4aa1d65678edb01e7ec53e0b7bf619230eea4ff9d1556fb44357f57af52

Observation b63ec2b0-961e-4470-96ae-ee7e7d6f1cf7 · outbound

This paper cites Suppose we find ten working jailbreaks against each of two models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Suppose we find ten working jailbreaks against each of two models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.514526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.923218Z digest=sha256:d0cad4c7d3375983bc7932358577867cc81be442a03f11402b5f9f5612d1ca5d

Observation e1511973-57fb-4595-a23c-3e1d34c46a7b · outbound

This paper cites Finding a jailbreak that works once is easy; proving it works reliably is expensive.

AI Security Leaderboard: Methodology, Results and Minimal Standard Finding a jailbreak that works once is easy; proving it works reliably is expensive

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.504661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.926614Z digest=sha256:eeaf0c57944d22d47ad934d19f3127e5f2fbff76f4431b600ac1853ab925c898

Observation b7a4f3f6-6596-4d57-830c-85d2addbc8f7 · outbound

This paper cites A smart attacker does not run every candidate over the full sample.

AI Security Leaderboard: Methodology, Results and Minimal Standard A smart attacker does not run every candidate over the full sample

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.494329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.929731Z digest=sha256:218dc331bb83bc3b17676565106ac1aa38ad9ce905ffc1576fbb833cccc35a47

Observation 33117418-4b40-46b2-92dc-61512cbe38de · outbound

This paper cites Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain).

AI Security Leaderboard: Methodology, Results and Minimal Standard Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.484182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.932922Z digest=sha256:6684e77d14d9149727b03e580186bccd2a24a9087f7bf12977f9829f73ded8cf

Observation 16677385-b223-4e30-bbce-beb5e9ad92c4 · outbound

This paper cites Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%).

AI Security Leaderboard: Methodology, Results and Minimal Standard Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.473005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-08T01:04:40.936182Z digest=sha256:91925034946b9102b8b7de5962cd82ba123a5e13cf2a65a573387e213d312dd7

Pith citing papers

No inbound Pith citation observations are available.