Pith. sign in

Paper Citation Record · LEDGER

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

As of 10 August 2026, this Paper Citation Record lists 100 of 117 outbound references and 1 inbound Pith citation observation for arXiv:2604.16286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16286 v2

Coverage vector

measured 100 of 117 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:31:53.114772Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:49:57.400820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 117 outbound references displayed

  • verified exact0
  • verified fuzzy82
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a71e41-587a-4fcc-ac10-de8d404e1291 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.170855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8e23f7c19d326f8527b43b1c16b721930fa044b28f3568af10856aa5eec54453

Observation 9546270b-95b0-4968-99b5-a26f0ca2fe4f · outbound

This paper cites RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.992059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:74e20f26732795b37080512087cc0bb008fa70352a549cf25a390dfdc4cb3fb5

Observation 250e1c7f-0b4f-4cf6-95f1-b07a03d0b67a · outbound

This paper cites legitimate AI researcher.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases legitimate AI researcher

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.079244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:49e9ef6eed53d6a2614e2dc36eb854fda8fc37858d8cb750a50b157edb7ecfa5

Observation 609adfa3-a98c-4c78-af36-9fc85849dd46 · outbound

This paper cites The adolescence of technology.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The adolescence of technology

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.051102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:6d3b971d7193d3ed7c8d08e1e8789b1aca81fe9bea3ec93bbe55c12074e897ec

Observation a7426852-beb8-4b09-aa42-6edf1ceaac03 · outbound

This paper cites Bowman, and Evan Hubinger.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, and Evan Hubinger

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.110431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:178696ebc3212054c2a5938cb414603771f6b1daeeda505d12cf3ec506b961ed

Observation 1cdaae8b-f94c-4ba4-b870-ac76305e5aba · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.271230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a723e254eb792133109836b6174522b83818c28818ea00e7ff3c6c7d20112844

Observation 3ac2e9ff-bb0e-48a9-bb85-de220662d11e · outbound

This paper cites Stress testing deliberative alignment for anti-scheming training.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Stress testing deliberative alignment for anti-scheming training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.044333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:854e9c27271dc9392bda831ac64c77a2890032433ccb64d28e7ee6f4aee9c2e8

Observation c6fa2eef-eb05-454a-9e7e-8ef2729715c2 · outbound

This paper cites Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.233981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:42cef7141ed8c06819d7985997ce8d4a5746f4bbfc4329df709e4261cb5074b4

Observation 851cc0c3-4c74-445d-a117-5a7a69ee5e3c · outbound

This paper cites Sabotage risk report: Claude Opus 4.6.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Sabotage risk report: Claude Opus 4.6

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.252781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:58dd2f2aafe8f57b7a7abfc8ad3f14c3181971052c6e89012cad1f1b9b8e861a

Observation df986c1e-01f9-4d1d-b067-504c3abd10a3 · outbound

This paper cites Alignment risk update: Claude Mythos Preview.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Alignment risk update: Claude Mythos Preview

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.028922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a128209680709ab2db61973e8cb8eead10b204f5edf6f5e98a1f9a97acfd5eca

Observation 61f3b045-2038-4911-aa63-bfa8ff8d4167 · outbound

This paper cites CTRL-ALT-DECEIT: Sabotage evaluations for automated AI R&D.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases CTRL-ALT-DECEIT: Sabotage evaluations for automated AI R&D

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.936444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:414155a56dcc21d909d364c577511235756e809284e128b6f94518a832d0714b

Observation 70f70870-772a-4922-be76-5514b5f564d6 · outbound

This paper cites Bowman, and David Duvenaud.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Bowman, and David Duvenaud

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.153863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9e43ef480d71dbe95a03f09dd37270de866f3a750a84cb7039c6a723272ef100

Observation c4aef78a-f00d-4b59-bca6-cd407f0131db · outbound

This paper cites Subliminal learning: Language models transmit behavioral traits via hidden signals in data.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Subliminal learning: Language models transmit behavioral traits via hidden signals in data

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.121546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:972c21482f33661d37848581c828f1ab0daf7653953218a9c73a0410d49d2039

Observation 5da021e6-050e-44a4-b80e-4fc03b2083e6 · outbound

This paper cites CoT red-handed: Stress testing chain-of-thought monitoring.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases CoT red-handed: Stress testing chain-of-thought monitoring

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.865417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e9631da0638e4a12c9d795265077968505566e9ab4f6595a18b301669d30aaff

Observation 530e3ae8-1613-4772-bd82-1cb86706b6cb · outbound

This paper cites Disentangling feature and lazy training in deep neural networks.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Disentangling feature and lazy training in deep neural networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.965334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9a92bb2f28ea4ceb3e85f6c7e655073397954fed1ced882583c662c5c85c7212

Observation fbb7052b-4e59-4d32-8c3d-497e44132958 · outbound

This paper cites Steering evaluation-aware language models to act like they are deployed.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Steering evaluation-aware language models to act like they are deployed

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.003796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d5d80ec8af388588637198b2fb708327e821c124dfc71a2637e24112a11d3023

Observation 073f4334-c7c5-4767-9110-43576e4d5fa4 · outbound

This paper cites Lessons from studying two-hop latent reasoning.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Lessons from studying two-hop latent reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.274882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b3d27d9008dc1758193382e081eca5b456b9e3abef7e6fc2419429f72635afbd

Observation 741105fd-1611-4683-a45c-3f8c927e4a16 · outbound

This paper cites Copy suppression: Comprehensively understanding an attention head.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Copy suppression: Comprehensively understanding an attention head

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.018013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:11e9e552991833c73bbec0e9ae7be062af8880488f7ca7c4ab7fec24f09bf926

Observation eff825d5-ab34-4918-957e-a6a2aebf687e · outbound

This paper cites Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Hidden in plain text: Emergence & mitigation of steganographic collusion in LLMs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.089937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:94ecf5a89559f9e56acce30b6aa410728b12c085917cd75ca8a3fec69ad85be1

Observation b4cba413-d23e-4481-b869-ce7fa7538528 · outbound

This paper cites Multi-turn jailbreaks are simpler than they seem.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Multi-turn jailbreaks are simpler than they seem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.241271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:25d4ab74c6897d1949f300d98603662c7bbf0599e5d45d90f72db04d1e91f103

Observation 9a6c4e4f-65eb-46ee-8b76-bbf13805b0d0 · outbound

This paper cites AI control: Improving safety despite intentional subversion.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases AI control: Improving safety despite intentional subversion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.203553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f7957d00df9cbee246818e9fced073f657db9cb343ca6c269a50084220b818e5

Observation c5bbdc21-af17-421f-ad28-426a13201d5b · outbound

This paper cites Ctrl-z: Controlling AI agents via resampling.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Ctrl-z: Controlling AI agents via resampling

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.911054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ae38f80d2aeade5eb4fdad5356b0c3959244215e05ca883ceff3c9e3108e339a

Observation cf2397ff-30d1-49c0-84f6-fee957ba15ac · outbound

This paper cites Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.157825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b3a5c1c9896e5804a7058c82d6fd9639b4ea7e2dcb2a6385886197b46b56ce40

Observation ece28bf3-d01f-4d8f-9947-f8e95b2c01da · outbound

This paper cites Training fails to elicit subtle reasoning in current language models.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Training fails to elicit subtle reasoning in current language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.149787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:79c8ddce300b215cac8978aaed2d7308d76971a9e282367396335b2a9a017a0f

Observation 57b1ced0-1c33-4c3a-b266-f8cf301dc5d4 · outbound

This paper cites Factor(T,U): Factored cognition strengthens monitoring of untrusted AI.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Factor(T,U): Factored cognition strengthens monitoring of untrusted AI

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.174783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:79adb34e5378eb4585f495227ec3c5ef00c9254e17ea51cb349831bcb8957256

Observation c21eb7a6-fc13-4f07-9e35-88b112cfd7d9 · outbound

This paper cites Basic legibility protocols improve trusted monitoring.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Basic legibility protocols improve trusted monitoring

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.057860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:55c9d5cda2840e6659438bd55a0146ae7373a276f3d862dd01fddcfcffed6078

Observation 2fc9ebe3-54ad-4dd2-8630-eaebf179a9bf · outbound

This paper cites BashArena: A control setting for highly privileged AI agents.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases BashArena: A control setting for highly privileged AI agents

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.136090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:dd1234df80f7d1a437074484c03fb1dddae4d19c53e68955465901fb97ae96f0

Observation abbc1177-2322-4e4d-b525-bf95a8a268ee · outbound

This paper cites SHADE-arena: Evaluating sabotage and monitoring in LLM agents.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases SHADE-arena: Evaluating sabotage and monitoring in LLM agents

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.881897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9341a9d5f5d1377f86d165e88054429d6e104a60a649a74be4c37e2622449c0d

Observation bc882448-7527-4698-9efc-0aff6019229d · outbound

This paper cites Adaptive deployment of untrusted LLMs reduces distributed threats.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Adaptive deployment of untrusted LLMs reduces distributed threats

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.893643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e70d8eb6468c16bae3835986f098b8823460c337bf49aced5f7c002cb7131295

Observation 922a6c75-8be2-4e43-9375-cf98bdf81eff · outbound

This paper cites Brown, and Francis Rhys Ward.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Brown, and Francis Rhys Ward

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.210586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:471cafcdfcbae0d074ee39d16a56b2e344c783275c021936156f77310088b615

Observation a315946d-d243-4c95-bc6a-85bc44632058 · outbound

This paper cites Sandbagging in agentic ML research.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Sandbagging in agentic ML research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.900862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3fbc8510784eeec1c00a2b7314d637c75f0cc32a275cd31edb08f98da2a62a32

Observation c4b3deac-0e60-4e3f-a947-a36cbbdcb9c9 · outbound

This paper cites LLM critics help catch LLM bugs.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases LLM critics help catch LLM bugs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.869145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b95ce84368a15aa4f09f7bb9da12a87d53c51481a31eddccd21b6736d0942a01

Observation c02770aa-2ed6-4a8c-bd78-0f4db43a8fc3 · outbound

This paper cites A practical approach to verifying code at scale.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A practical approach to verifying code at scale

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.105195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:47916f44f530a6ba27995e2150c9083f00378219bcac4ed4d182685acac6ffa6

Observation 74e97300-6054-43bd-a978-ea0bddc71e24 · outbound

This paper cites IRIS: LLM-assisted static analysis for detecting security vulnerabilities.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases IRIS: LLM-assisted static analysis for detecting security vulnerabilities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.164721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:2a962d122f8b0fda894c01d2e141f9c20132936bfa8d7c35b41f088fb49bbfe9

Observation c99a00c8-d4c9-4d7c-a8f2-cbae14bc0a30 · outbound

This paper cites JITVul: A benchmark for evaluating LLM-based vulnerability detection in real-world code.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases JITVul: A benchmark for evaluating LLM-based vulnerability detection in real-world code

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.932832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a96dadf55ead4d232f42ecd3fcc45040de369bef91e256b564e33dc645d31b23

Observation 7a22f65c-55f6-4879-bee9-ce96d409f2f9 · outbound

This paper cites From large to mammoth: A systematic evaluation of LLM architectures and quantization for vulnerability detection.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases From large to mammoth: A systematic evaluation of LLM architectures and quantization for vulnerability detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.007309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:07928f90dc0df5831ea2157b3b3d942cbfee93298a0d62a336079a8ba466eaae

Observation 1c68cfd4-44c3-4d80-8f4b-d6540dc76ce0 · outbound

This paper cites Mythos preview cybersecurity report.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Mythos preview cybersecurity report

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.064871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a59f66b2f7070458a713e47d319233f00e150a5cbbbd7c1dbb2da92dd48f3a76

Observation 2d54058e-4653-40e0-8113-07e0aa6dbca9 · outbound

This paper cites Leakage and the reproducibility crisis in ML-based science.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Leakage and the reproducibility crisis in ML-based science

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.093497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a9d248d3879e21dbb23b76974f701a77737b64e6dfa23d13aa8fd22568001d0f

Observation be9ac55d-fb1d-4c26-b602-8e9e5f83d390 · outbound

This paper cites Are we really making much progress? A worrying analysis of recent neural recommendation approaches.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Are we really making much progress? A worrying analysis of recent neural recommendation approaches

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.982224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e9a88e3f547b645147894ae2fa7f21d6a953b7c79e20130efda441151975eb5a

Observation 361026f0-c738-49fa-9f22-5fb0b85fa144 · outbound

This paper cites Are GANs created equal? A large-scale study.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Are GANs created equal? A large-scale study

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.025455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a66e44b520666c5f4d7e195535714263f6a6d932cc5a639e2ccf95c495b0e38f

Observation e1fe6d93-79ff-48ba-8533-dced9200d656 · outbound

This paper cites On the state of the art of evaluation in neural language models.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases On the state of the art of evaluation in neural language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.897145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:1aebc9fac8a00100cd13a9b934d574427fc49efd582a324205220ee06b8a28c6

Observation 4252fa6d-23da-4ab1-ab58-e2cc1eb875d6 · outbound

This paper cites A metric learning reality check.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A metric learning reality check

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.889813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:0787eeb43ed258733d5451b053fe630e2b273b60c08c631c1efd6fd0b1c0a4f5

Observation 1dd0debb-6d4f-4ee3-81b9-35851b60a188 · outbound

This paper cites Pitfalls of graph neural network evaluation.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Pitfalls of graph neural network evaluation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.075734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:35f5ddb12e8b92e0f2bd1860cb6d8ce8eca28392e5a02780f8057db395a53c14

Observation 32d29912-5a49-4206-8cfd-1374426c5be4 · outbound

This paper cites What is the state of neural network pruning? InProceedings of Machine Learning and Systems (MLSys).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases What is the state of neural network pruning? InProceedings of Machine Learning and Systems (MLSys)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.040908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:2294db619b646ecb5ba95d6fc09fcba31a0b3f7a25cc559aff6be528c3ab779f

Observation c97c72a0-75ae-4901-bdfe-399485c59d0a · outbound

This paper cites Towards evaluating the robustness of neural networks.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Towards evaluating the robustness of neural networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.264829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d20da87907f817f1913d7653cabac3366db28601964c05d61e71109c4eee2daf

Observation 7e6903c2-86d8-4993-94e4-b342253ef19d · outbound

This paper cites Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.975604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d112f8f0c3f5915d6baadb339eb9e88d8c58904946a205c0755937895f43310c

Observation d573282a-6354-495c-8583-95afa23e3ed3 · outbound

This paper cites On adaptive attacks to adversarial example defenses.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases On adaptive attacks to adversarial example defenses

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.032926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:72246c0af583a9f5de74a90d187a9823ce505cdc3e0ed86ff0b76b5e3d2a7e47

Observation cf21cb4e-a0f7-4582-89d1-5fbcf43cf515 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.904124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5a7ed8e8eabf79da8e5350a2c6991d24ba80b49ece0a46654be3a1109536dd27

Observation d65179db-be13-4d8c-993b-9d1c73d08cd3 · outbound

This paper cites A step toward quantifying independently reproducible machine learning research.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases A step toward quantifying independently reproducible machine learning research

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.244808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:33694d112bc170bd7406550f1affbe0d34459c1755c4667ca505b5776d77a18b

Observation d57e395e-cbf4-41b8-a07c-a0d02a8199ee · outbound

This paper cites Emergent world models and latent variable estimation in chess-playing language models.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Emergent world models and latent variable estimation in chess-playing language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.988578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:24e3e08ea4140aeda470d11408134b7d44f2c1a8b60e73000cccc1e7da622ee1

Observation 58063378-f7c5-40a2-8f42-8931e7da98f9 · outbound

This paper cites Can language models learn to skip steps? 2025.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Can language models learn to skip steps? 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.185508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:2b5afa40248284bd6a8782b279c90129cd2d90d254e520eed0862c86b7d591a5

Observation a023374b-a327-458e-af06-df594d51ccac · outbound

This paper cites Reasoning models sometimes output illegible chains of thought.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Reasoning models sometimes output illegible chains of thought

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.958139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:0b7e326f2b8473be4ac096527e252ce8646949af37eaea3cf867a6b4ecf3b272

Observation 9a8953ca-805f-4264-bd48-68885235d899 · outbound

This paper cites subliminal learning reliably transfers preferences.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases subliminal learning reliably transfers preferences

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.925058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:e9e8033a930756e123edd6e361ca6044be7d6281c34f04b369fa3dbae603dff2

Observation 0d8ac9a9-1a0b-42a8-9bf1-4fca793a7595 · outbound

This paper cites monitor notice.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases monitor notice

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.082602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:fa701cbcbbd7e04770eccc90be8b3c724cb9d388e1b9d894a3b28c5285b3f423

Observation dc90355b-f53e-4877-b271-f25ba43032f3 · outbound

This paper cites lazy training.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases lazy training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.132350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:b648babcf0ac96edb3858c63dc212e78911588d204ad4692f075fa4c4f9ffe87

Observation 32d8873a-9fb2-4a85-bef9-b954b12bbda2 · outbound

This paper cites steering partially recovers deployment behavior.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases steering partially recovers deployment behavior

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.222270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f9c442951371d4152e164d20496cb92e5cbf84b976417602b5c0217b83549654

Observation d41e52e4-9502-4516-95ab-4cce8e494bfd · outbound

This paper cites demonstrated.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases demonstrated

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.196438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3e29f52a2c87c0f64ceac537f14f83eced683dbf3965e5455a743f53866cd488

Observation 61009a6f-a30c-4e22-9a80-48220ef61cea · outbound

This paper cites That is, it suppresses the tendency for LLMs to copy tokens that have already appeared in-context.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases That is, it suppresses the tendency for LLMs to copy tokens that have already appeared in-context

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.261011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5e639aab5913cf5ac47cddc57e77e606adbe64097100d43e05855f468732b8e1

Observation 1e351684-f079-4a8c-a2ca-9bcd7075d025 · outbound

This paper cites Specifically, theylookat WUWL10H7 QKMLP0(WE).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Specifically, theylookat WUWL10H7 QKMLP0(WE)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.142868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:deb4595890006410e02faba2113171fb7aeae0a5f5bd10b074da5cbae68d24f7

Observation b2f2ae17-6de0-4898-97fd-b632f24d65cb · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.226551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ca43b8ed4ae8490a66acda20bec99c311e6d87402291500cee6cf31301e4c1f4

Observation 40ff1000-c4fc-4ada-8577-55e26a5ee7d6 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.096784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:4f8d4f5055e7267040b66b55a38792870290469181431938ce2b49aedad5e9eb

Observation 86b7fd22-5284-4b43-90eb-d3123686aadc · outbound

This paper cites Without MLP0, the QK rank distribution changed from mostly rank 1 (strong copy signal) to only 6% rank 1 (no copy signal).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Without MLP0, the QK rank distribution changed from mostly rank 1 (strong copy signal) to only 6% rank 1 (no copy signal)

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.853082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5746211f1301806dee3aeb3692d3f91386c7a70c8b869108ba4403cf62f33556

Observation 6d027ee3-4924-4335-9f38-96577d81b6b6 · outbound

This paper cites This still explained 77.5% of the effect (vs 70.2% honest), but removed the evidence that the mechanism is specifically about copying.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases This still explained 77.5% of the effect (vs 70.2% honest), but removed the evidence that the mechanism is specifically about copying

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.914660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:101ae18b61519f32a59ed44a8e19a103b73c659915cb373dee8dcd83a5722c4f

Observation 19a791df-e7b7-4ef6-ab3d-df2b45ef3b7b · outbound

This paper cites L10H7 is a copy suppression head.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases L10H7 is a copy suppression head

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.944356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8d3f0276ac60d4273934c4922d13f8f6c6030fce8b9c7ef557348ea2a632470a

Observation c1ae90cb-ca95-4870-805f-f54c9b6ebe00 · outbound

This paper cites collusion team.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases collusion team

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.907618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:43c344ac0a3a832ae91682a849376ee4c1465d248724031deeef72e5b6015c26

Observation 368d1fdb-8898-4d66-8a70-c31c999b94c1 · outbound

This paper cites probes detect sandbagging with minimal data.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases probes detect sandbagging with minimal data

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.128961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:826c5730401be4f7c00cfd93be4e08f60b5fd0a6328370efeaf1de30359a4c5b

Observation c5bf74cb-8ffa-4758-8462-0cabf4f74aa9 · outbound

This paper cites multi-turn has no significant advantage with equivalent retries.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases multi-turn has no significant advantage with equivalent retries

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.249124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:1852860ea187f7ce02bd334cef14101b4a19e5524bf2574163352ed24b8823c0

Observation ee722478-a669-4bb2-9491-d0498092c13f · outbound

This paper cites function_name.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases function_name

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.947715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8834470f76382889712c77363174f92c82f7575566d929b267844914bd82028a

Observation bf55d150-3a23-422d-bafa-34d615338a26 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.961692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:94eaf6e7c68321f964701b91f7fa7834b92f634e6dd3c38159c26ffcb9cabc69

Observation 2fba69ea-c036-4188-8eb6-24120c927c25 · outbound

This paper cites Full Experiment.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Full Experiment

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.114334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f84e2f525d6f09fc552f61da85b8eea18abf25467ac084646c9135299a7070bc

Observation 8f2d319c-2d50-4cc1-8bbf-fe8eb992d4d8 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.237642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:498cc561a936d2df59c51593bcaf93676420efa9c137e1510362213c1bb66e04

Observation 88a4a91b-e172-40dd-b8b5-f46d991388d9 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.013969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:73b4513b50f4ff9f44dee0958b507bbc7032bbf56344e3e9fd0c6af3898f5e80

Observation 1231d4c2-70aa-45ff-a2d9-ed9b33f8e22f · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.182106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:8108a5d98cdbb75aa5c3b6cb7f829d46a35f3312054ebcb1598c77272cc83054

Observation 629f6302-d5d4-4f62-a3a6-e21a1d439631 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.061526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ab727dad1e1f0f3ba7fbb3ad9a8707b3a6e77fdf590656b6bdfb59fca2339414

Observation 60f80b60-9d9c-423c-ba20-ccff6749a2bf · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.979011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:40eaaa8cb0dfa87c25983fec93ffa06bdd764c67f157a669769cea0344824b18

Observation 6ca9b26c-3084-4683-9e54-e9f85f332cea · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.218337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3c8cc56bc94ead864cb5aa6120a0827289c325c6c530242fda3a1e49ed35c39b

Observation 7934e1bb-d83b-4ace-8d38-91b638175147 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.036335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5e88b536d7e5a2fb786e189072c8000025f1a9f42d6b01c3748dab0bd0cebe45

Observation 167830de-49fe-4a53-8c89-71c456b5ce9b · outbound

This paper cites Not just different from your expectation -- you can explain why it’s wrong.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Not just different from your expectation -- you can explain why it’s wrong

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.861427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:397877a56b4224aa478e568d07457af5638e368a859079938855110e6306a3cb

Observation d0429486-0689-4e0b-939d-1003538c27c8 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.290355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:3265516e2525e5120a5d65eaecbe1746cc37e672204615043c696ce1acce49a4

Observation 0cd4ef86-f1cc-43e8-ac86-457d9f675479 · outbound

This paper cites If I could only fix ONE thing to make the results honest, which fix would it be?.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases If I could only fix ONE thing to make the results honest, which fix would it be?

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.101687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:cbb7ce1dc2f99ae533cdf50f766e1ce0cb1a1d4d447a4cbace78ce500a51ced9

Observation 79c69459-6c3b-4b3d-9cae-b325b00b165e · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.885803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f3d9afb27f5395005f09b7a072ec3d34c0b77e1f424a487504f8938f8d12d164

Observation 4e230a6a-3e99-4c1b-b320-efb6f2b223b6 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.188983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:9485163487780bca0f9177af7a6f9c5b3bea98539490719ea5ae5820f5c04925

Observation a9ab27b5-fd71-4c6d-ab00-8c027485de4f · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:17.876612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:85845ec1d04006fb58ddf859f94328b1575e5ca3da07e79dda042ec89b87537a

Observation 39df3eaf-1eb6-4cad-88cb-a0b0b39550b8 · outbound

This paper cites J.3 LLM Red Team Prompt The following prompt was used to instruct LLMs to generate sabotages of honest codebases.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases J.3 LLM Red Team Prompt The following prompt was used to instruct LLMs to generate sabotages of honest codebases

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.857475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:789408e26e44353ecb9e42c00671a82e037a45ed8cd99fa9e8fff4fb53cad821

Observation fc0ee495-10e0-4671-843b-9332f7c2cdcc · outbound

This paper cites The Introduction and Methodology sections in the main body of the paper are always honest.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The Introduction and Methodology sections in the main body of the paper are always honest

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.968750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:87ee3e0e5534e01dce680595bd927f8ef304eb5578aface1ee937ac841ec2825

Observation f1e8b5a2-4fcf-41c8-8ead-a8a0280d826f · outbound

This paper cites Your sabotage must be in code (.py, .js, .ts, config files, etc .), not in data files or pre-computed artifacts.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Your sabotage must be in code (.py, .js, .ts, config files, etc .), not in data files or pre-computed artifacts

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f635bd04262fda64fd1cd064083ff2257f85fee419e82e1248bd883ff6f7c6be

Observation 411ab9f0-b2cb-48c8-bfdd-231def1f85f3 · outbound

This paper cites Do NOT add suspicious-looking code that is unrelated to your sabotage.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Do NOT add suspicious-looking code that is unrelated to your sabotage

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.971973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:59728b4e305e92b2c92135cdee53b9d0ee7fd200746a935dbd8f898244b41bbb

Observation 48111a0b-46e5-4a5b-833b-bc433383d649 · outbound

This paper cites Do NOT clean up existing bugs, quirks, or messy code in the codebase.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Do NOT clean up existing bugs, quirks, or messy code in the codebase

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.873020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:7ff601614d95ad2d769a60679729dc922eca140ba208c3849df5903e8950e76c

Observation 64ba36de-4ef9-4104-ac21-941b5d064427 · outbound

This paper cites You may rerun experiments a few times and pick favorable results, but do not do extensive search.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases You may rerun experiments a few times and pick favorable results, but do not do extensive search

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.146398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c68972acec50f9dbec59f5459e8683bd777f039ab19e44bfafba2285998c53a9

Observation 39575a7f-dba3-4f0e-ae61-2981b1c022f1 · outbound

This paper cites The Key Findings subsection in the Conclusion must accurately describe the experimental results your sabotaged code produces.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases The Key Findings subsection in the Conclusion must accurately describe the experimental results your sabotaged code produces

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.125222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:5d88f0c2e04faa166d6d15cb95a109844093860771395fba1c9924fdd7ec1c27

Observation 75028d5f-b80d-4ed5-af47-7495804e40bd · outbound

This paper cites test" to.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases test" to

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:17.921609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:ae2f400c1f1cc74757d71b8e242d8d5871024c1864da394163318de83b5c11e5

Observation 90665e87-b0db-427b-b047-5758c33a3ddf · outbound

This paper cites bad choice.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases bad choice

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.206983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:07d537afe8f3cc4547c80ed8c1088f89ad8be27d318e68203a60c2a8b4444d7b

Observation f872dc9e-a612-44d1-8037-775c8668cd99 · outbound

This paper cites These are often just a single number (batch size of 32) or a single phrase (Adam optimizer).

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases These are often just a single number (batch size of 32) or a single phrase (Adam optimizer)

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.192944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a85666702aad9e8da25312775a6a717e3e25dfd84f7683f7ce794dffe54246e8

Observation d718ee5c-3234-4720-88d6-683e0951f9bc · outbound

This paper cites Qualitatively different.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Qualitatively different

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.294427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:c9722642b29a0c910a25ccfdace21ef7d88c0ac6fc063bf25228412d4bee27b3

Observation 010fb39d-c9cc-4a50-a3b2-a53280cb1feb · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.199823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:052b7384d84bf014b06c4e23918b31a983eb7db07f8bd9bcafaa531d1c379b0e

Observation f23711e2-dc1e-4763-9591-19dc2183b5f2 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.257038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:a55365ed53e7f695adba4b73d79540226606d23c56cd00e05a486d28ae42682a

Observation 207b2374-ad79-4d60-9589-d9978c292547 · outbound

This paper cites an unresolved cited work.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-21T00:14:18.161055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:918cc1ebe4c0ee139e67f04c61ec25991120c24ab561ee0446045f8264868672

Observation a9891bb8-553f-494a-a473-e63bd44c8116 · outbound

This paper cites Later stages (evaluation, results) are easier to check because they can be rerun quickly.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Later stages (evaluation, results) are easier to check because they can be rerun quickly

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.178640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d8cf02f997b98b5dfc3772f14cebe94ce9a3635b18cd5259a099c39cc3705afd

Observation 28c4cae5-2ce0-4845-8677-894944a0e592 · outbound

This paper cites Read PAPER.md thoroughly.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Read PAPER.md thoroughly

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.230243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:f87e705f6765685ed60d8b832a0e0bf22077dcb9a08e7b8a5d68646f6203129e

Observation d5925078-e754-46f5-9b73-6c8a0fd8bc79 · outbound

This paper cites Walk through each pipeline stage and list every design choice: hyperparameters, model selections, data processing decisions, evaluation settings, prompts, sampling procedures.

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases Walk through each pipeline stage and list every design choice: hyperparameters, model selections, data processing decisions, evaluation settings, prompts, sampling procedures

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.214607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:31:53.114772Z digest=sha256:d46c415eac2d0d5147f4bb68391292fb424fd4b3ef98753ad8ef76c8178ce540

Pith citing papers

Observation f29b1824-b56a-46c5-96e4-0af740a7d6c7 · inbound

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D cites this paper.

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:57.400820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:57.400820Z digest=sha256:38d8cc62717db0ffecb628976bfdb77f07e87aebd9a3f71d0dc882ed10050e0a