Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T20:31:50.043920Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 100 of 129 outbound references and 4 inbound Pith citation observations for arXiv:2605.12673.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-14T20:31:50.043920Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:35.095351Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-12T00:12:42.362881Z
100 of 129 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 891ab41e-bc43-41bf-b548-03096873f857 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Concrete Problems in AI Safety
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a7fba226-432f-4afb-9b10-69aeb0ab1979 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Alignment risk update: Claude mythos preview
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8c849738-2c8c-4f25-950e-b6b7edc8ddef · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Claude code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 971037cc-e3c1-4c2e-adfa-5b98ee9ae3a7 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Analyzing and improving chain-of-thought monitorability through information theory
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6db6f338-5a7d-4c70-b07a-8bd2834e63ea · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rewardhackingagents: Benchmarking evaluation integrity for llm ml-engineering agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6a176a57-f55a-4f09-acf9-c25e441b7e5f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d8bc6e99-076e-4134-b5b5-c55f67b5e467 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Adversarial reward auditing for active detection and mitigation of reward hacking
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8741e97a-d21d-4b4d-8341-40c8bcb021d2 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack What Will it Take to Fix Benchmarking in Natural Language Understanding?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1f887bd-a17f-4a92-b2ef-065878966ef9 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0ec8fc1c-28d0-4133-abd5-e6e0d7caed93 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack arXiv preprint arXiv:2502.17521 , year =
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b84fc53d-0b68-463d-b08b-8de0dee67701 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 24b201f8-68b0-46b3-aafc-03a13b094649 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reasoning Models Don't Always Say What They Think
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5b96a8e9-d100-49fd-9dda-f8e395fc79f9 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f57ad81-487d-4945-ab52-ff5ed9280a0e · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack The Benchmark Lottery
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eb042689-a284-479a-8f12-f6a3a55d4fd3 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dcf2e410-a702-4925-9ad1-6e2cc7f80457 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 688e7832-5307-4d62-912d-270b8db9bd2a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Benchmarking reward hack detection in code environments via contrastive analysis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 20557408-85e8-4ef3-bea9-6014f7709f51 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8231cbf-0c30-4838-ba88-0348a5a08f30 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Generative adversarial nets
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 970d21b9-f531-4e21-bc6f-1d3c77bd1dc1 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Problems of monetary management: The UK experience.Monetary Theory and Practice, pages 91–121
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 55f6a03c-a705-493b-a3f9-0f7671f381af · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 91f824d3-4cc5-470e-a739-a6a1d5fe2d9d · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 95ab70cd-b430-4086-8a80-4793071f6217 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Issue #14: Iquest-coder-v1
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 75d75803-1887-43d6-aa4d-943e7aa756fe · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cff115f9-91d8-4e04-92c3-93e4e40e7104 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a4fecea0-9a15-419d-b761-be0e5ca96581 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f30ec8bd-62c8-46a3-871c-e7a487cee1e7 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc9f9a25-82dc-4095-b9c5-322d5d6ad964 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a88031bd-cb72-47c5-bb47-926e1d8b84e3 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b4c7eac9-a96e-4e31-991e-2af5ca23690d · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ea3e3b66-fad3-4dd8-a431-b6b5a0afc290 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Diagnosing Pathological Chain-of-Thought in Reasoning Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7c30f080-14de-473c-9660-a283785d9642 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack AgentBench: Evaluating LLMs as Agents
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 45f6f51e-7648-46fe-9c3c-cc0e84f7161f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Natural Emergent Misalignment from Reward Hacking in Production RL
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b23de401-40ff-4c31-ae35-90626ce4c12d · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Gonzalez, Jingbo 12 Preprint FrontierCS T eam Shang, and Alvin Cheung
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c32812f1-7fb8-42ee-940e-4ce580779882 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b30275a-5e73-4e8c-a910-49f21a762280 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack GAIA: a benchmark for General AI Assistants
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ae9e7d26-2de6-4bdc-8c2f-e27c4060e10a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Introducing codex.https://openai.com/index/introducing-codex/
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 438b2977-8481-4dd6-a6d4-3c45a93a9207 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Why swe-bench verified no longer measures frontier coding capabilities
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 824532d7-400d-40ed-8919-f33015637cc8 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Proving Test Set Contamination in Black Box Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6637ef59-6403-4448-885f-50d32873c0e6 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack KernelBench: Can LLMs Write Efficient GPU Kernels?
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ffc46a41-e1c3-40a8-a160-731b1f3714e3 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Feedback Loops With Language Models Drive In-Context Reward Hacking
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9831abea-ee67-4253-a4c0-34ef70e079d2 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Frontierswe: Benchmarking software engineering skill at the edge of human ability.https://www.frontierswe.com/
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b7ef27b5-44f3-4959-a71d-bcadc872f6b8 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 13b2a7bc-488b-4033-9d77-e8a68d539c75 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Posttrainbench: Can llm agents automate llm post-training?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fc2b3a61-925d-4124-803d-5b8dac7f352a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack PostTrainBench: Can LLM agents automate LLM post-training?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d6b62d88-4d53-4aff-9442-31fa43fc5179 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a1f7e64c-02e0-4783-8fc6-6403b05fd57a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d4512d50-caf6-408d-91fd-53055ffd103e · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Defining and Characterizing Reward Hacking
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 209f8941-b88c-4e25-860a-102c191ffd50 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Detecting Safety Violations Across Many Agent Traces
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aa31985a-c5c4-4a1c-8cf4-12c62ae7e30f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack improving ratings
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 17b8a3d1-a9e4-40cb-9dbd-82c75f668e6d · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Recent frontier models are reward hacking
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c14b6047-9106-4091-a5f1-5cbb034ea167 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2ea1b0e3-a0d8-4623-8742-42e1973d8e79 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f8358352-54c7-4d15-a484-b46f2d7741a9 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 905d6597-1ec7-40cc-b0fe-10b17eb2335f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Detecting and Suppressing Reward Hacking with Gradient Fingerprints
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bca79f29-383c-45bf-aeec-fa5c741cfd68 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward hacking in reinforcement learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fc031751-6045-4792-a429-c80663b4a523 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Monitoring emergent reward hacking during generation via internal activations
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 129bda9f-d70a-4730-97fc-d03263e9b001 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 28f67635-80e8-445c-b423-2b71c8b42b35 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Investigating cot monitorability in large reasoning models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f7878bb6-d958-48d6-a10f-13d2cdb9def3 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4f495794-702f-4bb1-8343-5752ebf469ed · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b73a68c2-028a-43f6-8f0d-81d129b78d1a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b205cdf-96ee-41f0-87b9-3b0c39b8354b · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 610552bc-b577-4555-af5a-0e0021d293a0 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5545b94b-9fde-4482-99e4-00978f65bbff · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Netpress: Dynamically generated llm benchmarks for network applications
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9892f498-abd3-42e6-8b3b-f8dc35dc75ac · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing best practices for building rigorous agentic benchmarks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 54fc18c7-13b3-4f01-ba8a-8a9115d6c13d · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ea0ce51-533d-4e7f-bf2f-395262f6381f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack [exec(\"\
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ba73d78c-ee42-48ed-ac09-221cf49232f0 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack {name} benchmark github
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 47673d6d-bb9b-4c2a-abac-182d16123638 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2e5def7-b746-4fa3-a8bc-47e48a4633f0 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 7 8If the benchmark is well-known (e.g
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d0c82d64-3b08-40bc-8dcc-a4486a5fe175 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5198de44-7cb3-4395-bbf6-441549a2d6de · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b85deeaa-af8f-4286-a1d9-333e7bd619cf · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4714beb-dd3a-4113-af0a-3db315f5d17d · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6fc5e842-9b62-4f5d-b65f-929ecaf57aa8 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8ec067b8-c9b6-4c7a-ab80-4e19bc931d68 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e04210aa-ef86-412c-81eb-d8dbb5a4e75a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task_id_1
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ca9cc7de-6c67-42e0-a53e-e566e5933873 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 307d07ad-d9a2-495e-aa24-d94621d5a73c · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task": "<task_name>
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2d2c5ca2-b64f-4587-8be1-c70625988b5a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack If it’s a URL or package name, clone/download it
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2dea01d-7d8b-438d-8487-73ad40b591fe · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3b1cea2a-a412-4bf5-82d2-cabfff1562e3 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack This is critical -- every point where agent output touches evaluation code is an attack surface
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ff6ad4f3-6093-480f-b0e7-4aedf0b2eb71 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Report: 49- **Docker images / large files**: Does the benchmark require pulling large Docker images, datasets, model weights, or other heavy artifacts? Estimate total download size
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bbee520-003f-40ba-bd83-ac8fb7da8e7f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task_id_1
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85201882-31d6-43ca-96ce-04efc05db9d0 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6a7ac56d-a637-4e35-a07b-3d16d8f49736 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 326e90fa-5212-4f40-bf44-0f1282fad61f · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 233c198e-ac02-427d-93f0-db24c9e7758c · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 275- It should set up the environment (install deps if needed), inject the exploit, then launch the evaluation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation abd3adf9-501a-4497-a496-d68af95d6cbd · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 99c8d5b7-00ae-446f-8082-5d32de178fac · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 309- Examine the task’s specific evaluation logic -- some tasks may have stricter checks, different scoring paths, or edge cases the current exploit does not cover
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0150ac14-3ecf-4e65-9c62-c2da0613dc43 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6b4a04f6-1358-40af-9477-53587d0ece94 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e7ae5a2d-1d21-4d41-b9f3-74c388754970 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 316 317Each iteration should be a deliberate improvement -- do not re-run the same exploit unchanged
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9fdcdc83-c97f-41e7-880c-f4909279ad73 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6c058542-f296-479c-a109-35482807ce44 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e3b51c43-9b8f-4d8f-9617-5d161b24108a · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1a41981-224f-4419-a8c3-2dfaaeb0ed35 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task": "<task_id>
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bcf1007-d79c-4416-a94e-aa11d6fb7d38 · outbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack If test.sh aborts before reaching compute_reward.py (e.g., set -eon a missing file), Harbor still ingests the pre-written value
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6efe1fcd-1a88-46c9-bf30-345e8cf79660 · inbound
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070048b1-064a-4cc1-932d-4c002e77f329 · inbound
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 79f9d0c6-a0b2-4cbd-a4ed-ecaf38d7708b · inbound
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35ee17cc-9a66-49da-ab10-9323b7abab82 · inbound
Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.