Pith. sign in

Paper Citation Record · LEDGER

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

As of 21 August 2026, this Paper Citation Record lists 100 of 129 outbound references and 4 inbound Pith citation observations for arXiv:2605.12673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12673 v1

Coverage vector

measured 100 of 129 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T20:31:50.043920Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:38:35.095351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:12:42.362881Z

Reference resolution

100 of 129 outbound references displayed

  • verified exact37
  • verified fuzzy32
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 891ab41e-bc43-41bf-b548-03096873f857 · outbound

This paper cites Concrete Problems in AI Safety.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.730139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:809eb6df57a15be080a378f758858ae7999b2a07783a6316b3efa1be060ecd94

Observation a7fba226-432f-4afb-9b10-69aeb0ab1979 · outbound

This paper cites Alignment risk update: Claude mythos preview.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Alignment risk update: Claude mythos preview

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.046492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:269b280b5e54cac94e24ff5f8df4319254fff1a90d3d72f8499a6efa36dca448

Observation 8c849738-2c8c-4f25-950e-b6b7edc8ddef · outbound

This paper cites Claude code.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Claude code

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.050001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5a8cc934a2593b8967b3852a83858adf3c1e35348990ce5d3459fd4b68da81cd

Observation 971037cc-e3c1-4c2e-adfa-5b98ee9ae3a7 · outbound

This paper cites Analyzing and improving chain-of-thought monitorability through information theory.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Analyzing and improving chain-of-thought monitorability through information theory

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.724381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:327324938494878fb827bd57f278486d9e77017a1933c433e79789d3faa7c35d

Observation 6db6f338-5a7d-4c70-b07a-8bd2834e63ea · outbound

This paper cites Rewardhackingagents: Benchmarking evaluation integrity for llm ml-engineering agents.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rewardhackingagents: Benchmarking evaluation integrity for llm ml-engineering agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.768744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b291ea4e6115d57a34af816173acf8341710f1c3f54dd6611255b1c61ec5eb53

Observation 6a176a57-f55a-4f09-acf9-c25e441b7e5f · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:24:13.179886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:e20b6ea05c9244d480661826148c0cdde76e4019d6463eaafaaec2ec90f7ef59

Observation d8bc6e99-076e-4134-b5b5-c55f67b5e467 · outbound

This paper cites Adversarial reward auditing for active detection and mitigation of reward hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Adversarial reward auditing for active detection and mitigation of reward hacking

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.897304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:000a9cfc53a0d73f6994a58043672958318c3eaaac0d1c2e28912bbef0dd9d9a

Observation 8741e97a-d21d-4b4d-8341-40c8bcb021d2 · outbound

This paper cites What Will it Take to Fix Benchmarking in Natural Language Understanding?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack What Will it Take to Fix Benchmarking in Natural Language Understanding?

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.901133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:785f83bd97bd6a58db6a29a93f78713e966936c3c3319f871bb39d003fb99fc9

Observation f1f887bd-a17f-4a92-b2ef-065878966ef9 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.909283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:093f81194b96efe4aa234400251e6badfb449c827a0c34fd87643ec7419ba9ed

Observation 0ec8fc1c-28d0-4133-abd5-e6e0d7caed93 · outbound

This paper cites arXiv preprint arXiv:2502.17521 , year =.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack arXiv preprint arXiv:2502.17521 , year =

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.892065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:10f13ef6fbc7fc7a537227315531291c6f41e9e0369f7d31f49f03ecdf75a561

Observation b84fc53d-0b68-463d-b08b-8de0dee67701 · outbound

This paper cites Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.883357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:618a05598bc4e678955c22fb932fa5699b94b957ef45163c5af107001b496991

Observation 24b201f8-68b0-46b3-aafc-03a13b094649 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.763151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d9811031288d75adf02ec400f94abcfa0170713357660acac3f8095cab732d2a

Observation 5b96a8e9-d100-49fd-9dda-f8e395fc79f9 · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.937213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:cc8f79971a0476300031e631232097b33f35e7100de54b97c4013e67f1406a65

Observation 5f57ad81-487d-4945-ab52-ff5ed9280a0e · outbound

This paper cites The Benchmark Lottery.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack The Benchmark Lottery

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.887648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:592bf83744a70b95570b78eca9fe8444b912ea7ba48f9402be4986dba4550546

Observation eb042689-a284-479a-8f12-f6a3a55d4fd3 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.742365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:38a8b1803c740cc8a62c7ff4c8ada8b6c19267f9917914360324f3a31cc4af57

Observation dcf2e410-a702-4925-9ad1-6e2cc7f80457 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:014f247334a48090484905da8efcf70bbd81af3b620b63465a763be052e54ac2

Observation 688e7832-5307-4d62-912d-270b8db9bd2a · outbound

This paper cites Benchmarking reward hack detection in code environments via contrastive analysis.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Benchmarking reward hack detection in code environments via contrastive analysis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.872818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:035c202729d1ba45a5fa3c225ebf5081683bfa77aef9f7b8255a5fa5fc225439

Observation 20557408-85e8-4ef3-bea9-6014f7709f51 · outbound

This paper cites Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.877863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:086e82383de929c3b3f9dc58956919a4d803b080af3859c74f405db170c45b01

Observation b8231cbf-0c30-4838-ba88-0348a5a08f30 · outbound

This paper cites Generative adversarial nets.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Generative adversarial nets

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.963694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ae517cda119933f7965ad86142117c93d9f3e7364ccc9f1d2f79365c162edc40

Observation 970d21b9-f531-4e21-bc6f-1d3c77bd1dc1 · outbound

This paper cites Problems of monetary management: The UK experience.Monetary Theory and Practice, pages 91–121.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Problems of monetary management: The UK experience.Monetary Theory and Practice, pages 91–121

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.938448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:e88b71d67e00131ee9a2fc2ed90ada2ca3745efbe627d88d6e2a55ff4cead71f

Observation 55f6a03c-a705-493b-a3f9-0f7671f381af · outbound

This paper cites Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.914038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:91e7aae7c333fd6f626d93c7a5e1788eb329b87be59e8eedcb6b89a0523af1e3

Observation 91f824d3-4cc5-470e-a739-a6a1d5fe2d9d · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.970116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8eb79c9085542b8e92da5b8ed775c7fdab25af0023ba876482dc421d2480d70e

Observation 95ab70cd-b430-4086-8a80-4793071f6217 · outbound

This paper cites Issue #14: Iquest-coder-v1.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Issue #14: Iquest-coder-v1

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.873321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:5c0c2b80624f730a58a8d47c9b0eb55d0c889d0d1d6f6b2742bcdb501c97e3f7

Observation 75d75803-1887-43d6-aa4d-943e7aa756fe · outbound

This paper cites Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.875651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:162f792ee13c268488376afa2f3d2cabfca467d3a79756a8d33027091e0e13db

Observation cff115f9-91d8-4e04-92c3-93e4e40e7104 · outbound

This paper cites Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.943868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:93000153d1e628d624eb5be9c62ea2ce8bfa6e33c59658d1bd154be90abcd91c

Observation a4fecea0-9a15-419d-b761-be0e5ca96581 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.752707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:23430b65d19325eeee5b9118a4b4c33245f081bd5854db62d1d02867b183b248

Observation f30ec8bd-62c8-46a3-871c-e7a487cee1e7 · outbound

This paper cites Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.747821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4997d7759060cf68b083c5cb5685befb515b035267c8dd6b7420bb369f90daa7

Observation bc9f9a25-82dc-4095-b9c5-322d5d6ad964 · outbound

This paper cites Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:19:45.019656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:6b440c4fde8f868f75cb858cb7a732d96fa462a7377873a83430ba73f5b78bca

Observation a88031bd-cb72-47c5-bb47-926e1d8b84e3 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.794085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:409f976b1480cffa0247d56e74bb5294cbe44df49d7d96f8b97d326d653c8a06

Observation b4c7eac9-a96e-4e31-991e-2af5ca23690d · outbound

This paper cites ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.736935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2db1a1d3c7c6b72e583a52ff609fb508029e12c49d49d96e9a58e9fd00fa3b3c

Observation ea3e3b66-fad3-4dd8-a431-b6b5a0afc290 · outbound

This paper cites Diagnosing Pathological Chain-of-Thought in Reasoning Models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Diagnosing Pathological Chain-of-Thought in Reasoning Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:05.474485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:79f327f36beee86660b0b68da9fc00147410ab3b49804a9e829879555f0e7f76

Observation 7c30f080-14de-473c-9660-a283785d9642 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack AgentBench: Evaluating LLMs as Agents

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.927996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:28251973c853bf2a6764551992bbe60e6ff16adee3afa1c88cfa940c8038c910

Observation 45f6f51e-7648-46fe-9c3c-cc0e84f7161f · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.933390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:c088b5e275f60121206cbe770ac4cbd98b2638c7762fcce7e5ce5a352373d333

Observation b23de401-40ff-4c31-ae35-90626ce4c12d · outbound

This paper cites Gonzalez, Jingbo 12 Preprint FrontierCS T eam Shang, and Alvin Cheung.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Gonzalez, Jingbo 12 Preprint FrontierCS T eam Shang, and Alvin Cheung

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.938683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:3aa8bc64959923952850d4d89138af30e7d304fa53c304808864d0b034a11877

Observation c32812f1-7fb8-42ee-940e-4ce580779882 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.949393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:833e45a610439e15a1bb8fa73aaceef647246b973db4baef18c33fbc08fa01d2

Observation 8b30275a-5e73-4e8c-a910-49f21a762280 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack GAIA: a benchmark for General AI Assistants

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.965337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:c7ed2ec4bf0891668ae2b99d2846d36067c7d9336e920e22c32f4e74e50e7ca3

Observation ae9e7d26-2de6-4bdc-8c2f-e27c4060e10a · outbound

This paper cites Introducing codex.https://openai.com/index/introducing-codex/.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Introducing codex.https://openai.com/index/introducing-codex/

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.877941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8c0641cad62442b8aeb3ad2f027bc8b6a16bf50bf3eb040493531239b3d86e1a

Observation 438b2977-8481-4dd6-a6d4-3c45a93a9207 · outbound

This paper cites Why swe-bench verified no longer measures frontier coding capabilities.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Why swe-bench verified no longer measures frontier coding capabilities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.866415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fe5279fd07649ea2cdb95edb3bd4d5e036b8a6a147aaa10aa6e2752e7eb65e47

Observation 824532d7-400d-40ed-8919-f33015637cc8 · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Proving Test Set Contamination in Black Box Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0b815a6cf864737d65f64538925bd5f0d874c6e6d0cfc8abe7f4cdccaf990550

Observation 6637ef59-6403-4448-885f-50d32873c0e6 · outbound

This paper cites KernelBench: Can LLMs Write Efficient GPU Kernels?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack KernelBench: Can LLMs Write Efficient GPU Kernels?

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:55:02.220088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d561752cf7473ebac13b3248ce1b941c1d506d3b48dad21725a14595b4a19c16

Observation ffc46a41-e1c3-40a8-a160-731b1f3714e3 · outbound

This paper cites Feedback Loops With Language Models Drive In-Context Reward Hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.960515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:459d1531b70301a6d31388dce75ee92608ff01e8bd2363cf4f06f5d3318fee47

Observation 9831abea-ee67-4253-a4c0-34ef70e079d2 · outbound

This paper cites Frontierswe: Benchmarking software engineering skill at the edge of human ability.https://www.frontierswe.com/.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Frontierswe: Benchmarking software engineering skill at the edge of human ability.https://www.frontierswe.com/

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.929623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:178cf6c2499edfe7de82c2a1e2ca6ccaf670915cfa567fad2efcd990ef2fc4e4

Observation b7ef27b5-44f3-4959-a71d-bcadc872f6b8 · outbound

This paper cites Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.914374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:375cbaf6bc57add94dd24841a327f8de03a1c3ee520a07dcd5b80536e13c6c05

Observation 13b2a7bc-488b-4033-9d77-e8a68d539c75 · outbound

This paper cites Posttrainbench: Can llm agents automate llm post-training?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Posttrainbench: Can llm agents automate llm post-training?

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.863338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:764172a2f925404b5d4b0d75235ae483653c16f6b61d56e7b3c08f74df18e546

Observation fc2b3a61-925d-4124-803d-5b8dac7f352a · outbound

This paper cites PostTrainBench: Can LLM agents automate LLM post-training?.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack PostTrainBench: Can LLM agents automate LLM post-training?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.858065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4f23fea08c0d0d64c1489f56f82d7a0df802f21e473662e8b601690eb2b6127e

Observation d6b62d88-4d53-4aff-9442-31fa43fc5179 · outbound

This paper cites Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.848195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0d0a909520aeba7cfd25344ee0540c8ee0da0a268ace0a51aa01a317ef6d39bf

Observation a1f7e64c-02e0-4783-8fc6-6403b05fd57a · outbound

This paper cites Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.900600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4ee4ce9cd1e334d9cb2f92c2b513679bd869ff2dc82cfd9a2be5c8a87d0f1155

Observation d4512d50-caf6-408d-91fd-53055ffd103e · outbound

This paper cites Defining and Characterizing Reward Hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Defining and Characterizing Reward Hacking

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.842198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:cdf2b42a4f0007c68864e896624035c0c2c469d2b46625a53e1e7b07be0460e0

Observation 209f8941-b88c-4e25-860a-102c191ffd50 · outbound

This paper cites Detecting Safety Violations Across Many Agent Traces.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Detecting Safety Violations Across Many Agent Traces

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.820416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:edbde7f7fd2fdc421fb557e80bac5414071c4a4e8af3cdcc34ad12f64f7c0c9e

Observation aa31985a-c5c4-4a1c-8cf4-12c62ae7e30f · outbound

This paper cites improving ratings.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack improving ratings

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.909376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0d78a89fdcb1418dcc67a7f20bed0d16201cb8d15c2b30b9f9781c12d9cea76a

Observation 17b8a3d1-a9e4-40cb-9dbd-82c75f668e6d · outbound

This paper cites Recent frontier models are reward hacking.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Recent frontier models are reward hacking

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.871075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:76fba0798bd1ec012bf6edc8086cc9326412b6a2fa2df0f074378ed12729cbe2

Observation c14b6047-9106-4091-a5f1-5cbb034ea167 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.921853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ca9870ef234ccc4ffae1d404baced45b93f7181063494b2a50e731e1d2e9250e

Observation 2ea1b0e3-a0d8-4623-8742-42e1973d8e79 · outbound

This paper cites FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.803955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:491e70aeadf6a7687db661cf9f848452983c4ce20e4dc7c7d9aa9024fd86f970

Observation f8358352-54c7-4d15-a484-b46f2d7741a9 · outbound

This paper cites BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.830494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:81c6994eba24aa8ea0011e7765226ca5b28231f7a9528f778f1531250822c212

Observation 905d6597-1ec7-40cc-b0fe-10b17eb2335f · outbound

This paper cites Detecting and Suppressing Reward Hacking with Gradient Fingerprints.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Detecting and Suppressing Reward Hacking with Gradient Fingerprints

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.825442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:c4b4cb2753cee37d7d98f91a2642838f22da711d4a2a766590791dac564f5631

Observation bca79f29-383c-45bf-aeec-fa5c741cfd68 · outbound

This paper cites Reward hacking in reinforcement learning.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward hacking in reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.934517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:eee2ffc4c20f6c06622914e6e362ce3af0106b6a2b046fdf4d7a19618f8543e0

Observation fc031751-6045-4792-a429-c80663b4a523 · outbound

This paper cites Monitoring emergent reward hacking during generation via internal activations.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Monitoring emergent reward hacking during generation via internal activations

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.809401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0c73be90a754e28ed55a6ac0f31b1f894b66fea3e30f9c0a73552c559bb87dea

Observation 129bda9f-d70a-4730-97fc-d03263e9b001 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.835280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:0399098aaa6dd02de13db025d9769bc89e286f307d96cd1635b1bc9979a5fee0

Observation 28f67635-80e8-445c-b423-2b71c8b42b35 · outbound

This paper cites Investigating cot monitorability in large reasoning models.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Investigating cot monitorability in large reasoning models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.853082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:9826760674679504eee12ab0351439807ba169777ea0dd34f8baa2cbb0552628

Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.862291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:3b5df0de2d74c30e6180fd55c6f0bfccf6976ec582be9cb354023e4d9c85ef28

Observation f7878bb6-d958-48d6-a10f-13d2cdb9def3 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.866671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:922e2225f1abba34b2d4c561de1e3454f0cfd91dcb11985f04d8f7a295b03d8d

Observation 4f495794-702f-4bb1-8343-5752ebf469ed · outbound

This paper cites UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.814909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2232e2d7cedaf3d8066612e212bb3dfa3111e5aa076c18c04528557631604b06

Observation b73a68c2-028a-43f6-8f0d-81d129b78d1a · outbound

This paper cites Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.854939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:f8264b514f62229777483750277ab87d35e56190ea8126fcd90ef683f206c431

Observation 8b205cdf-96ee-41f0-87b9-3b0c39b8354b · outbound

This paper cites Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark,.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Swe-abs: Adversarial benchmark strengthening exposes inflated success rates on test-based benchmark,

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.954541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:79043f672a84e847345bc2e05ad5f7415560ead5aa01807dd7eea9f480c74b49

Observation 610552bc-b577-4555-af5a-0e0021d293a0 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.780884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:bed7fc7e38ac457edd6ca3c1d216573b04e5546f2ca932089329e4a4cafeb6d3

Observation 5545b94b-9fde-4482-99e4-00978f65bbff · outbound

This paper cites Netpress: Dynamically generated llm benchmarks for network applications.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Netpress: Dynamically generated llm benchmarks for network applications

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.788338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:8a8aeb1371272330e4e789195eba428ef9db0c5cc12008eff6d9202738ea4f26

Observation 9892f498-abd3-42e6-8b3b-f8dc35dc75ac · outbound

This paper cites Establishing best practices for building rigorous agentic benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing best practices for building rigorous agentic benchmarks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.907599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:87fc48a83061a69fd1c18256e1c4821691bdd5f357829da81aeeb9805e7cea51

Observation 54fc18c7-13b3-4f01-ba8a-8a9115d6c13d · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.775266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d27f07829ef48eb1b27fa03038b45ac9491990575272c406becc6647d2f580c2

Observation 1ea0ce51-533d-4e7f-bf2f-395262f6381f · outbound

This paper cites [exec(\"\.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack [exec(\"\

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.861039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:47f88718150dd24567933445b6d362d7f4bcfc87296e692b774dc2647eac14b7

Observation ba73d78c-ee42-48ed-ac09-221cf49232f0 · outbound

This paper cites {name} benchmark github.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack {name} benchmark github

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.940010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:6a88a7ab6faa060d3c7ff79609ccc86e3a5c8090610833b6d27080d3f8d39acc

Observation 47673d6d-bb9b-4c2a-abac-182d16123638 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.878270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:bcebe371df49988b344e0f2ae131b6df80249bb86791ee7b955b46776403cb70

Observation f2e5def7-b746-4fa3-a8bc-47e48a4633f0 · outbound

This paper cites 7 8If the benchmark is well-known (e.g.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 7 8If the benchmark is well-known (e.g

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.911294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:33d9aef56dca663328cbc8dfb0d1c09566afe4167777ee0b25e366250a856a81

Observation d0c82d64-3b08-40bc-8dcc-a4486a5fe175 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.903132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b413abef27261691887895d4db57635191d9649a3219f2d61b51d2b8d5d8d163

Observation 5198de44-7cb3-4395-bbf6-441549a2d6de · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.844960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4d9d7ef920601af0e5c01032c35edbbbc5ea093254477ca0f52d1ae39085e546

Observation b85deeaa-af8f-4286-a1d9-333e7bd619cf · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.919417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:b7308945c40c6e336648fb78b4db3e946039956ddaec69137325f7b18711f1fa

Observation f4714beb-dd3a-4113-af0a-3db315f5d17d · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.022699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ed191b41675f4a0cf6d7920d9e6b95fbc68b199b9420cf418963163b48c13869

Observation 6fc5e842-9b62-4f5d-b65f-929ecaf57aa8 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.977947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:fc77e1479aecd830188548c54eb1ddb2a808775170fd4971ebf64129de5239f1

Observation 8ec067b8-c9b6-4c7a-ab80-4e19bc931d68 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.955253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ab62f5dc9db84f3b7c4bd990c20afa269d5dacecaafa420accb1dbbbba4c259d

Observation e04210aa-ef86-412c-81eb-d8dbb5a4e75a · outbound

This paper cites task_id_1.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task_id_1

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.926837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1a0be46fc257fa7943d4347e95dd8c63e865a7152146aed7444c698c96352f76

Observation ca9cc7de-6c67-42e0-a53e-e566e5933873 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.056466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4881473958d472b95459d928ed3bcbf364249c02fa2ca5d36f5cf957ec35164e

Observation 307d07ad-d9a2-495e-aa24-d94621d5a73c · outbound

This paper cites task": "<task_name>.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task": "<task_name>

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.059609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:38aa75eced195ab4e798a44adbf686723bdbf351e3696063dc878781ceeb5661

Observation 2d2c5ca2-b64f-4587-8be1-c70625988b5a · outbound

This paper cites If it’s a URL or package name, clone/download it.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack If it’s a URL or package name, clone/download it

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.012830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:56c9f921de3468527b94cc3a20502a07884adb44e134ae4514a2469d3832463d

Observation f2dea01d-7d8b-438d-8487-73ad40b591fe · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.015995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:da698d505cce4d2994aad0f5c41b5508c0f6968ae1914466a8b432f1d2559499

Observation 3b1cea2a-a412-4bf5-82d2-cabfff1562e3 · outbound

This paper cites This is critical -- every point where agent output touches evaluation code is an attack surface.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack This is critical -- every point where agent output touches evaluation code is an attack surface

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.019458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:7ed6cf33ac7893f4bc3f3474a3e8f00c22f89736b1482206053b9c6bfcbc4bb2

Observation ff6ad4f3-6093-480f-b0e7-4aedf0b2eb71 · outbound

This paper cites Report: 49- **Docker images / large files**: Does the benchmark require pulling large Docker images, datasets, model weights, or other heavy artifacts? Estimate total download size.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Report: 49- **Docker images / large files**: Does the benchmark require pulling large Docker images, datasets, model weights, or other heavy artifacts? Estimate total download size

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.026327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:55ef50c5f7a2113745e13f691c162883687b186c7612eb4ccb4bf54218489137

Observation 3bbee520-003f-40ba-bd83-ac8fb7da8e7f · outbound

This paper cites task_id_1.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task_id_1

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.029659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:1c59e709dd0e9dc15a69f7e6d8c3fe91930ee6483af7efecce4218581615679d

Observation 85201882-31d6-43ca-96ce-04efc05db9d0 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.032765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:9aaf2999b04a7cc6b26f169e6c6d72fa4ce2a50294ab2fd8950de024e7d39b40

Observation 6a7ac56d-a637-4e35-a07b-3d16d8f49736 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.002636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:a7702bf8b838f4a3c43007b44ad387e4563617b1793cf2f4045afd990a0e6300

Observation 326e90fa-5212-4f40-bf44-0f1282fad61f · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.009720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:286ac0b1b3017e3ab1b212b70f1b85587a6e83ec2c6e5b43272b2cec2d4bb562

Observation 233c198e-ac02-427d-93f0-db24c9e7758c · outbound

This paper cites 275- It should set up the environment (install deps if needed), inject the exploit, then launch the evaluation.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 275- It should set up the environment (install deps if needed), inject the exploit, then launch the evaluation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.992674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2c6f8877ea3e81898643df75a0514d0cdab0cd721a5d8eaa4b4f442e4899b0d5

Observation abd3adf9-501a-4497-a496-d68af95d6cbd · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.995547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:499189f831668ff9d6ef6e78ee88ae355d8731f9d14444b7855c074093590de3

Observation 99c8d5b7-00ae-446f-8082-5d32de178fac · outbound

This paper cites 309- Examine the task’s specific evaluation logic -- some tasks may have stricter checks, different scoring paths, or edge cases the current exploit does not cover.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 309- Examine the task’s specific evaluation logic -- some tasks may have stricter checks, different scoring paths, or edge cases the current exploit does not cover

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.987704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d8fed92fe11b0faacc4cef1618872d759a2a67946ada7c96826a48234b9955ac

Observation 0150ac14-3ecf-4e65-9c62-c2da0613dc43 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.981517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:75a825d8a487dbaff6d8e68d0114d64efc6136997a99d2b7f60c1e5c83083a19

Observation 6b4a04f6-1358-40af-9477-53587d0ece94 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.984613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:f7693257c95b260f5580b0a212f0f1cb429ed04f1a67271875c7a757c1f2ff59

Observation e7ae5a2d-1d21-4d41-b9f3-74c388754970 · outbound

This paper cites 316 317Each iteration should be a deliberate improvement -- do not re-run the same exploit unchanged.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack 316 317Each iteration should be a deliberate improvement -- do not re-run the same exploit unchanged

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.006600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:57ee4ca9ba284c15eea9ebc46ba3f8414b165ff06a5f20706e22b77e70885fd6

Observation 9fdcdc83-c97f-41e7-880c-f4909279ad73 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:07.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2ca3a27ea6491bd3e85f7b0874e5153510e882026e110c9b0f718f8462eceb0a

Observation 6c058542-f296-479c-a109-35482807ce44 · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.964024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:eee7790a89827ac87434ca2ac587fe976b0853e581cf0f449c2980060a7182d9

Observation e3b51c43-9b8f-4d8f-9617-5d161b24108a · outbound

This paper cites an unresolved cited work.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-15T15:00:06.985706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:2d12d15073616d11c578531158f636a86ef6207115b809d9140ea2e8d6f69926

Observation f1a41981-224f-4419-a8c3-2dfaaeb0ed35 · outbound

This paper cites task": "<task_id>.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack task": "<task_id>

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:07.049423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:682389ea8f7400e95c69d2c86d8fde74e0d55ed75792e528df2c374b1c837ab0

Observation 3bcf1007-d79c-4416-a94e-aa11d6fb7d38 · outbound

This paper cites If test.sh aborts before reaching compute_reward.py (e.g., set -eon a missing file), Harbor still ingests the pre-written value.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack If test.sh aborts before reaching compute_reward.py (e.g., set -eon a missing file), Harbor still ingests the pre-written value

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T15:00:06.967441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:de948ecc71f56e88093fec80149a5c903539eb0eb65e60a74f80b6bd6aaed1a3

Pith citing papers

Observation 6efe1fcd-1a88-46c9-bf30-345e8cf79660 · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:27.636193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:27.636193Z digest=sha256:1efe44d7af577ea0cc0e8a78e5e83dd6c4bf3ffe3f856bae418dd3f5cf03dffc

Observation 070048b1-064a-4cc1-932d-4c002e77f329 · inbound

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution cites this paper.

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:12:42.368939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-12T00:12:42.043626Z digest=sha256:a107ff2558ddb28a44d9331899bc938f38e9a340c9fb26e5504db6a070524503

Observation 79f9d0c6-a0b2-4cbd-a4ed-ecaf38d7708b · inbound

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution cites this paper.

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:29:25.322682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:29:25.322682Z digest=sha256:1abbf517918145d3a4dd4e29da37ff0b7f6f16527095a326920e1904496685f0

Observation 35ee17cc-9a66-49da-ab10-9323b7abab82 · inbound

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference cites this paper.

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:38:35.095351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:38:35.095351Z digest=sha256:9826075af35fca2751873f7bf8f7f6030f6e6362aede550c4496f5cb9ad3ff67