Pith. sign in

Paper Citation Record · LEDGER

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

As of 4 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 1 inbound Pith citation observation for arXiv:2604.28093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.28093 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T05:10:58.065201Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:25:08.066021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact6
  • verified fuzzy4
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 433f73ab-51e8-409d-8b71-d0aea40d42b2 · outbound

This paper cites Concrete Problems in AI Safety.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Concrete Problems in AI Safety

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.967971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:ef6bc499b2157cd51e10c3f8e794ddd1708036da9ce0f5783653bae53a7c0af2

Observation b94a18b1-ec0b-404f-afae-70e1f4292d29 · outbound

This paper cites Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.961371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:07e9fc2ce2f6de287d2af3d54d85c4fab594aea0ac77f17c8de79c679203550b

Observation 7bd7c221-8fb6-484a-97c1-c3ecd34cc25f · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.977699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:0a4c3b911180441e7ba5a91c9a6caab8c38abbe503fcf3d6cca9f88c3679b974

Observation 4d46291f-fd14-46bd-8ccb-5e1e21ef324d · outbound

This paper cites Dell’Acqua, E.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Dell’Acqua, E

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.983915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:d98d4caa99dc69fd2783284c6e8204ea99de7dcfd90979cb1c4769ceb1c018fd

Observation b4af3104-1a7f-4deb-9d2f-1669321c1666 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:fa2fcb8a8ab06dfffb939315067b3db2d2501974650a694ea4fb1e2652981ae7

Observation f422a5c1-db34-40e8-8818-136a45f0867b · outbound

This paper cites Krakovna, J.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Krakovna, J

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.974312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:973f989c8330c3f11eb7419f350902ffe56cf71a7b0f2fb1f7854a2cebab3324

Observation 3db5aac2-0534-4751-a061-16d7699ed47a · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.971759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:3fed53e93636f90a0b90c38fc07a14a70e2fa77f80519d86970c52cdd165aecf

Observation 1b7cf7a6-cd43-4d3a-949f-0897f251da6e · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:36:30.975678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:a5a926d23be6eb2137b88830341ddfeaf3a51be547f13b75f66fbbb9302dfc83

Observation 8a086f59-ba23-44d2-b88a-926bbcb348a0 · outbound

This paper cites Launching the OpenThoughts- Agent project.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Launching the OpenThoughts- Agent project

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.987471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:b531b0b4b41b0cad66ceba29ddfe46943fddd0f596580d0b5017fec69b893227

Observation 7e9de2b1-74b7-4445-abf0-287fde749d88 · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.967565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:59395040ddab494e3e956d6a4ad4fd76f27597e180021aad7fa4dee7d96deadd

Observation eb01340d-9a2c-4242-9c66-0f175b648b52 · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.990589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:568b3dd37869f6fc06efb9876d09acb1676f3d37b7bb01b358add478f45a8330

Observation 31434c35-f1cf-4fd8-a5dd-cc7b48aa5e9d · outbound

This paper cites Von Arx, L.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Von Arx, L

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T10:49:02.980718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:54435565465ed3e19d14677d096e51d7a679a2fd46a251b007744aa55986970b

Observation e9ceb1b8-34f6-46a6-8db0-6a882ed17d2a · outbound

This paper cites Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.979968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:5fcdfb1814004af718bf06b05c8fe0bfb5fd8b04042d78bb6e3697c3f0e64835

Observation 6e298221-0052-4e78-bc0d-20742c6049ab · outbound

This paper cites an unresolved cited work.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-27T10:49:02.970934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:708b85b7855ba0d6af4f893060d44839423cf1ee9ce59505d78c9af8e0db5c1b

Pith citing papers

Observation 03c64f5c-1682-4b91-a664-c94a02a0bffa · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:08.066021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:08.066021Z digest=sha256:a7eb75c41a49a6cd749ccd38350717573d99bec4bbbc40495a34024715236960