Pith. sign in

Paper Citation Record · LEDGER

Evo-Bench: Can Language Models Improve Agent Harness?

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2608.09096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09096 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:22:26.870215Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21a47d1a-1229-447d-8e70-15b6cd4707a0 · outbound

This paper cites HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry.

Evo-Bench: Can Language Models Improve Agent Harness? HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.628134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.628134Z digest=sha256:e0ad8b60b66c34268082792cf28255d7cf08edeb54151520ed7dbb3768e6a1b1

Observation b94458d9-0d25-41da-bf49-88c57a490cd2 · outbound

This paper cites 11 Evo-Bench: Can Language Models Improve Agent Harness? DeepSeek-AI.

Evo-Bench: Can Language Models Improve Agent Harness? 11 Evo-Bench: Can Language Models Improve Agent Harness? DeepSeek-AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.636498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.636498Z digest=sha256:24d88cd99c59c1e07468356880f96979208761f77e5675c14b62fa9b2d943ce7

Observation 192edf02-eef4-49db-98f1-2d6ce7af8406 · outbound

This paper cites SIA: Self Improving AI with Harness & Weight Updates.

Evo-Bench: Can Language Models Improve Agent Harness? SIA: Self Improving AI with Harness & Weight Updates

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.644350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.644350Z digest=sha256:7267167a23a5e7fd3c3ce90120c9a1b5f5525b0530357793df626dddc6a9bdb6

Observation 542b4358-d5d5-4cfc-baaf-7ee47c9cd90d · outbound

This paper cites Automated Design of Agentic Systems.

Evo-Bench: Can Language Models Improve Agent Harness? Automated Design of Agentic Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.648963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.648963Z digest=sha256:6f1824763c93328245b17e270adc957b43825df696792e534a2028ad086d26a7

Observation 5f998b68-8bfa-4d4e-b156-2a4bcfd2b75a · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Evo-Bench: Can Language Models Improve Agent Harness? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.653121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.653121Z digest=sha256:4abbde2fe54ea6d562ce1de828e47508d4ecb2cbd89927e408173370b8011549

Observation 7d91f3af-8392-4574-a5e8-d0ebd3653329 · outbound

This paper cites MiniMax Sparse Attention.

Evo-Bench: Can Language Models Improve Agent Harness? MiniMax Sparse Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.661868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.661868Z digest=sha256:384d30e132e4453f830957d3cd489c309fa020a82feb4fc2f646d57cbc987f79

Observation 9a3a6329-b9da-45e3-85f6-f93c83c20bf1 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

Evo-Bench: Can Language Models Improve Agent Harness? Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.665740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.665740Z digest=sha256:71356e85253f7228a60f1bef8811eb89dec106fb9a1a4307bb8461b28ef76260

Observation 60e17c74-e933-49df-965a-8b850b4db3d5 · outbound

This paper cites ClawEnvKit: Automatic Environment Generation for Claw-Like Agents.

Evo-Bench: Can Language Models Improve Agent Harness? ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.670161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.670161Z digest=sha256:94266d7d42fdb02eda03a9c5e951ce7a07c39b2a8238c9e6fec1e29350eeee05

Observation f1e14da7-cf84-4e04-974c-c61c67ca037d · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

Evo-Bench: Can Language Models Improve Agent Harness? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.673955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.673955Z digest=sha256:031b85b93eddcb8afe1004f98c1ff2ee1013d8bd49571e309761a3e5a7052b95

Observation 0dc64604-d4fd-4563-9004-aaf627cdceee · outbound

This paper cites The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?.

Evo-Bench: Can Language Models Improve Agent Harness? The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.679367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.679367Z digest=sha256:f9b0f207a55c96b4dfce8d568a9194f39da700057fb0e276c41af7273f754de9

Observation 41b0c0d4-d6b0-45ed-9a06-eb006d0cb72c · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Evo-Bench: Can Language Models Improve Agent Harness? AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.683106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.683106Z digest=sha256:4134e4645c88400214de2bb5d5030f0c12014afa93dd1299bdf63e51310d6e92

Observation ae91a44e-5ae1-41f6-b4ef-1a20507e0424 · outbound

This paper cites 16, 2025; cloud-agent research preview announced May 16,.

Evo-Bench: Can Language Models Improve Agent Harness? 16, 2025; cloud-agent research preview announced May 16,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.686828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.686828Z digest=sha256:3aa3401eced7b2ee9c21f3af21b2b2181ec296e40263974cf86a8b6cc7e79924

Observation 98c760b8-a68b-4f7c-acaf-fef3972e3371 · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Evo-Bench: Can Language Models Improve Agent Harness? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.691164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.691164Z digest=sha256:0d35c89b7c3a931d913361c96a79429c279cb99e12d1d966eaefcee152f656cf

Observation c95b22f6-3d20-4ced-b076-616f405caaf0 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026a.

Evo-Bench: Can Language Models Improve Agent Harness? Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026a

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.696626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.696626Z digest=sha256:9732edfc38973760d90944f61141368d20fa1dfe74592dc9955759317b632479

Observation cfc401ee-8430-4d30-90e9-ab59b4f4e620 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Evo-Bench: Can Language Models Improve Agent Harness? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.700377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.700377Z digest=sha256:b8f7d9f529746442cf8d7595d9d9bb5e0b2608231ea7567d57cd8ad81ccd122e

Observation f739b9d8-5d58-4ebe-a8d2-e9c415888516 · outbound

This paper cites Gemma 4 Technical Report.

Evo-Bench: Can Language Models Improve Agent Harness? Gemma 4 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.705407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.705407Z digest=sha256:b6d42e727147db980525ad7dfb2ed6abc65ea33d3c80d03525368ea69bd0cc84

Observation 0174f3a2-cb08-46a2-8fcf-1240ab09803f · outbound

This paper cites MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling.

Evo-Bench: Can Language Models Improve Agent Harness? MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.710180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.710180Z digest=sha256:8405cc0c4a05f4dd574035c54d5309a9b2449d0999142d0c4c2b9a5f74c87ac7

Observation 108dbe65-6a23-495c-9c94-9e21f0963f76 · outbound

This paper cites VeRO: A Harness for Agents to Optimize Agents.

Evo-Bench: Can Language Models Improve Agent Harness? VeRO: A Harness for Agents to Optimize Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.714960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.714960Z digest=sha256:ea1de0720c5dcf071fdfb2b3ae5aaa3dedd233cc6eec5e0d4000058e0f8cee6e

Observation def9e4d5-b117-45cf-9016-6c6da82b2285 · outbound

This paper cites Apex-agents.arXiv preprint arXiv:2601.14242,.

Evo-Bench: Can Language Models Improve Agent Harness? Apex-agents.arXiv preprint arXiv:2601.14242,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.730007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.730007Z digest=sha256:0c320bd6de3693034c431a74146dd0a418aede7875a388e24538e38d4829a740

Observation 4308d9a8-9172-47a9-b55f-c1eacca0b19e · outbound

This paper cites Rethinking the Evaluation of Harness Evolution for Agents.

Evo-Bench: Can Language Models Improve Agent Harness? Rethinking the Evaluation of Harness Evolution for Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.761061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.761061Z digest=sha256:5b5bb865c218c10956e61f4f250b18ff44298b9754237786c78aaac201b1f4a6

Observation 6c557bd6-57e4-4451-8b03-2bd022e647fa · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Evo-Bench: Can Language Models Improve Agent Harness? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.804133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.804133Z digest=sha256:e96823e89645340f7c658736b695a2b6baef127e092b4055c9afe4b5592ad449

Observation 40910431-7f02-4f95-ab99-34b730e701de · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.843491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.843491Z digest=sha256:45f605f8fa7751d546013a3367e6f2a3e287464652c96bf0a47f673e49fb73a8

Observation f1477e90-55bf-4874-b626-25d1c27f2edf · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evo-Bench: Can Language Models Improve Agent Harness? $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.848374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.848374Z digest=sha256:a23eb3bbb650c1fdf322107d4d04df02a3f81a74f9e92f97c09aacacb461a09f

Observation f0df5ac0-0714-4f99-9677-2070dfceb393 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evo-Bench: Can Language Models Improve Agent Harness? $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.851933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.851933Z digest=sha256:cb860fd601ecdd8a41b3304eaac2324549475206539ddec8a3da0d2e63823fb3

Observation 3206dd85-11ef-448a-a448-3c52425eeef9 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Evo-Bench: Can Language Models Improve Agent Harness? GLM-5: from Vibe Coding to Agentic Engineering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.856839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.856839Z digest=sha256:939851d338f4530dd1e153c6dbe7ce08f0e46ee9391f4c329845a71264c0c610

Observation ef8392bb-4f72-42ca-baca-929458df42b2 · outbound

This paper cites Self-Harness: Harnesses That Improve Themselves.

Evo-Bench: Can Language Models Improve Agent Harness? Self-Harness: Harnesses That Improve Themselves

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.860842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.860842Z digest=sha256:f084c09c9358e24d5f2d1c656fd80317adc17d37a8c03b68c148bad6a74a3a5c

Observation 77191e59-d1ca-4303-b48b-1e2363aa591e · outbound

This paper cites 16 A.2 Details of the Policy Harness.

Evo-Bench: Can Language Models Improve Agent Harness? 16 A.2 Details of the Policy Harness

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.865501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.865501Z digest=sha256:ab98fd3f457df7e98c08cbe897a53634d3eb2b50db8c9178a3a170d3209c684d

Observation 3f3e5601-2092-4b19-952a-7e640d960733 · outbound

This paper cites Domain tools, planning, memory, and verification are left for evolution.

Evo-Bench: Can Language Models Improve Agent Harness? Domain tools, planning, memory, and verification are left for evolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.870215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.870215Z digest=sha256:40346adaee981fc95a7bbd683b4758786958b857147f1e9a3e66040d1768e3ba

Observation f98ced87-c1ac-48b5-ab92-2d30adcc738a · outbound

This paper cites SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment.

Evo-Bench: Can Language Models Improve Agent Harness? SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.657509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.657509Z digest=sha256:068c77f9c14e3a98f2ec0727bc61b3dd338f106261277327898e51d74414e18d

Observation a1d1bcf8-d99b-489d-a183-23f4ca6af170 · outbound

This paper cites EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer.

Evo-Bench: Can Language Models Improve Agent Harness? EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.640162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.640162Z digest=sha256:f3c0b4bd5d0b9a9998ea962c842ccc80400dee3371712812d9d579116a5966d1

Observation 65bbff88-96b1-4ed7-974c-958d79cbae98 · outbound

This paper cites 24, 2025; general availability May 22,.

Evo-Bench: Can Language Models Improve Agent Harness? 24, 2025; general availability May 22,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.620002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.620002Z digest=sha256:8b0841ba6f7f1fac2a30195a529aaf917e8b88bd468f09a0fbe33f186ad0d28f

Observation fffa250f-dc3b-462d-8df6-8ba7d009ff76 · outbound

This paper cites Mle-bench: Evaluating machine learn- ing agents on machine learning engineering.

Evo-Bench: Can Language Models Improve Agent Harness? Mle-bench: Evaluating machine learn- ing agents on machine learning engineering

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.624677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.624677Z digest=sha256:5ce012fdfa5124bfa52b41b55c7a0b9544d17e8640e6b0d4812a3d3af36ec1ac

Pith citing papers

No inbound Pith citation observations are available.