Pith. sign in

Paper Citation Record · LEDGER

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2605.06125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06125 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T09:04:06.347354Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:47:01.373765Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact9
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c56448ec-fe67-4be4-b095-fed53ac09e85 · outbound

This paper cites TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:12.022865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:9dc11923946f000ab4b53c60e7d5351f79e71b9d13beb215624768056bf5fa41

Observation 0325c06c-79f0-4450-b662-d01a85937647 · outbound

This paper cites SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:12.077714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:c45cf1dd548912e6e6da84f4f1e8da316c7b5770175c0f4aa383112ac7073e8a

Observation 95194f36-9e8e-441c-991a-63883f4a566a · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:12.004352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:051f39e399bc170476f92d82d766439088e3cb82c0d01b185611abd3156e6c22

Observation e5097a72-fd71-451d-b850-25b82c93f511 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T13:48:53.849224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:d40ce60ddce9ae0de162a4c9c37f910c5a9391b5b4ce76fb42c87ee7cfc1b0c9

Observation e46be76a-84d0-46be-87f5-0fc62ff52bef · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution GLM-5: from Vibe Coding to Agentic Engineering

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:12.010118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:8087bcaf8971b55256cffcf4b690f98fb5517333e333351d4090aba0960764e9

Observation 36d55959-6cc0-4892-bcbb-70d17e6c3419 · outbound

This paper cites Just, R., Jalali, D., and Ernst, M.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Just, R., Jalali, D., and Ernst, M

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:28:03.244773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:ff0316ab2ebf5d67cfd500ae8d6dbade64feef8b9e8f55c6909d3bbb26f5a948

Observation c58597ae-8aeb-4413-b338-28b08539beb4 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Kimi K2.5: Visual Agentic Intelligence

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:11.984427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:aa42d5113fa1bcdc484727ccb2bcc3f723f358c02e55c24c7aa736ccba2a46ec

Observation 9ce0dc4d-c8a0-457e-9668-a9a46654f7be · outbound

This paper cites Qwen3 Technical Report.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Qwen3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:26:12.040198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:5749e42360e2b1a8618c05c548769ca1a3ad2ff08caede74790e42b0e4c91c7c

Observation db750414-5a0e-476d-8032-6b5f0586f0f6 · outbound

This paper cites SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:26:12.030598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:01f47b2322c6a3eddd180e1b448c8c348580773c363873ced54d1459975aec91

Observation 911de558-d813-44b2-9a25-b6b6c46b441b · outbound

This paper cites Exotic Topological Phenomena in Chiral Superconducting States on Doped Quantum Spin Hall Insulators with Honeycomb Lattices.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Exotic Topological Phenomena in Chiral Superconducting States on Doped Quantum Spin Hall Insulators with Honeycomb Lattices

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.976093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:eff3b5645655050e3d730f99de3f8435628c525e119c9d704d34b08a92620a31

Observation 2ae7aeed-dde3-43fc-ad02-536464b3ff4d · outbound

This paper cites Advancing Code Coverage: Incorporating Program Analysis with Large Language Models.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Advancing Code Coverage: Incorporating Program Analysis with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.944425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:eca552e559924e5a92086154993836f56e04248483dc305ff704104680020cc4

Observation caa121fa-d441-4354-8d44-8ccb4580c19a · outbound

This paper cites Unit test up- date through LLM-driven context collection and error- type-aware refinement.arXiv preprint arXiv:2509.24419.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution Unit test up- date through LLM-driven context collection and error- type-aware refinement.arXiv preprint arXiv:2509.24419

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:11.962909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:382f82e3ff15439bc44431329c6ccdf90de201ef71e0c41d7dc63c26677aa594

Pith citing papers

Observation 549c5b4b-52d0-4b48-9297-33de4b6f62c9 · inbound

MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs cites this paper.

MultiFixer: A Coordinator-Proposer Based Multi-Agent Framework For Fixing Multi-Hunk Bugs Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T12:47:01.373765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:47:01.373765Z digest=sha256:e977db4b05ba2a91cb5f7384840b133a71aa12ffe94d702922e4071a193d1c14