Pith. sign in

Paper Citation Record · LEDGER

SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.12952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12952 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:01:27.898502Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ebc691d-d01c-46f6-bb7d-91bd5fee1899 · inbound

COFFE: A Code Efficiency Benchmark for Code Generation cites this paper.

COFFE: A Code Efficiency Benchmark for Code Generation SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T11:01:27.898502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:01:27.898502Z digest=sha256:b70903f921b1d40bccacfdf18fbb3eb7f39ef5261ba1bdeb460c273aba9b99f2

Observation 7d6d7dab-37ac-48d0-9168-dbbca55e98f7 · inbound

Agentic AI Systems Applied to tasks in Financial Services: Modeling and model risk management crews cites this paper.

Agentic AI Systems Applied to tasks in Financial Services: Modeling and model risk management crews SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T19:23:59.900714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:23:59.900714Z digest=sha256:466518e0720aedb9a58c5b69dddb3285ad52604651a3f94660cd2140251b6f95

Observation 33dee1b0-7eec-4ef4-9ec5-702ca5051ece · inbound

CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification cites this paper.

CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:26.034044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:26.034044Z digest=sha256:873ab4316f2590dfd814dd828c92a7197614f90cf5714aa832674dbf05be291a

Observation 1342ece8-4ec1-4b9f-8338-032bfcab9e21 · inbound

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification cites this paper.

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.898783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:31:36.233977Z digest=sha256:21c2b088da53b85a441ce90ed94f2fad4f1bd1fe9047bd9dba64a6abb53e62f9

Observation 7294be2b-68e9-4aed-94fb-8c8b77eda287 · inbound

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification cites this paper.

FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T18:43:45.603737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:43:45.603737Z digest=sha256:388db8768be42e82bb55134f57159748e9f48cc3286285b84b7f9c58d67f3aaa

Observation b1886534-7030-43f2-93f3-4d1fbe8413af · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.436996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:38:12.144252Z digest=sha256:d9e2f256bf646bf71ed17314eaf8a71f86f187600c5a75c40a98996adc83ad8d

Observation b80b8a8b-75e6-4795-8fd5-67884241e464 · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:44.566904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:44.566904Z digest=sha256:a7927bf9dbb344d1fc3b372e09a091703b8a5683d270bedbaa44ef48f13579ef

Observation e77c6c09-3d33-4557-90f5-5153c535b6d6 · inbound

Vibe-Coding: Feedback-Based Automated Verification with no Human Code Inspection, a Feasibility Study cites this paper.

Vibe-Coding: Feedback-Based Automated Verification with no Human Code Inspection, a Feasibility Study SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.439528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:17:56.139148Z digest=sha256:66b1953d9db9aa75b020c6a1545336e67a77f52eb101a29505cf3fe7ace83935

Observation 0325c06c-79f0-4450-b662-d01a85937647 · inbound

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution cites this paper.

Breaking, Stale, or Missing? Benchmarking Coding Agents on Project-Level Test Evolution SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:12.077714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T09:04:06.347354Z digest=sha256:fd94e234d9931c16166e46ac209096449989dbcb0aae503f1244129eebbde6d5

Observation dac780e1-9e67-47ba-99ec-4ba292dc3769 · inbound

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents cites this paper.

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:01:39.711128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T04:58:10.803420Z digest=sha256:212c5a99389f80f1e4f8614da2e63a0eb6a2651c43240797a297fe9448d1e0bc

Observation 9dadea3b-079a-4243-927b-1cc5a0081008 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 149

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:02.945964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:d54cc687bb0de5248aff17594d5836ddc2ac93a46cc79cdbc75c8c7e1dda0a16

Observation 4b051cd0-eb9d-4680-9106-654078c7e260 · inbound

ExplainBench: Evaluating Code Explanations from Agents cites this paper.

ExplainBench: Evaluating Code Explanations from Agents SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:36:23.269519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:36:23.269519Z digest=sha256:97b676754af5c9e568ec458fb6b2429848e48c08c92eb769b3749a9ffc6492b7