Pith. sign in

Paper Citation Record · LEDGER

RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.07832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07832 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:55:46.635284Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c9ca7b5-642e-46a3-8476-9a6b21702808 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.452526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:c9bd916acd4ea62a1a2426dbc2f17eca9ff971b11ae19cc2d6784eccf0c41ac7

Observation 6934c22c-1f5c-4c9e-8eb2-ce0e23850d7a · inbound

Evaluating Code Reasoning Abilities of Large Language Models Under Real-World Settings cites this paper.

Evaluating Code Reasoning Abilities of Large Language Models Under Real-World Settings RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:34.165493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T21:23:44.762007Z digest=sha256:0127fc458d9095e1b3ff7f7f7690297f394a9fe38fdea5447c0b6eb5a0dbb0d1

Observation 9c7a203f-a7cb-40dd-bd56-b021951e9c3b · inbound

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair cites this paper.

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:49.880043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:05:32.117271Z digest=sha256:95124482e2f80d4c24decee34f48ee3d39ed54f1b2762f33e74fd88d118782df

Observation cdd6732f-5408-437a-87e4-3ba143d72931 · inbound

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair cites this paper.

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:21.871897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:58:41.120953Z digest=sha256:f5e3732fda2bfa4a691f9d7afa2c1881c9b19cf87def0b6c2affaf462773145e

Observation 9229fe0f-68cf-44af-b4c6-d7cc00c0e079 · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:56.777009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:e548858c7377678874500f10fc7a57ed3cc6bc4e72dbeefde78a278b2a9575e6

Observation 5d898c85-7161-4c56-9e36-4a9da96ff680 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.109425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:4365e7be0daa147922be20d0f0717ead7906ba9d33bf5ebed3d0f003a3b23352

Observation cab455ad-1a6a-4a53-b9a2-d30927ee470b · inbound

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search cites this paper.

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:54:05.931493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:51:59.101930Z digest=sha256:574588320df7610bf1b02c83af170dd85a024712613ae99c0b164a2f106684a4

Observation d9c2560a-8ebe-49ee-8a6d-6bb4b26e09fb · inbound

SmellBench: Towards Fine-Grained Evaluation of Code Agents on Refactoring Tasks cites this paper.

SmellBench: Towards Fine-Grained Evaluation of Code Agents on Refactoring Tasks RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:56:59.492628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T00:50:43.941064Z digest=sha256:6894398328d29413f292e91280d249ac1bf74b67a8dd4597fa5ea9f81e515e43

Observation 6c53dfeb-cd32-439a-ad2f-0b8bd088aaea · inbound

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation cites this paper.

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T05:40:26.231504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:40:26.231504Z digest=sha256:5db1d801cf4207bebe5009c9f453d38478aeea77e13468da67c9b5d40bc5f65e

Observation 87664a74-9708-45fa-8451-4e797bf6d067 · inbound

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation cites this paper.

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:46.635284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:46.635284Z digest=sha256:2009645377095b147acee93812ee4ca11e54f75be532af3d5a6f1b028239619f