Pith. sign in

Paper Citation Record · LEDGER

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

As of 11 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2507.09063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09063 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:50.185042Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:25:09.639122Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:59:37.262688Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50d5f51c-86b7-461d-bbf0-380d477824ac · outbound

This paper cites CodeMirage: Hallucinations in Code Generated by Large Language Models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments CodeMirage: Hallucinations in Code Generated by Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:42.937431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:42.937431Z digest=sha256:7626c29fc496f33714614fe54963db46893a39e6089c3d5fff71708a312a7442

Observation 9297bddf-9034-42d8-a918-4095eea14f66 · outbound

This paper cites Aider code editing.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Aider code editing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:53.176750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:43.101216Z digest=sha256:8f54909a0c5457235c18417a5e1658bfcde1b8cec5d64638e11358ccf2569b6d

Observation 6da50e90-6837-4278-b2a5-d5a6b69f218d · outbound

This paper cites Introducing claude 4.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing claude 4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:53.007049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:43.534078Z digest=sha256:e7e64377f0c5997d1aa3dca4f6d0bec5d70a2192d3570c6449425fc68a0aa631

Observation 3890ee47-51a6-443c-b208-cb69d2cba843 · outbound

This paper cites Meet devin, the first ai software engineer.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet devin, the first ai software engineer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.841914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:44.557141Z digest=sha256:df25743082fdf25460656b5adc4efb10c15e6cc0b2a5e20467142b811ded9ac7

Observation d879eca9-43f8-45dd-88cc-d681e778c0de · outbound

This paper cites Envbench: A benchmark for automated environment setup.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Envbench: A benchmark for automated environment setup

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.672129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:45.312916Z digest=sha256:619d3e73866d11213c76a1d536579ad6b46a7bddcb5839bc0877528793e5141a

Observation e48daeb8-4b10-4042-91ba-34afbfcdc666 · outbound

This paper cites Code completions with github copilot.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Code completions with github copilot

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.561243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:46.040321Z digest=sha256:9b17f46111c49106d645fd1670704507ed33d613db1688693f9d27b73d6e4dd8

Observation 033ba43f-9853-492f-a666-53e636a432df · outbound

This paper cites Meet the new github copilot coding agent.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet the new github copilot coding agent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:46.588514Z digest=sha256:73b620b58009144afc44036922bbbbea5c6693227fe87e593899b9688917958a

Observation e125de41-40e0-4d9b-ac1e-2b951395e8a3 · outbound

This paper cites S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.856121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.856121Z digest=sha256:c337d7e1d101338acf248e0b2a09a623420b28a336eb49a94539d5a7b0b06f92

Observation 37340459-c733-4faf-be3e-941b902b7aa6 · outbound

This paper cites Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.990851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.990851Z digest=sha256:c14bd8490885abb54e065ca8258e4e0665d2b6810f64d1119c45951306eb5947

Observation 64aadbcc-7bd0-4e5e-b9fd-a6a96feeaaf6 · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.207136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.207136Z digest=sha256:4ccfaf80b1da481f01129342ac518f295ce019d2c7d41a3c58407f8543372aa9

Observation 189efb8e-726d-4d8c-a012-748ab8cbcddd · outbound

This paper cites LADs: Leveraging LLMs for AI-Driven DevOps.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments LADs: Leveraging LLMs for AI-Driven DevOps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.452716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.452716Z digest=sha256:55dd9f8cc498e215392099dcb29fed2de6ed78a410f18ccb975f5e59b92f2273

Observation 2e45cd78-68a0-416c-b7a8-3838f9ccc14d · outbound

This paper cites Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.747783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.747783Z digest=sha256:93d0bd53abd357b1ffcd0486f47317305d8023fa5f499924539dc14ab2d009d1

Observation ef28d23c-67e6-42d2-ae3b-65b2ce6c1acb · outbound

This paper cites ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.901841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.901841Z digest=sha256:de4166763762364758d181d7a3c53a46317643be0f95fd65355b4cde08845b91

Observation a1b86a78-b459-40bb-ae34-1b76186d72e9 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Long-context LLMs Struggle with Long In-context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.218778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.218778Z digest=sha256:ce98e545c01553d2cb2d19743761053867e4175566549a95800f71a1ea7315b3

Observation 172e390e-a508-4fed-81b8-56cdc56880a2 · outbound

This paper cites NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.471016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.471016Z digest=sha256:da8025e135dd6931c57e40deea91d2ef5083a40bd4fd907108a04b5cfafe4ed6

Observation b22a4f2a-f6e4-41ac-ba8b-904495d780e1 · outbound

This paper cites GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.894069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.894069Z digest=sha256:1457ddb6df83fbb796b1d8e9462b3ebc759cc416b453058194b71ea5e1e3005a

Observation 1f1c29eb-498d-45fc-8586-90fdbe2c509a · outbound

This paper cites Agent B ench: Evaluating LLM s as agents.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Agent B ench: Evaluating LLM s as agents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.185296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.409404Z digest=sha256:7a216008d9e62fdc2b65f5dade5c01e5332812196c44ffa332e4d3b6387e6a0a

Observation 465e73db-8c85-484f-bde7-400bdb029460 · outbound

This paper cites OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:49.556866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:49.556866Z digest=sha256:36c9c616d13f0fd88f7273043c4511781d421aea2196dc18c4032ef0a72e546a

Observation f4d4ff97-6a2b-4a42-9c02-08faba3b1908 · outbound

This paper cites Beyond pip install : Evaluating llm agents for the automated installation of python projects.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Beyond pip install : Evaluating llm agents for the automated installation of python projects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.886387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.657176Z digest=sha256:4530333f32f65cb8513baf482c0fa57e798a7ba97ea686799b6222141fda8f80

Observation 873b81fd-f42a-4b8f-b0ec-65c6c81668f8 · outbound

This paper cites Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.730388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.760261Z digest=sha256:3e5164ece61318333f61a08b3b485724a2d26f55c5f1ed87bd2a1e84ee2f4005

Observation d1f09481-964e-487a-a003-74c3ccfc8592 · outbound

This paper cites Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.560978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.838672Z digest=sha256:ae4a769037481b30008bc2560f6210ab193890ee1c7f9f4e997083ee53f6bfb1

Observation a745daca-bc41-4f02-b73e-483b5074a561 · outbound

This paper cites Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.393493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.949818Z digest=sha256:7134990e4b4aa879f13b09833b3d2ce67d151a569d6063760e9ce1738c40fccb

Observation 1f59e4bb-8f10-4883-9797-937ad07ba375 · outbound

This paper cites Introducing codex.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing codex

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.532796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:50.015354Z digest=sha256:19655fba18612699e99c84797d56df36a67bfe5d8162665a112db09b5f97b7b8

Observation 6c10f197-0d49-4891-b04d-e5f550e034ff · outbound

This paper cites Tool LLM : Facilitating large language models to master 16000+ real-world API s.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Tool LLM : Facilitating large language models to master 16000+ real-world API s

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:50.078862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:50.078862Z digest=sha256:6e2d6edcfb2ded74d6e73e5d9d02b44788784ce293b8f98efc7f2b3edd5b479a

Observation 14cdc954-7a61-4ff9-ae2b-9710a708ce93 · outbound

This paper cites Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.106233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T18:08:50.185042Z digest=sha256:7c1dc878759cd9f2792c3d563c98677b61252a450f768c9b8e5c1a306abf23ea

Pith citing papers

Observation ffbff7db-862b-4a39-a151-2dd1eed36654 · inbound

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge cites this paper.

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:08:41.972644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T17:04:07.449733Z digest=sha256:4469ed9a38604d8191624d99d1092a6ff7f10c89272c9af217b856a115ce0a19

Observation a46cff9f-fb2b-450e-a320-c89f9805d8d8 · inbound

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment cites this paper.

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:16:49.263539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T05:35:12.027697Z digest=sha256:66407be2958e496454701c2a5512dd7291f32b851923c3b27480cd719632d29f

Observation 019662f6-356d-4c7d-9f4c-df963549c716 · inbound

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer cites this paper.

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.264275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:02:37.404821Z digest=sha256:537a89b5dd9e6fb8d38b9d8f872a0ffe08305813f363fb487331683685642d25

Observation d0300806-5c46-477b-aae5-08da90ed168a · inbound

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents cites this paper.

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:48.608177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-30T01:19:54.081755Z digest=sha256:42d047750590e137e6ae2ca4f61d8324d141e972fa33fc87f54d2d3b3028807b

Observation e5abfd60-86ad-4964-a30f-6ce9e6e87fca · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:09.639122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:09.639122Z digest=sha256:330e9dd67ed68c0f41d83b68f5676e8f2d5d11704938a0d7d1b92a044ac518f8