Pith. sign in

Paper Citation Record · LEDGER

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2507.09063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09063 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:08:50.185042Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:25:09.639122Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:59:37.262688Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50d5f51c-86b7-461d-bbf0-380d477824ac · outbound

This paper cites CodeMirage: Hallucinations in Code Generated by Large Language Models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments CodeMirage: Hallucinations in Code Generated by Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:42.937431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:42.937431Z digest=sha256:47094453565b97cda9ac30ed43bfd58319909cf8ae956a1b3353242f55d1e8a2

Observation 9297bddf-9034-42d8-a918-4095eea14f66 · outbound

This paper cites Aider code editing.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Aider code editing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:53.176750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:43.101216Z digest=sha256:d6d2710b51a741b1794b1fc67b8c5d7af4e623cbce0e6f94de52ca57a148032d

Observation 6da50e90-6837-4278-b2a5-d5a6b69f218d · outbound

This paper cites Introducing claude 4.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing claude 4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:53.007049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:43.534078Z digest=sha256:9b5d1cf9d6b87b496dac5996b2ddd503c1b5fca11633bf3b87f872774d656d4b

Observation 3890ee47-51a6-443c-b208-cb69d2cba843 · outbound

This paper cites Meet devin, the first ai software engineer.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet devin, the first ai software engineer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.841914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:44.557141Z digest=sha256:cc3d65f28d14509613040de66e16c72fae4809b6b9c1c62a90592ed513d82f7f

Observation d879eca9-43f8-45dd-88cc-d681e778c0de · outbound

This paper cites Envbench: A benchmark for automated environment setup.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Envbench: A benchmark for automated environment setup

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.672129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:45.312916Z digest=sha256:b1b5049af51aa9eea25abbf1395a77130194c25f63b19ae6ed2c8cfb63b43b0c

Observation e48daeb8-4b10-4042-91ba-34afbfcdc666 · outbound

This paper cites Code completions with github copilot.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Code completions with github copilot

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.561243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:46.040321Z digest=sha256:77fd9b46fd129c72efc2c8a83bf47f894103e31998c68fcc07a9d4dd8f4cc9ad

Observation 033ba43f-9853-492f-a666-53e636a432df · outbound

This paper cites Meet the new github copilot coding agent.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Meet the new github copilot coding agent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:46.588514Z digest=sha256:7acf231699444854116e23b8ca5823c127893e9ba062c56210f056930f95bb57

Observation e125de41-40e0-4d9b-ac1e-2b951395e8a3 · outbound

This paper cites S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments S table T ool B ench: Towards stable large-scale benchmarking on tool learning of large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.856121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.856121Z digest=sha256:8a59c2b95932349f3b207966a4f64c5605088e54340f821a48d101ce1b543e07

Observation 37340459-c733-4faf-be3e-941b902b7aa6 · outbound

This paper cites Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:46.990851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:46.990851Z digest=sha256:9f8c3d83b2d500b748e611e870f0e8f61a6bedc0afce6e5023147f4dcbbd3700

Observation 64aadbcc-7bd0-4e5e-b9fd-a6a96feeaaf6 · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.207136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.207136Z digest=sha256:1c54d6b0d07ce487d7f5a993d50a3c2b3ea57dbc27b2fea6e91bb0ffaab45d0d

Observation 189efb8e-726d-4d8c-a012-748ab8cbcddd · outbound

This paper cites LADs: Leveraging LLMs for AI-Driven DevOps.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments LADs: Leveraging LLMs for AI-Driven DevOps

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.452716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.452716Z digest=sha256:124e12a295d18f3d6706b9b5f0dff376d1414e36beca1263c55c1d49eae1ed0a

Observation 2e45cd78-68a0-416c-b7a8-3838f9ccc14d · outbound

This paper cites Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.747783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.747783Z digest=sha256:47cef4f0f70d44bf44a36424dfe4778013401f53439a17583bbbaef078eaeaed

Observation ef28d23c-67e6-42d2-ae3b-65b2ce6c1acb · outbound

This paper cites ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:47.901841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:47.901841Z digest=sha256:6fcc376f22b941d08f47083eab0c6cbb545f712a99a8e2c75e06ba34fba55f68

Observation a1b86a78-b459-40bb-ae34-1b76186d72e9 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Long-context LLMs Struggle with Long In-context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.218778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.218778Z digest=sha256:ba6f21e0b8d6f8d17d0c6b89a9b60c165b3b11ca087cc1fd1eb909ce2785eb93

Observation 172e390e-a508-4fed-81b8-56cdc56880a2 · outbound

This paper cites NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.471016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.471016Z digest=sha256:65592f778a8a31d7285d6f209e381589e291d3c1ff82d9a9743b7cf8f8552658

Observation b22a4f2a-f6e4-41ac-ba8b-904495d780e1 · outbound

This paper cites GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments GitGoodBench: A Novel Benchmark For Evaluating Agentic Performance On Git

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:48.894069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:48.894069Z digest=sha256:5dacd107845186aa5ca2078eedd97313bf5256573fdbdc296b90efa2ce1f8a4d

Observation 1f1c29eb-498d-45fc-8586-90fdbe2c509a · outbound

This paper cites Agent B ench: Evaluating LLM s as agents.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Agent B ench: Evaluating LLM s as agents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:52.185296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.409404Z digest=sha256:15a906163982095a1db784bae22d6ac74e65c503b4c77389ebc38e0f9fcb3ff0

Observation 465e73db-8c85-484f-bde7-400bdb029460 · outbound

This paper cites OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments OpsEval: A Comprehensive IT Operations Benchmark Suite for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:49.556866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:49.556866Z digest=sha256:f86fa86bdc94631599b590ff1745079834d227cdb56c5931d2164942cd93b2c6

Observation f4d4ff97-6a2b-4a42-9c02-08faba3b1908 · outbound

This paper cites Beyond pip install : Evaluating llm agents for the automated installation of python projects.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Beyond pip install : Evaluating llm agents for the automated installation of python projects

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.886387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.657176Z digest=sha256:7041ffd472772888ddcb94f24a5fe4ec3c79ad22bc3ca1bdb871fe1ec6013e2e

Observation 873b81fd-f42a-4b8f-b0ec-65c6c81668f8 · outbound

This paper cites Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.730388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.760261Z digest=sha256:80ba81d11c5588afc7d83725d33658902ff5583ee30c95aedf497dfb1c5523a4

Observation d1f09481-964e-487a-a003-74c3ccfc8592 · outbound

This paper cites Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Mitigating Configuration Differences Between Development and Production Environments: A Catalog of Strategies

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.560978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.838672Z digest=sha256:fcecb83fa82972c6c46ca8d3655189636f2ad79fb62af01ad81c4ca8e4d8835c

Observation a745daca-bc41-4f02-b73e-483b5074a561 · outbound

This paper cites Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Identifying Factors Contributing to Bad Days for Software Developers: A Mixed Methods Study

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:08:50.393493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:49.949818Z digest=sha256:2a36814bf5f176a0727023b26f6f928f1edc55c0209ec013f81da5e3667133d9

Observation 1f59e4bb-8f10-4883-9797-937ad07ba375 · outbound

This paper cites Introducing codex.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Introducing codex

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.532796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:50.015354Z digest=sha256:678d64d8893c8a36d38f5df36ef0da57d427238f8038cc2c6817dd240d9cd94b

Observation 6c10f197-0d49-4891-b04d-e5f550e034ff · outbound

This paper cites Tool LLM : Facilitating large language models to master 16000+ real-world API s.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Tool LLM : Facilitating large language models to master 16000+ real-world API s

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:08:50.078862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:08:50.078862Z digest=sha256:a32ed4bf08b43717ff94bab3e4d3418ad55280c4b3ea4fd73e1326f894fdf32c

Observation 14cdc954-7a61-4ff9-ae2b-9710a708ce93 · outbound

This paper cites Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models.

SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments Retrieval models aren't tool-savvy: Benchmarking tool retrieval for large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:08:51.106233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:08:50.185042Z digest=sha256:e84a55ddd78798fcf9cae6f34a696d27121782dff1a20746dcd1736c9ca58f7f

Pith citing papers

Observation ffbff7db-862b-4a39-a151-2dd1eed36654 · inbound

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge cites this paper.

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:08:41.972644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T17:04:07.449733Z digest=sha256:e066db028df0a4b2ea31b9cbe444d8f516711d316a06fc018fbb33f614212107

Observation a46cff9f-fb2b-450e-a320-c89f9805d8d8 · inbound

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment cites this paper.

DeployBench: Benchmarking LLM Agents for Research Artifact Deployment SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:16:49.263539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T05:35:12.027697Z digest=sha256:3145b1ab83ddfe97dece75a873ecf3f899efa7e1337c427d7e0c4b8042f2f0c6

Observation 019662f6-356d-4c7d-9f4c-df963549c716 · inbound

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer cites this paper.

Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.264275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:02:37.404821Z digest=sha256:90cb4db74259cf41ed6533a94054fe2b8d691771bc27cc7c690b5804be62a2cb

Observation d0300806-5c46-477b-aae5-08da90ed168a · inbound

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents cites this paper.

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:48.608177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T01:19:54.081755Z digest=sha256:d1fc704cfb33c6020e7c4a54be25d5a7ecd66ea3e0e2aa7e4b4e853bfb89bc4f

Observation e5abfd60-86ad-4964-a30f-6ce9e6e87fca · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:09.639122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:09.639122Z digest=sha256:8d8c4942571c604855c2c8798bd16c9261481c88e7935f7d87cbf0b825bd05d8