Pith. sign in

Paper Citation Record · LEDGER

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00172 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:17:45.470828Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3c96934-e189-4598-8b84-cb32a3653808 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.198051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.198051Z digest=sha256:517be7c72a1de56d070bb5c808f9e9c286663bc4a8a37864123ad73266919c0c

Observation edf39d07-da64-4c83-a760-260036ed7813 · outbound

This paper cites MARPLE: A Benchmark for Long-Horizon Inference.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents MARPLE: A Benchmark for Long-Horizon Inference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:17:45.870546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:17:44.420152Z digest=sha256:7daea26ce754cfe25a371b3d7d449377a89311233d022e7404d74bf8eb73b5e2

Observation 0db5c82e-6129-49d0-b34c-14192f4f7637 · outbound

This paper cites an unresolved cited work.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.778572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.778572Z digest=sha256:814c36656c179e9154c437100027de95b9dac284dc5f52f2aedde4fdf1d2a49e

Observation 9d2ff84f-c45c-443a-a488-d72430be49dd · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents AgentBench: Evaluating LLMs as Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.888628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.888628Z digest=sha256:f5621cc6d3ece293b40e2d5225294889e89e2cf0807ae58082f2339f1f80ca5e

Observation d664ad6a-ee01-489f-9aac-0f635ddb0cc3 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.997683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.997683Z digest=sha256:0a9dbf14b8d3501a5647030512a9957e6610e38bd1c9b1e9c74d5892ffaf0ce9

Observation 2ff36b65-ec3a-4205-b6ec-6d73fc436d05 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.190655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.190655Z digest=sha256:44cb7b500dc53e8f6bc7eccca90afbe88e870c51008ff90db641bd9a3ef74c64

Observation 13c666ef-947e-41ce-be2f-5852f06b73e0 · outbound

This paper cites Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.278095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.278095Z digest=sha256:12e33fe2634f07ad7ea9d1c54b44b6dfd77e9d18f0e29392e7ea6fd4e667ca03

Observation ab9b37ec-79fd-46c5-afcc-461cd431d942 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.374967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.374967Z digest=sha256:52012824b71538597ea32f8cdf9b27d203a7641851ab3cf422a80820502a59da

Observation 34b13d74-94aa-43d4-ae23-797e0a045e74 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.122364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.122364Z digest=sha256:c9e6bc2415a6de205a8816588c6f9fa756d03e4aa4ff34c3d8ddce0ea2c09959

Observation 7ff258c3-e82b-42e2-a58b-d3e9d85fff9c · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.675554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.675554Z digest=sha256:cf94b2019843fc3677a00a1a5d1db314b63952865c1a98f09458bffebd176798

Observation 0eafdb86-590c-4383-b8d7-23735c733308 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.470828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.470828Z digest=sha256:bdffb1de28c892b0c68348ae28526093b426decb36423d4b56a72df3a4a191ef

Observation 444f1190-f0e7-4f3b-b90d-ce3c4286dd53 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.287496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.287496Z digest=sha256:bb092d3a58539ee442cc09a27ad9ca185664b09910c36196ce80445ac1994de0

Observation d8fe9f11-f709-4ea8-b24b-29bb36acf662 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Measuring AI Ability to Complete Long Software Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.566544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.566544Z digest=sha256:1504d4356a5880d612bab73b3b041cfdcbb83b18e32f6405dca64a474b3ba730

Pith citing papers

No inbound Pith citation observations are available.