Pith. sign in

Paper Citation Record · LEDGER

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation

As of 14 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 2 inbound Pith citation observations for arXiv:2502.05714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05714 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:20:22.558735Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T05:11:01.323314Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2bdf5a32-745e-4ed2-bb03-33e0cf363d7b · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation Code Llama: Open Foundation Models for Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.514245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.514245Z digest=sha256:d579845b36cc1b1d5853d6286ecf3f8b0457e8f011797f83ac2252c4694dc69b

Observation e2ed5352-a667-4e0d-bdb7-bada423db203 · outbound

This paper cites LEGO-Prover: Neural Theorem Proving with Growing Libraries.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation LEGO-Prover: Neural Theorem Proving with Growing Libraries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.519772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.519772Z digest=sha256:2aaf18c757f67402ce1be6d074c0875508c9a8b38ca5812d43abe2d587a28e54

Observation 3900a402-1c6f-4551-854a-ad7422fb89b0 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.529338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.529338Z digest=sha256:0b7a236642415f6e7f0cecc86a542822cb8f87189078aa4aa940124632f39d7f

Observation e1c46de1-6506-47fc-9c4d-d88503dc6464 · outbound

This paper cites The Llama 3 Herd of Models.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.534347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.534347Z digest=sha256:b1fe620366138e79e39e24d0511eab0aed23d2a0b3e826c0af56620c13235d7b

Observation b1c27942-5e92-4b3c-a41d-e90b90ae2043 · outbound

This paper cites DafnyBench: A Benchmark for Formal Software Verification.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation DafnyBench: A Benchmark for Formal Software Verification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.539045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.539045Z digest=sha256:9e00d6b7c9dd2c562081dfe732e26ad2397a4b5b12f3f50ebdb508cc8c1657dc

Observation bbccd834-a36b-4fff-a4a6-72acc026921e · outbound

This paper cites GPT-4 Technical Report.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.543642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.543642Z digest=sha256:fa036af159d4a51d6992b1b2f1c7af420f5e0e0fa1fe963686c312a913ce927b

Observation 3c559fc6-56aa-4a49-a686-c81480d1d42f · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.548525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.548525Z digest=sha256:cdc795fee17b0aada5750e385c1274625fe48791fc33976981136f3a7eae8832

Observation 776b766b-eefc-400e-b66f-571f0ae33cad · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.553511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.553511Z digest=sha256:cdb4741281e76c13f33eb153105d9ac4006459b30f3b0ed6dbe90c9add9a2d4d

Observation 195d5823-d506-47d3-add2-a512adb2cbfe · outbound

This paper cites AutoVerus: Automated Proof Generation for Rust Code.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation AutoVerus: Automated Proof Generation for Rust Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.558735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.558735Z digest=sha256:d501b640b44221d3d71f2a9d6ab19435955fa4bcd6dc5f158e007b01ab6b0fb7

Observation 09b4de5c-6218-427a-aa9a-1f067e8ee89c · outbound

This paper cites Learning to Prove Theorems via Interacting with Proof Assistants.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation Learning to Prove Theorems via Interacting with Proof Assistants

Reference 1891

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.492751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.492751Z digest=sha256:58c10ea89901ba1edad1cfa95d2384a483ae8f2ac11b0fe72aa78995bd5bf52b

Observation dde48156-2f44-4ff4-ab47-205fc87ac575 · outbound

This paper cites Mining the Archive of Formal Proofs.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation Mining the Archive of Formal Proofs

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:20:22.815459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T18:20:22.486469Z digest=sha256:099b915b19592e6560c3e81f066c90cf1c09f581682baae1fcdb5670a33764d7

Observation d2d9c1e6-329e-4ef9-b217-d6ec2760dc85 · outbound

This paper cites DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.498151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.498151Z digest=sha256:aaa348af5afb454fdb61873e2bede9c8ee99049f31493726280702e81c5040ed

Observation 5f673435-009d-4a72-bf16-afe52d5b67e0 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.503535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.503535Z digest=sha256:1efa69f2db40908de3d715c8db8eaa22adb3cd674590ac8b23c175f9b732bf5c

Observation b9e72d7c-0d3d-4756-8328-6f89ce593838 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T18:20:22.508815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:20:22.508815Z digest=sha256:e77438ed4b571e0db65e1b7d094ca9ca58daee6b15804b9f2a84bcd4b6f9e4c5

Observation b3793567-4abc-487a-bbf6-9df12cf02119 · outbound

This paper cites google/discover/blog/ai-solves-imo-problems-at- silver-medal-level/ (visited on 10/30/2024).

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation google/discover/blog/ai-solves-imo-problems-at- silver-medal-level/ (visited on 10/30/2024)

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:20:22.799322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T18:20:22.524655Z digest=sha256:b00a78612163f883b57a906ad737af2658a196e1f7d17e725d2b3aecadc52d68

Pith citing papers

Observation 8b5825af-ab1a-4734-99f0-1f068be75cee · inbound

FVSpec: Real-World Property-Based Tests as Lean Challenges cites this paper.

FVSpec: Real-World Property-Based Tests as Lean Challenges Proving the Coding Interview: A Benchmark for Formally Verified Code Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.162743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T17:05:13.012431Z digest=sha256:f273a631b0ef50cb21764accbfcb453221d06dfcc8ba121a9b59f88103eaa6a1

Observation a308fcb9-490a-4986-bae3-307e7ee548f5 · inbound

Vero: Can AI Agents Build Formally Verified Software Repositories? cites this paper.

Vero: Can AI Agents Build Formally Verified Software Repositories? Proving the Coding Interview: A Benchmark for Formally Verified Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T05:11:01.323314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T05:11:01.323314Z digest=sha256:8163b5f3678192c9e67f6f6d4179faaa86ed3ff1bef5ab8d048dcae531486e83