Pith. sign in

Paper Citation Record · LEDGER

TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2410.00752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00752 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:25.979207Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T03:45:21.868843Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 02e7e399-a2e0-4019-8f17-b1cbe94ed763 · inbound

MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms cites this paper.

MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.871439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T03:42:36.946141Z digest=sha256:05fd0c9f9652e3fdd7ecfbda9f9aee9b376fdf8dac4254a3663fc05ee7896133

Observation 8a0d8b4d-4b60-4060-812c-68bf083baa74 · inbound

CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification cites this paper.

CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:25.979207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:25.979207Z digest=sha256:f7894bfd35b0a14b324d654aa5ed7ee4c2cd8ad9a776200d3d9f3c3c50d7c62a

Observation 23b90f6f-747b-4e68-8d8f-7eaf5d0ef2fa · inbound

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution cites this paper.

SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:27:56.442733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T10:27:56.185943Z digest=sha256:234653e4c90ed76a47fdbbc5ece70c123069f7f26a239b17326e5895539f69d0

Observation 8b3bcd12-303f-47a2-9af5-9897f9af9525 · inbound

Why Do Multi-Agent LLM Systems Fail? cites this paper.

Why Do Multi-Agent LLM Systems Fail? TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:42:58.684087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:42:57.561468Z digest=sha256:8be17603bc778fc7d9e39b0be89c2de679115868cb4ee71efe4b07ee410073fe

Observation 24efad3f-0eec-4e62-8e7b-b58fb412704c · inbound

HardTests: Synthesizing High-Quality Test Cases for LLM Coding cites this paper.

HardTests: Synthesizing High-Quality Test Cases for LLM Coding TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:56.138659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:56.138659Z digest=sha256:403f22f487012f85c581845041ddfe83a8077b2bad8692d1417f1f43e0f0e4a5

Observation a611a773-c084-4928-a60e-ded7b64f1ecb · inbound

FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation cites this paper.

FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:25.396246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:25.396246Z digest=sha256:7d04085e36767053792d06b2193cfe269b5d7b46544acce0d9e74d78516b5bfe

Observation f921c16d-44f4-44eb-8cfe-1af72535ad51 · inbound

SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation cites this paper.

SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:37.827435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:37.827435Z digest=sha256:88a09b4eeca2370b10fc2bfc4a6864077ab1107433ad1ac30c39587988d55a49

Observation 049e1bf8-c0f5-4740-8a45-f9f4d428a18b · inbound

Benchmarking LLMs for Unit Test Generation from Real-World Functions cites this paper.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.159085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.159085Z digest=sha256:fdc0f96786f6d0531c81f295c4c123cbb2b9a5b2ea76c29e94d018f847585c48

Observation b6a57fc7-2497-4244-8c4b-3494b0fa165b · inbound

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts cites this paper.

ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:46:52.591939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:46:52.591939Z digest=sha256:1df30952c1412604ef2db1d831a55b2b01fc379edfc6138faddb783438611e41

Observation 5fc08513-02e2-4c76-bc92-3ea9fc658c4e · inbound

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios cites this paper.

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:28:24.498603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T20:24:40.939455Z digest=sha256:6f99b7cbb4f8f952a6f62808847637f72418a6974c4f7e5e0e230d04af0426fd

Observation a2ff6eb6-907b-48d0-ab31-42dc42c9de5c · inbound

Planning to Explore: Curiosity-Driven Planning for LLM Test Generation cites this paper.

Planning to Explore: Curiosity-Driven Planning for LLM Test Generation TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:55.781002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:50:29.242287Z digest=sha256:1d86cd080f59d991a31792a081f6a0143a6fff2330a3b548eedfb5a095d16fec