Pith. sign in

Paper Citation Record · LEDGER

Benchmarking LLMs for Unit Test Generation from Real-World Functions

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2508.00408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00408 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:15:37.403984Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T21:36:14.478007Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T21:38:18.794785Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 909ff480-c21a-4d51-8576-4053065cded6 · outbound

This paper cites An orchestrated survey of methodologies for automated software test case generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An orchestrated survey of methodologies for automated software test case generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.491272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.126101Z digest=sha256:7a22f47aff942f5d064d5675b22bc3e2e9b91b40d84e361aef82d85e5d34639a

Observation ea7abe96-a976-4ac4-9c89-dbba8875d68b · outbound

This paper cites A survey on model-based testing tools for test case generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions A survey on model-based testing tools for test case generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.307863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.133773Z digest=sha256:af178cf1cb94b8267d74b45b36f11a38b1ec887dfff6333b437ab59c9a7bcd93

Observation 800f75a8-58ec-41cb-b846-af53ddb8ef4b · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An empirical evaluation of using large language models for automated unit test generation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.191143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.141935Z digest=sha256:a550899aaab9155361771f3c6d91694c346569eb57bd5986bacd7d97dd5e55e9

Observation 20765719-ae57-483c-b8bc-a4c241f1158b · outbound

This paper cites Measuring the Influence of Incorrect Code on Test Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Measuring the Influence of Incorrect Code on Test Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.146926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.146926Z digest=sha256:13db123dda25ebc03efc7f3739cbb03558f9b0de00cc3389d6671bc6ff2c0811

Observation 29603517-e653-4063-81b1-b10fab08a9be · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.152164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.152164Z digest=sha256:93329f21e29c75d8d1ae87aa0dd4cf2ae424fc86f716cb418b15b27712916571

Observation 049e1bf8-c0f5-4740-8a45-f9f4d428a18b · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.159085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.159085Z digest=sha256:fdc0f96786f6d0531c81f295c4c123cbb2b9a5b2ea76c29e94d018f847585c48

Observation ae226fdf-6f5f-40f1-8101-baf4c422bc97 · outbound

This paper cites A survey on unit testing practices and problems,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions A survey on unit testing practices and problems,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.047530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.169201Z digest=sha256:45312dc487c5ec0f7ab97220caaa5ed816c9dfcea87a690a06b2b94efc7eb6aa

Observation eb0478e7-befc-4ddf-8116-e9c4b62f79d2 · outbound

This paper cites DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model.

Benchmarking LLMs for Unit Test Generation from Real-World Functions DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:15:37.951397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.174272Z digest=sha256:a38c1d12fb6442c20540d1eb049c62faae93329c91ae0f94ad595baebedcb645

Observation 0a4d2216-eb5c-4bda-9b3d-eba1c356a96d · outbound

This paper cites CodeT: Code Generation with Generated Tests.

Benchmarking LLMs for Unit Test Generation from Real-World Functions CodeT: Code Generation with Generated Tests

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.179054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.179054Z digest=sha256:54cea091c4ca18dd726dcd193eb9055fc5547986b3cc55011d830889d84edb6c

Observation 4d65a213-27c3-4f68-8908-30991d5de08c · outbound

This paper cites CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.184821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.184821Z digest=sha256:943c2e7db3d40647f06b58a08183772a4530db613dde594e8d598d3901ac5aef

Observation 9ae9a8b8-c5cc-405d-a5bb-098ae768dbdd · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.190831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.190831Z digest=sha256:a8bf8fba8f53969f9e11742b488feab4e367726bf5765354bd84c52eb6937257

Observation 4d6aaf3f-2819-4815-bc29-b6e02d53aab1 · outbound

This paper cites Mercury: A Code Efficiency Benchmark for Code Large Language Models.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Mercury: A Code Efficiency Benchmark for Code Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.196364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.196364Z digest=sha256:029be43fe3968d95f89937199d00fb5eca4aeea2be6d2db26ed14625dc579406

Observation 9e438847-c64c-46cb-b037-17e03cdcd579 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Reflexion: Language agents with verbal reinforcement learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.926580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.201965Z digest=sha256:54c8903e4bdb1628cae3ef31a757a6c4ad6f6c74ab13807c654f4bdd5426f8eb

Observation a9d9dcd7-daaf-4541-90a9-e715e922b3fe · outbound

This paper cites Kernelgpt: Enhanced kernel fuzzing via large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Kernelgpt: Enhanced kernel fuzzing via large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.772241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.207221Z digest=sha256:cb0b6c9887c4e683b68ed050e796e21bd296c67d1f85edf3f689df7a9dcb358a

Observation 71555cd2-b041-4621-b1d4-4a882b6319a2 · outbound

This paper cites Universal fuzzing via large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Universal fuzzing via large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.645203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.213034Z digest=sha256:fdf762ee6f2d3f45e7ca155eb65e0aa2a097ee675cc691e87f8252dc245f8c16

Observation 5694b4f1-d54c-491b-9b6a-c62ab0566d60 · outbound

This paper cites Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.532618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.217762Z digest=sha256:8b4f990cfada1327a9187d86677fcb95bd1cc9cac2a0ad7e5acc0ca00f0e2082

Observation 3e6756b4-c2fc-46d2-a198-6f38cec3352f · outbound

This paper cites Whitefox: White-box compiler fuzzing empowered by large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Whitefox: White-box compiler fuzzing empowered by large language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.395371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.222301Z digest=sha256:8131f4e53f167dd456fbe35714c7910f18a6290dd497d88a35dd2344ea030af3

Observation af101ac6-f2ce-4913-b4ee-00732bf7ef31 · outbound

This paper cites Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.250652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.227948Z digest=sha256:d565829b3f60eb65db68733997a7f4ef156a2cdac393b9bdba6131c448f930cd

Observation b4f36b78-d7bc-454c-b966-ae36971a537d · outbound

This paper cites Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.233882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.233882Z digest=sha256:3346f6f4e04533d257e29fcfcc45a0cf2b52b0eff6d0a469eaf2d69e89c28b34

Observation b2705c7b-8835-4789-9913-b08f624d00c1 · outbound

This paper cites Swt-bench: Testing and validating real-world bug-fixes with code agents,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Swt-bench: Testing and validating real-world bug-fixes with code agents,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.117904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.240076Z digest=sha256:725d7193a4fa8582ccdb0b60ded5aefeb944f33a7939b26bdce29042fa230616

Observation cf2f71a3-4925-48f5-82d3-fce2de10f34b · outbound

This paper cites TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.245997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.245997Z digest=sha256:9711f54c827b617bc8fd91a23e81280b2d6bc1054b433f5875f0cc7ad2e8df96

Observation 83562be3-cda6-4a35-a212-1d2279ecdaa8 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Benchmarking LLMs for Unit Test Generation from Real-World Functions SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.253546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.253546Z digest=sha256:1ab84292caab1883ef9d120f0ff6ec78a67348f0daac1aabf1ca4b2dccfae0ee

Observation b5ec29d1-8d54-4483-b0bf-bfbcd156afee · outbound

This paper cites LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks.

Benchmarking LLMs for Unit Test Generation from Real-World Functions LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.260327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.260327Z digest=sha256:c2287d1536a5750dcc9f5e91c91768906c687ef601c526643f6d1fdd02b9b0db

Observation ff5e4032-eae5-4b3f-b5ab-053d5a41b488 · outbound

This paper cites Large-scale, independent and comprehensive study of the power of llms for test case generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large-scale, independent and comprehensive study of the power of llms for test case generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.266586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.266586Z digest=sha256:451b66bd3dca81c2330917ca9014daf4bfc5e1789077f89cb96a87ae1e8c6b10

Observation 02536c0a-22c9-4bfa-bade-7b3fda8814aa · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions StarCoder 2 and The Stack v2: The Next Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.271977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.271977Z digest=sha256:01219b073e5e475d0034bb02b59f0ca0b959bbead292cafdab5fe00b76bf210c

Observation 41ed3b87-906b-47ac-ad8d-ac26dfdcda2c · outbound

This paper cites Software engineering (ed.),.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Software engineering (ed.),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.927642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.277072Z digest=sha256:d2e2c3edbaf3eed3aa0cf6e8b0858f37ce64cd36f8bd8980f41a1ae8043955d8

Observation 0f116071-e015-4802-95ed-7e87226c6279 · outbound

This paper cites an unresolved cited work.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T10:15:40.759349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.281642Z digest=sha256:6bf11b8bcda5a63155317afdc78c6705834985750fafb1f16b7e962e5c43646d

Observation ff03e296-9140-4fed-a2d7-f911343c657d · outbound

This paper cites A survey of unit testing practices,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions A survey of unit testing practices,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.593196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.286525Z digest=sha256:652966ef2ee6a29368a3fa75c1de44a3b99626aa2158cf358abb4ffbecfe201b

Observation dd1ddfd1-1f53-49bc-a13d-38bbb1428f03 · outbound

This paper cites Meszaros, xUnit test patterns: Refactoring test code.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Meszaros, xUnit test patterns: Refactoring test code

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.291794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.291794Z digest=sha256:58495c956ee516486d0101987019234ece92da59e4ee53c47df93166c2717c73

Observation 924cc713-f094-4d54-8340-4c419023b4d5 · outbound

This paper cites Automated unit test generation for evolving software,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Automated unit test generation for evolving software,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.417405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.301252Z digest=sha256:084cb55a5941b27bf13960cf04b1e6d1d8a5fe13c9ae6ce94cd04487af2a29d3

Observation 5c3066a3-18ad-4b10-a7b0-67c9f92525d7 · outbound

This paper cites Symbolic execution and program testing,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Symbolic execution and program testing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.256742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.309549Z digest=sha256:cc421be76a53d40851017ad4f760c19cbbeb4ff26d0ff6ae9949c26bdb0c10e0

Observation b384488a-e4aa-44a9-a408-8a17902f543d · outbound

This paper cites The s2e platform: Design, implementation, and applications,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions The s2e platform: Design, implementation, and applications,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.097392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.314862Z digest=sha256:7ad2747c3e275c48c06c348cd5c668ad0f9dc295042e76ee853f7cbac4679c82

Observation d1d81501-34b4-491b-b7f3-9ee0d69f9d12 · outbound

This paper cites Search-based software testing: Past, present and future,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Search-based software testing: Past, present and future,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.946637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.323001Z digest=sha256:dda4ea9ed914d48e4b3fac54693a5eb9e3fcaa6d8029f1c3c1d88bdcd7780b07

Observation 83da7827-a392-4379-9e6e-6a8fd88a1d87 · outbound

This paper cites An empirical study of the reliability of unix utilities,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An empirical study of the reliability of unix utilities,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.791995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.328908Z digest=sha256:b44092400adadccc6809ea7093cf4342fa8e227ecee643f8bdec92c9beda405d

Observation e86156dc-c893-471e-be4e-40841930b032 · outbound

This paper cites Large language models for software engineering: A systematic literature review,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large language models for software engineering: A systematic literature review,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.627099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.338955Z digest=sha256:1463e6f2ab2aff3e51463b2d592aa21008c1b4bcc460e39ae8a56c5e435d60c9

Observation 93e459f2-955d-469c-9a37-1fe3996eea03 · outbound

This paper cites Mutation testing advances: an analysis and survey,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Mutation testing advances: an analysis and survey,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.446470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.344655Z digest=sha256:19d410539eabc5780f2eb1e451de99053b94a01d78b73a5c46986a18c20e40c8

Observation 2c7014cb-cf71-48d7-a63b-a9737d135658 · outbound

This paper cites An analysis and survey of the development of mutation testing,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An analysis and survey of the development of mutation testing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.314473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.352791Z digest=sha256:29af57f4a823c77499c4b3e35330e50ccb0c83eb99582e4d22666fdac550347f

Observation c2fd3083-4b86-4351-bcac-e386a66b740e · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Benchmarking LLMs for Unit Test Generation from Real-World Functions BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.358395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.358395Z digest=sha256:447da2b07b7fb32540095ecf366dfc3832c475e9f981487d8c55b9aa963511f7

Observation fa05e55a-0e56-4278-a613-8fced4f459b2 · outbound

This paper cites Using large language models to generate junit tests: An empirical study,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Using large language models to generate junit tests: An empirical study,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.183460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.363660Z digest=sha256:2ee601652d1950d69a1120bbaf02450288f5864fba627151c2dd543b878a92b6

Observation 6cbb6b53-f7a5-4458-88b7-b559f9465ca1 · outbound

This paper cites Using github copilot for test generation in python: An empirical study,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Using github copilot for test generation in python: An empirical study,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.012047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.369111Z digest=sha256:00b6b49f8e5c8940c9b259dfd3fa3aeb7a3901a82956090c1755c94fb98444ee

Observation f0f93134-cd37-4036-be2d-f5ed052d9b88 · outbound

This paper cites Testspark: Intellij idea’s ultimate test generation companion,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Testspark: Intellij idea’s ultimate test generation companion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.873452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.373948Z digest=sha256:16ef5a60bf0f8d38033db8a20359b44604d4a5bebaefa98169d0c284a5aeb588

Observation bed90727-2ed0-47ba-af57-ab730671c8fc · outbound

This paper cites Code Generation Tools (Almost) for Free? A Study of Few-Shot, Pre-Trained Language Models on Code.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Code Generation Tools (Almost) for Free? A Study of Few-Shot, Pre-Trained Language Models on Code

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.378469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.378469Z digest=sha256:71c3c738e75bc8e4ba5d5be3ef4777c60db7e95199c1e0d1511f471ae352930c

Observation a015a862-9b21-4eda-8a81-7bf7dcde4a6c · outbound

This paper cites Retrieval-based prompt selection for code-related few-shot learning,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Retrieval-based prompt selection for code-related few-shot learning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.684641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.383535Z digest=sha256:042137349eabc4fe9c7e67e8ff50fca6f3fff9a8d386a335ccd948992a51a3e6

Observation 85510848-0265-4bfc-be73-f7a94fdf1be4 · outbound

This paper cites Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.565708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.388650Z digest=sha256:0e4d52d250b6f8ffeac6152b833765ed43d38be9f2e60a5ee147f8c073f59a39

Observation 2b5f2b33-ba1b-4734-b40c-8a316f39248c · outbound

This paper cites Testart: Improving llm-based unit test via co-evolution of automated generation and repair iteration,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Testart: Improving llm-based unit test via co-evolution of automated generation and repair iteration,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.374832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.393606Z digest=sha256:8336b35bdfbb3292af64b72818beb253655de6973f942b3a7323687118e7e251

Observation 8e54c550-ad65-43a7-9ee1-ceffe40b417a · outbound

This paper cites ASTER: Natural and Multi-language Unit Test Generation with LLMs.

Benchmarking LLMs for Unit Test Generation from Real-World Functions ASTER: Natural and Multi-language Unit Test Generation with LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.398844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.398844Z digest=sha256:d18cae2b4ac618d3125b1f56925ea0bedf9611f775ca05196de2d7b4b33d44c4

Observation 0e0e4403-ed56-48be-aa69-1e2fce5c4585 · outbound

This paper cites Effective test generation using pre-trained large language models and mutation testing,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Effective test generation using pre-trained large language models and mutation testing,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.187543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T10:15:37.403984Z digest=sha256:7023ece50797069df26d03bb908e8bc52580915698af7c6b91738109b5e22dea

Pith citing papers

Observation 3c6eb1fb-54b4-4124-8224-b7e758e4d17a · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning Benchmarking LLMs for Unit Test Generation from Real-World Functions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:38:18.796047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:ae826f7a987d063de351107298f3847ee7288c95e3f3f3bf947cb84466960441