Pith. sign in

Paper Citation Record · LEDGER

Benchmarking LLMs for Unit Test Generation from Real-World Functions

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2508.00408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00408 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:15:37.403984Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T21:36:14.478007Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T21:38:18.794785Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 909ff480-c21a-4d51-8576-4053065cded6 · outbound

This paper cites An orchestrated survey of methodologies for automated software test case generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An orchestrated survey of methodologies for automated software test case generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.491272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.126101Z digest=sha256:f8f86a54c9cf3683149b994bfe0ad510a4e2ad67336e1910c0219b82ab8b44e5

Observation ea7abe96-a976-4ac4-9c89-dbba8875d68b · outbound

This paper cites A survey on model-based testing tools for test case generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions A survey on model-based testing tools for test case generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.307863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.133773Z digest=sha256:7223b984c1d8f2999da2ab4c53bb1ec1724d0c0e8269ee90890bea6ec206572c

Observation 800f75a8-58ec-41cb-b846-af53ddb8ef4b · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An empirical evaluation of using large language models for automated unit test generation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.191143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.141935Z digest=sha256:2032425367b6adb8b6adc0ffdfc1235b824d7cad5b3bbd7a1204d759abc15904

Observation 20765719-ae57-483c-b8bc-a4c241f1158b · outbound

This paper cites Measuring the Influence of Incorrect Code on Test Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Measuring the Influence of Incorrect Code on Test Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.146926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.146926Z digest=sha256:eef9360789d08007c3b23186c5b92dc0c748ce268e92bd16a86eee4d431bc343

Observation 29603517-e653-4063-81b1-b10fab08a9be · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.152164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.152164Z digest=sha256:2ddd67d21226ab3cdc2eb4293deba49d06a6b3c567462f4cb3c6758539eea475

Observation 049e1bf8-c0f5-4740-8a45-f9f4d428a18b · outbound

This paper cites TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.159085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.159085Z digest=sha256:e306ee40dfac278a678aeaa8f047db43d1ecec1452ba58c9937a9b199fbff306

Observation ae226fdf-6f5f-40f1-8101-baf4c422bc97 · outbound

This paper cites A survey on unit testing practices and problems,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions A survey on unit testing practices and problems,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:42.047530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.169201Z digest=sha256:3581d7b6bcd73cad74061ebb0c189a2dc480779604983b5b39f7be7065e469ce

Observation eb0478e7-befc-4ddf-8116-e9c4b62f79d2 · outbound

This paper cites DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model.

Benchmarking LLMs for Unit Test Generation from Real-World Functions DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:15:37.951397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.174272Z digest=sha256:ffd802fff49f6e9fd88e829b02f43866801d5c991e944ace1e626a981f74075e

Observation 0a4d2216-eb5c-4bda-9b3d-eba1c356a96d · outbound

This paper cites CodeT: Code Generation with Generated Tests.

Benchmarking LLMs for Unit Test Generation from Real-World Functions CodeT: Code Generation with Generated Tests

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.179054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.179054Z digest=sha256:54cea091c4ca18dd726dcd193eb9055fc5547986b3cc55011d830889d84edb6c

Observation 4d65a213-27c3-4f68-8908-30991d5de08c · outbound

This paper cites CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.184821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.184821Z digest=sha256:c85059807cfc03904e58b0ff8a606fa13da289b2aaeafe46f0c7f2326496e48d

Observation 9ae9a8b8-c5cc-405d-a5bb-098ae768dbdd · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.190831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.190831Z digest=sha256:70b6da60bc07899877b64b0e96f221a07c0498cc72a4b859db7f0af7bfc38fdb

Observation 4d6aaf3f-2819-4815-bc29-b6e02d53aab1 · outbound

This paper cites Mercury: A Code Efficiency Benchmark for Code Large Language Models.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Mercury: A Code Efficiency Benchmark for Code Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.196364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.196364Z digest=sha256:3888febf4851a9338e257356296c9818938567b64cc2b1a780f62cbdce6a2aca

Observation 9e438847-c64c-46cb-b037-17e03cdcd579 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Reflexion: Language agents with verbal reinforcement learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.926580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.201965Z digest=sha256:0132f7e77f38488a809d6a464bffc70991488b8a0b4f632bb53221698f0a4c5f

Observation a9d9dcd7-daaf-4541-90a9-e715e922b3fe · outbound

This paper cites Kernelgpt: Enhanced kernel fuzzing via large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Kernelgpt: Enhanced kernel fuzzing via large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.772241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.207221Z digest=sha256:00177d71604f036e899699a56478965d298ebcf25b211d33dea5ebfb18005d05

Observation 71555cd2-b041-4621-b1d4-4a882b6319a2 · outbound

This paper cites Universal fuzzing via large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Universal fuzzing via large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.645203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.213034Z digest=sha256:9ca5c5945862c18c1e5ae4194d710a02d683b1a4df0d56bd3d954af0059beed4

Observation 5694b4f1-d54c-491b-9b6a-c62ab0566d60 · outbound

This paper cites Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.532618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.217762Z digest=sha256:b349940e4e838f7ff3e0b1d4b8a06592adeec7fc5cb2a15d3f66480bb0b1293c

Observation 3e6756b4-c2fc-46d2-a198-6f38cec3352f · outbound

This paper cites Whitefox: White-box compiler fuzzing empowered by large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Whitefox: White-box compiler fuzzing empowered by large language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.395371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.222301Z digest=sha256:6af74e979d6756832728192b4ce5a4ff916e60e1725b06b6d4759d1c8b95179d

Observation af101ac6-f2ce-4913-b4ee-00732bf7ef31 · outbound

This paper cites Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.250652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.227948Z digest=sha256:b53b62e019e1cd9267049f9d989b6844ac666580d846fada8a47d5e7662e37e3

Observation b4f36b78-d7bc-454c-b966-ae36971a537d · outbound

This paper cites Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.233882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.233882Z digest=sha256:65939b2773eff3d15b14b787885b2b88f7643add50bece222d560253d72b7de4

Observation b2705c7b-8835-4789-9913-b08f624d00c1 · outbound

This paper cites Swt-bench: Testing and validating real-world bug-fixes with code agents,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Swt-bench: Testing and validating real-world bug-fixes with code agents,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:41.117904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.240076Z digest=sha256:18704ec7807b2f6aee48fb32a9141e3d7496e44bc145944794b2c9e428bfb1f3

Observation cf2f71a3-4925-48f5-82d3-fce2de10f34b · outbound

This paper cites TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models.

Benchmarking LLMs for Unit Test Generation from Real-World Functions TestBench: Evaluating Class-Level Test Case Generation Capability of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.245997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.245997Z digest=sha256:a03277c1f541c3f879771ceae1da2f31cb75456bbe0d1744a2f251b73b7ba835

Observation 83562be3-cda6-4a35-a212-1d2279ecdaa8 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Benchmarking LLMs for Unit Test Generation from Real-World Functions SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.253546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.253546Z digest=sha256:6c6290816deb1c7408bfd3d4849a2adee003694fe1cd2379b1dad71d72d3f0dd

Observation b5ec29d1-8d54-4483-b0bf-bfbcd156afee · outbound

This paper cites LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks.

Benchmarking LLMs for Unit Test Generation from Real-World Functions LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.260327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.260327Z digest=sha256:b265a133d60bd3c89cae3b46f5f4987438f9976f6dc6ec16c9b81e7babea286b

Observation ff5e4032-eae5-4b3f-b5ab-053d5a41b488 · outbound

This paper cites Large-scale, independent and comprehensive study of the power of llms for test case generation,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large-scale, independent and comprehensive study of the power of llms for test case generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.266586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.266586Z digest=sha256:451b66bd3dca81c2330917ca9014daf4bfc5e1789077f89cb96a87ae1e8c6b10

Observation 02536c0a-22c9-4bfa-bade-7b3fda8814aa · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Benchmarking LLMs for Unit Test Generation from Real-World Functions StarCoder 2 and The Stack v2: The Next Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.271977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.271977Z digest=sha256:01219b073e5e475d0034bb02b59f0ca0b959bbead292cafdab5fe00b76bf210c

Observation 41ed3b87-906b-47ac-ad8d-ac26dfdcda2c · outbound

This paper cites Software engineering (ed.),.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Software engineering (ed.),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.927642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.277072Z digest=sha256:30636e94d745f6b1cc5c7b3544f02caebab4818b840a01a9e2d091cea033833a

Observation 0f116071-e015-4802-95ed-7e87226c6279 · outbound

This paper cites an unresolved cited work.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T10:15:40.759349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.281642Z digest=sha256:d454ce08c647ce4b6554ca51f70409181747a38e39c23aa0bb5376255e7b0c14

Observation ff03e296-9140-4fed-a2d7-f911343c657d · outbound

This paper cites A survey of unit testing practices,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions A survey of unit testing practices,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.593196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.286525Z digest=sha256:dd1af2cf8d68153f38c3b1b57c6546e8d7aa02630e3fd342d06ffc26b0b6fad4

Observation dd1ddfd1-1f53-49bc-a13d-38bbb1428f03 · outbound

This paper cites Meszaros, xUnit test patterns: Refactoring test code.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Meszaros, xUnit test patterns: Refactoring test code

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.291794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.291794Z digest=sha256:58495c956ee516486d0101987019234ece92da59e4ee53c47df93166c2717c73

Observation 924cc713-f094-4d54-8340-4c419023b4d5 · outbound

This paper cites Automated unit test generation for evolving software,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Automated unit test generation for evolving software,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.417405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.301252Z digest=sha256:49c9d13d36c1d10b8bfb2c4403a67de00d4158c7304218291b1b50b9d070a701

Observation 5c3066a3-18ad-4b10-a7b0-67c9f92525d7 · outbound

This paper cites Symbolic execution and program testing,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Symbolic execution and program testing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.256742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.309549Z digest=sha256:1d581250611fe207e93360c38e12e836ab81fe93adffc28d105240154443dd7f

Observation b384488a-e4aa-44a9-a408-8a17902f543d · outbound

This paper cites The s2e platform: Design, implementation, and applications,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions The s2e platform: Design, implementation, and applications,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:40.097392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.314862Z digest=sha256:eb870ad5e7a7b9f2b02d1f992cc08b6c1819e51d142d0aefebfb3baabac0d353

Observation d1d81501-34b4-491b-b7f3-9ee0d69f9d12 · outbound

This paper cites Search-based software testing: Past, present and future,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Search-based software testing: Past, present and future,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.946637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.323001Z digest=sha256:bd0d44346da5af35c45aa98da1077d96e7cc95b16ce77d063104e37cc1877e9f

Observation 83da7827-a392-4379-9e6e-6a8fd88a1d87 · outbound

This paper cites An empirical study of the reliability of unix utilities,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An empirical study of the reliability of unix utilities,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.791995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.328908Z digest=sha256:3790783f9328bbf57fcdb030c778d2266e9d5628cf94c4cf4628fbdb5301ba5d

Observation e86156dc-c893-471e-be4e-40841930b032 · outbound

This paper cites Large language models for software engineering: A systematic literature review,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Large language models for software engineering: A systematic literature review,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.627099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.338955Z digest=sha256:3d0d568be45d7b140d25192325b60e857010c3b9d0a5223c933458e9edfd0869

Observation 93e459f2-955d-469c-9a37-1fe3996eea03 · outbound

This paper cites Mutation testing advances: an analysis and survey,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Mutation testing advances: an analysis and survey,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.446470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.344655Z digest=sha256:0b6a54a8612728b365421ac34820c5911f1104054fea458285136565ed087566

Observation 2c7014cb-cf71-48d7-a63b-a9737d135658 · outbound

This paper cites An analysis and survey of the development of mutation testing,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions An analysis and survey of the development of mutation testing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.314473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.352791Z digest=sha256:07a53a24d05b190e3b5ce221db925d721f78e10ca676a44c5d53d6600ebd196b

Observation c2fd3083-4b86-4351-bcac-e386a66b740e · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Benchmarking LLMs for Unit Test Generation from Real-World Functions BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.358395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.358395Z digest=sha256:f8186561dc30d0cc1fb73b8a6599b5917cb3ffa83cb8ec2295a54705de0c10a7

Observation fa05e55a-0e56-4278-a613-8fced4f459b2 · outbound

This paper cites Using large language models to generate junit tests: An empirical study,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Using large language models to generate junit tests: An empirical study,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.183460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.363660Z digest=sha256:92f3d7a905d15dc81b18488131568185fa70a0891334c4cd40e33078954b563c

Observation 6cbb6b53-f7a5-4458-88b7-b559f9465ca1 · outbound

This paper cites Using github copilot for test generation in python: An empirical study,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Using github copilot for test generation in python: An empirical study,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:39.012047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.369111Z digest=sha256:7e6627393fd58fabe1e89c8db4a8cb89b3dbdc2e4b9a60fbb68227c88d4fd996

Observation f0f93134-cd37-4036-be2d-f5ed052d9b88 · outbound

This paper cites Testspark: Intellij idea’s ultimate test generation companion,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Testspark: Intellij idea’s ultimate test generation companion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.873452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.373948Z digest=sha256:c44b863af6318ec54cbc04f3e2e64c9ee3d8249c6f8ae57dd9f00f1476e10f0e

Observation bed90727-2ed0-47ba-af57-ab730671c8fc · outbound

This paper cites Code Generation Tools (Almost) for Free? A Study of Few-Shot, Pre-Trained Language Models on Code.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Code Generation Tools (Almost) for Free? A Study of Few-Shot, Pre-Trained Language Models on Code

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.378469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.378469Z digest=sha256:796ea8d176b98e1f6c9c8e56e9968e1527b190367a3f5633716202f9839c175b

Observation a015a862-9b21-4eda-8a81-7bf7dcde4a6c · outbound

This paper cites Retrieval-based prompt selection for code-related few-shot learning,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Retrieval-based prompt selection for code-related few-shot learning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.684641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.383535Z digest=sha256:6bf42a3f55ed69f52c8503552c5104a9d82f175459a0e0eb9950ac81786d6137

Observation 85510848-0265-4bfc-be73-f7a94fdf1be4 · outbound

This paper cites Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.565708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.388650Z digest=sha256:3fb44403676cbd10c32f2aa2e65dd60e2928fcf151a767a8be615e475ad77134

Observation 2b5f2b33-ba1b-4734-b40c-8a316f39248c · outbound

This paper cites Testart: Improving llm-based unit test via co-evolution of automated generation and repair iteration,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Testart: Improving llm-based unit test via co-evolution of automated generation and repair iteration,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.374832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.393606Z digest=sha256:74645de489e1a5e0c40c31998d594824ee1a29633e7da6e2d28a3d293155f217

Observation 8e54c550-ad65-43a7-9ee1-ceffe40b417a · outbound

This paper cites ASTER: Natural and Multi-language Unit Test Generation with LLMs.

Benchmarking LLMs for Unit Test Generation from Real-World Functions ASTER: Natural and Multi-language Unit Test Generation with LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T10:15:37.398844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:15:37.398844Z digest=sha256:f34812eeea8c087ed96f215e4667f2282bb4027ae10634bf0542c4a81d36afc6

Observation 0e0e4403-ed56-48be-aa69-1e2fce5c4585 · outbound

This paper cites Effective test generation using pre-trained large language models and mutation testing,.

Benchmarking LLMs for Unit Test Generation from Real-World Functions Effective test generation using pre-trained large language models and mutation testing,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:15:38.187543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T10:15:37.403984Z digest=sha256:74bfdfd3e70dd9d7b78640eedbc7b6fc20c1a012f3df7d92e4bbe13b5c5ce71d

Pith citing papers

Observation 3c6eb1fb-54b4-4124-8224-b7e758e4d17a · inbound

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning cites this paper.

TestDecision: Sequential Test Suite Generation via Greedy Optimization and Reinforcement Learning Benchmarking LLMs for Unit Test Generation from Real-World Functions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:38:18.796047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:36:14.478007Z digest=sha256:37e931127e505b661ffcc060e0bbda2b2def22dfd43e76dfed1150b1510b4c2f