Pith. sign in

Paper Citation Record · LEDGER

Comparative Evaluation of Large Language Models for Test-Skeleton Generation

As of 8 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.04644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04644 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:59:44.334749Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a56c354b-4765-4186-9b04-75a0ca420820 · outbound

This paper cites • DeepSeek-Chat: A domain-specific model optimized for developer tasks, with enhanced per formance on programming-related queries and structural code outputs.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation • DeepSeek-Chat: A domain-specific model optimized for developer tasks, with enhanced per formance on programming-related queries and structural code outputs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.977040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.344770Z digest=sha256:b54a2563e72d8e1655dda2a354bcfe2377a12654244a381ac76a2d512fc1ddb4

Observation 1f18e828-bd31-46e6-8bd6-5b05e1cf722d · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.969241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.410462Z digest=sha256:28ec516204b0bfb969edff0493edf17b38c98fc0cab79af1b66a50d45097bdf3

Observation 48b3af5c-2e3f-45bc-8270-30cbed800f3d · outbound

This paper cites This metric captures how completely the model covered the methods of the class.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation This metric captures how completely the model covered the methods of the class

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.961628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.507627Z digest=sha256:8e61e2d5472d03832a24834f78f8e2dfed6d0bc2a9e82970a3023cd4cd46aa27

Observation 486d227c-6b63-4476-8e9e-092fa731a294 · outbound

This paper cites The reviewer assessed each generated skeleton using six dimensions, scoring on a 1–5 scale:.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation The reviewer assessed each generated skeleton using six dimensions, scoring on a 1–5 scale:

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.954259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.630098Z digest=sha256:42faf7aa7599ec76d45708a2e535e9de1c24ba7442bae6d58f5f6a838b7fa5c0

Observation 2aa9f794-00e1-49da-bd61-c0556f896be5 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.946549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.722768Z digest=sha256:cde8d0b94ee75e9972f834ca1cb2403089a086b5586abeccad4e66c9bffbc0ea

Observation 23c1419b-ba06-44f3-89b2-18278065f3e9 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.938308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.807785Z digest=sha256:14594c5a86c83102a477a3e0f3e0c7a1e738c0505fb9026d75f7d0079c549ac3

Observation be35ffbd-2315-4430-b548-def0f3e66a13 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.931073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.881018Z digest=sha256:c2d36ab457c331e7bfecb8447153a5877772d3f234766ed3476faac3e37fefbf

Observation 54b6d876-ec38-424e-a6c7-00f8b0975f9f · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.923339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:41.972822Z digest=sha256:793fcbbec0f4d472505613a578fd44a07e7e0e349660045b7a02c03c2906c86f

Observation 7968bccd-2a55-494f-a818-0eeb245f31bb · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.915965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.063694Z digest=sha256:15d29ed69160f9e8650ee8cb5e488f4a1f166b88cc05645ecfd04403f16721aa

Observation 3cdd46fe-230b-403c-81ec-48642f7b9676 · outbound

This paper cites The expert review focused on identifying semantic misinterpretations, structural flaws, and usability within the skeletons generated by the models.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation The expert review focused on identifying semantic misinterpretations, structural flaws, and usability within the skeletons generated by the models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.908762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.131373Z digest=sha256:765dd3bb5f15fbc09700e0c708d620c2aa70f9064ef09ab44cacacefa420dccc

Observation b85c90f0-a531-4c50-8718-a442114582a2 · outbound

This paper cites This included appropriate usage of RSpec describe blocks and clear distinction between instance and class methods.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation This included appropriate usage of RSpec describe blocks and clear distinction between instance and class methods

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.900460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.194904Z digest=sha256:091ac5562f734e749d4c5a9ea4be51c5fd51109bcce105f6efbda440a135abe5

Observation 0f6af92f-122c-406a-bab3-00aeff581380 · outbound

This paper cites Its test skeleton was praised for its well -organized structure and strong alignment with RSpec idioms.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Its test skeleton was praised for its well -organized structure and strong alignment with RSpec idioms

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.891842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.268192Z digest=sha256:889e028902c3e7509b72fa762cd4d523b711c604205dbb0665d16a31cc1c1eae

Observation 98978fa1-f833-40e1-a54b-a47155aff6f8 · outbound

This paper cites Llama’s was the cleanest and easiest to maintain.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Llama’s was the cleanest and easiest to maintain

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.884101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.376257Z digest=sha256:6eab6aea36df417aa507ce7907ae00caf6dac59d784a8d1d00f34c9dc3cb9eda

Observation 3f890ebc-a1d5-4762-b228-9b6fe29c5b27 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.876794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.510525Z digest=sha256:21763240e21191e7f5db7b1a944e1e834e3f391dfe58e55eb15979e71f7332b2

Observation e4329897-e4db-4b47-9278-f75e5cae2c60 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.869064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.605560Z digest=sha256:f68a332e11783d5753b27787fc0918bc16ab30f2616da9a1af1b0a6fc39d2963

Observation 29b6aaa1-5873-40a0-91d1-64f80750467b · outbound

This paper cites Models that produced clean, readable code, like Llama4, were deemed more suitable for collaborative workflows.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Models that produced clean, readable code, like Llama4, were deemed more suitable for collaborative workflows

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.860914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.740858Z digest=sha256:0b2d9a14c23890dd9422054c8186d77a33df2804030d17925bf91548658caba7

Observation e0e3f8a9-0755-4317-9ddf-74f8d609fb38 · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.852855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.846810Z digest=sha256:9216a04cfdb5b5be0831d5e698b5a642e106e1a8172249e676f7671ad67495b9

Observation 7a433ddd-0f37-47ab-bfae-26a6d7f05a1e · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.845038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:42.940550Z digest=sha256:891742e5b48e9888635667c7bcd74adac5d6b94923651912d4e7223a99df8c99

Observation 87812ec4-1834-44d7-8c1e-46d95612554a · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.837008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:43.062723Z digest=sha256:9a19824e0c12c9dd577621a3a5ddb7d9c34a494e712a2a12a107d6c5154ddde3

Observation 862e439c-4c30-465b-8671-195159d70a0c · outbound

This paper cites While some models generated verbose skeletons with full coverage, these outputs were often dense, unstructured, or syntactically flawed.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation While some models generated verbose skeletons with full coverage, these outputs were often dense, unstructured, or syntactically flawed

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.793706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:43.203333Z digest=sha256:5e09104495b710f33e0fe1dbd63fa05eb228b1ed4eb9a491847ce9ff0d828d40

Observation 142814e7-915c-4348-9883-7f7ede077fad · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T05:59:45.673151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:43.295770Z digest=sha256:e88a54f33bbc457fb41f54db73352234962d411b2f817bd563c5bfb090574b36

Observation 88317820-8781-4781-877a-8906c4741ca8 · outbound

This paper cites Automation of Test Skeletons Within Test-Driven Development Projects,.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Automation of Test Skeletons Within Test-Driven Development Projects,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.381587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.381587Z digest=sha256:a5da078fe0d8c31b54fcae75b65fcdcb54707d455e2e2077cbe93073a13f6e8d

Observation c1ecfe73-34cb-4784-bcef-bebe459a42ed · outbound

This paper cites Software Testing with Large Language Models: Survey, Landscape, and Vision.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Software Testing with Large Language Models: Survey, Landscape, and Vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.507568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.507568Z digest=sha256:7a508b81b3d7ded572daa0958172d440b49897d81986217ca489651ea94db763

Observation c57c1dd8-5430-40dc-aee1-fa33e674ced4 · outbound

This paper cites Large Language Models for Software Engineering: A Systematic Literature Review.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Large Language Models for Software Engineering: A Systematic Literature Review

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.512605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:43.653363Z digest=sha256:e2a73f252cc426c1ff4e451c84796ca705f685513a93db774e64b6a22be9b664

Observation 6baeb417-25f0-4f31-bc11-249234aed318 · outbound

This paper cites Intelligent Software Testing: Harnessing Machine Learning to Automate Test Case Generation and Defect Prediction.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Intelligent Software Testing: Harnessing Machine Learning to Automate Test Case Generation and Defect Prediction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.371492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:43.741012Z digest=sha256:60d6c15659e89a91c78d40a0d2dd774e0ea89ffb55515651b80fbe048440c910

Observation f1ef20a3-8da8-4d4d-8f12-4e68ed204cb4 · outbound

This paper cites TESTEVAL: Benchmarking Large Language Models for Test Case Generation.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation TESTEVAL: Benchmarking Large Language Models for Test Case Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:43.828934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:43.828934Z digest=sha256:c80b78d09da47925cfd01c43ed7d846ce1df70a9c899865662e92ae24d548b03

Observation 9c1dff1c-6bb4-492c-94a6-4b25e24507f4 · outbound

This paper cites Evaluating Large Language Models for Software Testing. Science of Computer Programming.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Evaluating Large Language Models for Software Testing. Science of Computer Programming

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.264082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:43.963435Z digest=sha256:32baae033f91f0299c071c65713a958f27275722f9d8947b347a750e675f43e5

Observation b804f479-7dab-4b25-a69b-f7f736243ff1 · outbound

This paper cites EvoSuite: Automatic Test Suite Generation for Object- Oriented Software.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation EvoSuite: Automatic Test Suite Generation for Object- Oriented Software

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T05:59:45.091153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:44.085058Z digest=sha256:cdab1eecdf22c1bbddee8e9fe562d331e7e4cf989ffde4872f2b81585bf15e27

Observation 158ed5fb-e8d8-41a8-ac15-c6a5a0a0941b · outbound

This paper cites CUTE: A Concolic Unit Testing Engine for C.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation CUTE: A Concolic Unit Testing Engine for C

Reference 30

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T05:59:44.786201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:44.229379Z digest=sha256:356b529edb57147ee14226f8aa472e158e42cc4c51d0480a6c576c511e43176f

Observation 780e796e-02ff-4c64-8b1d-df185da2c10d · outbound

This paper cites Improving Automated Test Case Generation with Learning-to-Rank Techniques.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Improving Automated Test Case Generation with Learning-to-Rank Techniques

Reference 31

Resolution
verified exact
doi, observed 2026-08-05T05:59:44.462803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:44.294754Z digest=sha256:a5721d9b3a42b38aa3c4a56c203b31f494786f450a06d19a800839f10ebc95d6

Observation 57918f8e-d878-48fc-9b2d-54d82ff63d7b · outbound

This paper cites Intrinsic Nonlinear Hall Detection of the N\'eel Vector for Two-Dimensional Antiferromagnetic Spintronics.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Intrinsic Nonlinear Hall Detection of the N\'eel Vector for Two-Dimensional Antiferromagnetic Spintronics

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T05:59:44.620488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T05:59:44.334749Z digest=sha256:6be04b6fb861af90a4f46facc5d20ae504ebc5a6064a9cdbe0eb5e3da3de9a00

Observation eec26237-07d0-495b-a655-4b8974e6c9ba · outbound

This paper cites an unresolved cited work.

Comparative Evaluation of Large Language Models for Test-Skeleton Generation Unresolved cited work

Reference 419

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:44.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:44.163389Z digest=sha256:a4e6d3f6f00a0fcede06b1f0c7b36806541863fdb06a051910f7ad83a1dad897

Pith citing papers

No inbound Pith citation observations are available.