Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

As of 18 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 3 inbound Pith citation observations for arXiv:2506.07594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07594 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:35:41.214905Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T16:24:25.357338Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 23247134-ec7a-462b-b92b-b0da15036b50 · outbound

This paper cites On the relation of test smells to software code quality,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study On the relation of test smells to software code quality,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.756482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.943703Z digest=sha256:810a40406cc148b36606f534ab4a0bc4f0b48cb4348c3f55c911045225473333

Observation 8a415b7c-ea53-4703-807f-b2a336d20b4f · outbound

This paper cites On the diffusion of test smells in automatically generated test code: An empirical study,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study On the diffusion of test smells in automatically generated test code: An empirical study,

Reference 2

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:35:42.283784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.949345Z digest=sha256:6a2b68e8bdcfd70ad9ea06133d704ea2a3bab8ab724f309ccc735d7ce249863a

Observation db152661-eaa6-4851-905d-3fa62b9f4660 · outbound

This paper cites When and why your code starts to smell bad,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study When and why your code starts to smell bad,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.741730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.954889Z digest=sha256:2b2cf684b34e42e226a341728444bf42fa8a043efde424ba6f28629b2440618c

Observation 688e61d6-4cdb-49a5-a2de-f687eabf87dc · outbound

This paper cites An empirical analysis of the distribution of unit test smells and their impact on software maintenance,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical analysis of the distribution of unit test smells and their impact on software maintenance,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.726187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.960193Z digest=sha256:fff6feb0f69ded25f5b1e26511894f688c0ea5569045db2804fac480c4c8eeab

Observation 568dc405-c62c-401c-b339-7d00346a4871 · outbound

This paper cites Just-in-time test smell detection and refactoring: The darts project,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Just-in-time test smell detection and refactoring: The darts project,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.966381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.966381Z digest=sha256:031a59a9551d87628bf8b0e0e23e9b040e374bb66a92606c279a8b3bb81f1aae

Observation 1d6ec493-2c87-4666-9895-45d04d52dc64 · outbound

This paper cites GPT-4 Technical Report.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.971440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.971440Z digest=sha256:35a85afcaa46ff5aaadaf2b364b275bc39370d0a6d490c4d000ea9437c839e28

Observation 748b762e-2e35-4b29-a32f-bc2578dbe139 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.977343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.977343Z digest=sha256:d7dc5d7579f3652e28456ea1a9647fef60ba425e17905798c709298cfbaee8d4

Observation aea9bf46-fe6a-4274-9c69-d0eb6c0509fe · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:40.982398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:40.982398Z digest=sha256:045b8c154620388fb0c19178d3ff22731300a4b04e5fc2af8b8e72c98d4204f1

Observation 08a1ec8d-69e3-4d06-b149-1248e6a9c634 · outbound

This paper cites Codebert: A pre-trained model for programming and natural languages,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Codebert: A pre-trained model for programming and natural languages,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.710348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.987859Z digest=sha256:c8f23bf2cf737634b56ddd5a864f34d85da853666ec900cbc6224a232709a81c

Observation 37b2e45a-2e62-4589-a29e-30195993cc57 · outbound

This paper cites Codexglue: A machine learning benchmark dataset for code understanding and generation,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Codexglue: A machine learning benchmark dataset for code understanding and generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.692702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.992839Z digest=sha256:7a54acfeb5bb5bbe0f6754e48c0f677540679e16661f3636ca7162f19bdf7dd0

Observation 84c2c9f9-5d92-4b0f-858b-61b0e668b247 · outbound

This paper cites Top programming languages - the state of the octoverse 2022,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Top programming languages - the state of the octoverse 2022,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.677548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:40.998380Z digest=sha256:046e489e2f923a82089607f7ca94609cc452895e9742fcc9121ab03e3c08714c

Observation fe91d1d0-56e7-4e29-8824-7615b7467435 · outbound

This paper cites Utilization of pre-trained language model for adapter-based knowledge transfer in software engineering,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Utilization of pre-trained language model for adapter-based knowledge transfer in software engineering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.660795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.003516Z digest=sha256:cd25a6f9fb00f892e6381da7af95421893eac8f57708107560d0808857805e5c

Observation 11ba22fa-d781-4eb2-b61f-2334d7482892 · outbound

This paper cites To- wards efficient fine-tuning of pre-trained code models: An experimental study and beyond,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study To- wards efficient fine-tuning of pre-trained code models: An experimental study and beyond,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.645541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.008583Z digest=sha256:adf9f9fddad7cb74df6121a86bd7202efd43a157587931ce241daa4744011b20

Observation d2584707-6a9a-40f0-b8f8-2338e8f740ac · outbound

This paper cites An empirical comparison of pre-trained models of source code,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical comparison of pre-trained models of source code,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.013423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.013423Z digest=sha256:7bd4172c20ba99972ce900f8fe93e6fcddbdd0eed2b21d740b75acf747d27e7c

Observation ab33be99-7a39-4325-9b52-4b05844df9d9 · outbound

This paper cites (2024) Testsmellsrefactoringbyllms.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study (2024) Testsmellsrefactoringbyllms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.629396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.018130Z digest=sha256:515b0f3271e659ff76b4fbae4ed918a267d5883d4b57339154b3bd8253774288

Observation abc5f614-3046-456b-a24d-e7157091aecb · outbound

This paper cites Large language models for software engineering: A systematic literature review,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Large language models for software engineering: A systematic literature review,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.023442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.023442Z digest=sha256:994c03b28b934a39de07ce92662287fcdfbe51a75f83fc03d89233ee9dfd0c8c

Observation 9007f747-60b9-4a80-9dc6-1896b7c41e46 · outbound

This paper cites Software testing with large language models: Survey, landscape, and vision,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Software testing with large language models: Survey, landscape, and vision,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.028343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.028343Z digest=sha256:bafc0c97778d8661351314f49d96651f73ff26ff1f10145d8ae486a27ee15363

Observation 604f551a-1ac6-4e77-876c-0e1f253f2d38 · outbound

This paper cites An empirical evaluation of using large language models for automated unit test generation,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical evaluation of using large language models for automated unit test generation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.034055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.034055Z digest=sha256:5ac306b8b8bd2f36dd3c6ae4b746d14af01b3c1e42890569bee3848c9da6614f

Observation 16a0c54f-72b7-452f-ad91-f24e0a7b3b19 · outbound

This paper cites Automated test case repair using language models,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Automated test case repair using language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.592740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.038870Z digest=sha256:f0cdd3abe0967a2a51439b8b7dba8a9a5a741a461e6c7f4d660c61e9b3a2eb0c

Observation a607be27-a5ec-42e0-9f50-e90ee04bff8a · outbound

This paper cites Chatunitest: a chatgpt- based automated unit test generation tool,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Chatunitest: a chatgpt- based automated unit test generation tool,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.577039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.044486Z digest=sha256:8c0ee15bd79e2cf88964c83dcbf9e93d9e53df0b0b6dcfcfd0e121ee924f472f

Observation 7501baa6-8233-41ba-922e-bc3d59d8bb78 · outbound

This paper cites An empirical study of using large language models for unit test generation,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical study of using large language models for unit test generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.561428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.049310Z digest=sha256:06031962fc3f2268eb5c4a8ff9695d5201b82af086ba84e476c6407feb809c23

Observation d4444700-dfc5-48e5-8936-21cab8d709c6 · outbound

This paper cites Towards an understanding of large language models in software engineering tasks,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Towards an understanding of large language models in software engineering tasks,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.545326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.055046Z digest=sha256:1b9ef4a2db47182019d3e2aae2dd5bacf015dd5b5dd3175688b102e592857988

Observation 01643684-ce9e-417e-b4ce-ee1b3ecd2056 · outbound

This paper cites Pynose: a test smell detector for python,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Pynose: a test smell detector for python,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.060347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.060347Z digest=sha256:01434753dcf80de95651a5d2c1f8165262f653f4c4c0d1c6b521fcd879fd89ac

Observation 37e1047a-9964-4933-ad72-31fc5fd62b2c · outbound

This paper cites Tempy: Test smell detector for python,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Tempy: Test smell detector for python,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.065639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.065639Z digest=sha256:114f7275a2ef9ad132ee10c41f395d024decb9979065478c09ef802b1b15def8

Observation 1808a497-898a-4853-89d7-46f6bb0b93c4 · outbound

This paper cites Handling test smells in python: Results from a mixed-method study,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Handling test smells in python: Results from a mixed-method study,

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:35:41.834355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.070373Z digest=sha256:5a2b9ece6a27f6dc8ffd00d18387849cd646125d0900992531c3ae9a5308a1ce

Observation 0b467529-eab7-4e27-910f-cd8f9cd3ffd2 · outbound

This paper cites A trend analysis of test smells in python test code over commit history,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study A trend analysis of test smells in python test code over commit history,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.525234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.075592Z digest=sha256:d45a279c0cf056c8957d6f97513f70947ed616031cff052bb2707c44d04f8b81

Observation a7830eb5-2748-452a-9657-0adedbb77461 · outbound

This paper cites Pytest-smell: A smell detection tool for python unit tests,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Pytest-smell: A smell detection tool for python unit tests,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.080279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.080279Z digest=sha256:0a69a54c8d9a79f7febb74a8dcf0a56877c24d181959cbb90b05ce26253876e5

Observation 15ad84c6-b314-495e-a34f-c35c84b732d8 · outbound

This paper cites A trend analysis of test smells in python test code over commit history,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study A trend analysis of test smells in python test code over commit history,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.085461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.085461Z digest=sha256:33c287ba5d081ca64be471b8dc9036c27bd58053a7799abe5deb809d8fef471a

Observation 6e9b218f-0a68-451f-b6e1-53b2cb1d204a · outbound

This paper cites Detecting test smells in python test code generated by LLM: an empirical study with github copilot,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Detecting test smells in python test code generated by LLM: an empirical study with github copilot,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.090496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.090496Z digest=sha256:8f6e532bb9ba8b0cbc83b9a1f9d803a5f40e2361b443050154ec12659a77ed88

Observation 27036ac3-7f44-4dc0-90e4-a0c565ff228e · outbound

This paper cites Tsdetect: An open source test smells detection tool,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Tsdetect: An open source test smells detection tool,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.095249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.095249Z digest=sha256:a171c7e52f88b7407641262e493e4f2b65b625e736a6bcb517b8cdb5c7dcc046

Observation f37741b2-b5b4-4428-a81e-8f656088678b · outbound

This paper cites The secret life of test smells - an empirical study on test smell evolution and maintenance,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study The secret life of test smells - an empirical study on test smell evolution and maintenance,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.507134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.100530Z digest=sha256:7432517888b4c0c8c55e08e91466df2a1cdad446e0f8d74df0920d4ca94794d4

Observation c5a472aa-b24b-487a-893f-dded75dc24ff · outbound

This paper cites An empirical investigation into the nature of test smells,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical investigation into the nature of test smells,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.490301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.105536Z digest=sha256:9f3910d072458bdecca159209d148210cd9a52e6765787490a1a932573631c97

Observation 55a94a4e-6461-4fba-9507-0c1234c75a4d · outbound

This paper cites An empirical evaluation of raide: A semi-automated approach for test smells detection and refactoring,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study An empirical evaluation of raide: A semi-automated approach for test smells detection and refactoring,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.473239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.110958Z digest=sha256:cd91df11465bcdc68f40d5fda6740f7f868ec9a026b1313942bcfd15d9bca255

Observation 4c6232e1-13ef-464f-9272-7c5832301035 · outbound

This paper cites Machine learning-based test smell detection,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Machine learning-based test smell detection,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.457599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.116048Z digest=sha256:add1a3a080c505fbdc5787fb34ec86c2f061e71f8fedba62d2504a153b2148e0

Observation 17333b22-3370-4470-ada3-de15e04eaf6e · outbound

This paper cites Ml test smell detection - online appendix,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Ml test smell detection - online appendix,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.441334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.121328Z digest=sha256:e2ef61b31d99201f7027c92a9307523a1286b93f683ac709e604cfcd9b21304f

Observation 979a075e-875f-4c4e-817a-651e9631ee66 · outbound

This paper cites The Prompt Report: A Systematic Survey of Prompt Engineering Techniques.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study The Prompt Report: A Systematic Survey of Prompt Engineering Techniques

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.127695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.127695Z digest=sha256:7d2489d3fe63df49e49185a06da042603dba6b9c53e331e46b1c17acfadb2244

Observation 9e1f8995-92c8-409a-9f5e-50e8760143c9 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.133174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.133174Z digest=sha256:35a4740abb7bd680f880889f3949f84fe5fee4aec46fed7987d637f0b5354b2b

Observation e92c0d03-5ce9-4438-8ffd-bad42eb089bf · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Finetuned Language Models Are Zero-Shot Learners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.138681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.138681Z digest=sha256:61d5ff46d9105702e42d559ae3ff3ed294027319a257d6a2260651af5d4771a0

Observation bcf2c254-a14a-4662-9088-e3d66ed00209 · outbound

This paper cites Language models are few-shot learners,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Language models are few-shot learners,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.143500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.143500Z digest=sha256:3f769330b785c44b67c8c56eaf2731055da662636e3b45f1bef4decfb81937f6

Observation e8bcc4d9-42d0-42c7-bd7c-674bfb4baf9a · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.154116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.154116Z digest=sha256:6b1bb9d68e40cf8e6a310dcf4fbe326fbdc114c3492ba0eeb0475eb3f3f0ce4b

Observation f6690eac-0686-4d46-8b51-481bc0203440 · outbound

This paper cites Enhancing zero-shot chain-of-thought reasoning in large language models through logic,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Enhancing zero-shot chain-of-thought reasoning in large language models through logic,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.414785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.159476Z digest=sha256:b8c268410b49ff57b95e319f27cfe7254d01bea5854a26c30dd9b266feae3e3e

Observation c01ba966-08cf-47cb-af39-6179098f4152 · outbound

This paper cites Wilcoxon, Individual Comparisons by Ranking Methods.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Wilcoxon, Individual Comparisons by Ranking Methods

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.164376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.164376Z digest=sha256:57fff426e6a8e49ecad39f2adccbc9b752e11552b30b700a34fa57a3269662c6

Observation 4b784f3d-ec10-4cef-ab7d-6c9848bebb76 · outbound

This paper cites Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.169640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.169640Z digest=sha256:3547c2fb3480bf289d8fb06ab32835a5d325ba136554015c7334a18cabf69dcf

Observation 203e52a2-3eb4-4dd5-93e2-9ec74aabd054 · outbound

This paper cites Towards effective validation and integration of llm-generated code,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Towards effective validation and integration of llm-generated code,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.399120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.174703Z digest=sha256:cfb01aaf0ae7612266b1a26ad2e030fcd5430b87ed3962547b5aee93f995c809

Observation 4d3c0078-843f-4abe-a90e-af4cd3796809 · outbound

This paper cites Challenges and opportunities in integrating llms into con- tinuous integration/continuous deployment (ci/cd) pipelines,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Challenges and opportunities in integrating llms into con- tinuous integration/continuous deployment (ci/cd) pipelines,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.383731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.179595Z digest=sha256:ff83554c1c32d6f3bb38a53e6157c951739fb316db98d4ad2a6e2baa1a8da111

Observation 3866c2cb-de0c-4740-9325-25b67ecdb778 · outbound

This paper cites Next-generation refactoring: Combining llm insights and ide capabilities for extract method,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Next-generation refactoring: Combining llm insights and ide capabilities for extract method,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.367755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.184895Z digest=sha256:daca01f877200c52b237de6a84b3e6ebb93693d4493259f68a1a65d6ad770271

Observation bcf1d462-afeb-49f6-8b97-81e8497ba694 · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.350942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.189548Z digest=sha256:9623dac20879d65374930d5b9d90fc2f92615d4f5802d962c31349549ba700b7

Observation ba98147c-6e57-4c7a-86bd-39625fad115e · outbound

This paper cites Autorefactoring: A platform to build refactoring agents,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Autorefactoring: A platform to build refactoring agents,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.335039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.194572Z digest=sha256:33f3086b5bca5b56d7b6227412390fbc74e468c8ef17afce241e64182628cf5c

Observation de96ca94-7c8f-43c9-8533-e95f24ffaad4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study DeepSeek-V3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.200114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.200114Z digest=sha256:574c1809caaaf14090bfb052d99baad6ae803882742d97615609f677ea087482

Observation d33ee6e0-9751-4ba5-9b5b-b0d791ab5ddc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.204978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.204978Z digest=sha256:6d3fd8d90fb45b5ec1d1bef7d71c9c001958da189f67e9167d135ab7b4635f01

Observation d9736758-fc93-45ee-819c-908bea04db40 · outbound

This paper cites Runeson, M.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Runeson, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.318400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.210159Z digest=sha256:41aa9f99af126b401b2a403097c31ec4b860c92abe3b2ecd088bc554855af3b6

Observation 9a477ee6-450c-4e12-8886-ead2dd72cacc · outbound

This paper cites Qualitative methods in empirical studies of software engineering,.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Qualitative methods in empirical studies of software engineering,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:35:42.302511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:35:41.214905Z digest=sha256:89335ba8f9fcf5e8a3b63a67a318b345acd94e8d96f93260d0901659da575b40

Observation a9133213-2008-441f-ab73-53fa2c8f35fb · outbound

This paper cites Language Models are Few-Shot Learners.

Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:41.149159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:41.149159Z digest=sha256:c7e701ca575a699bcfa761f7089fdb99f9e77eaaafe628f1358350bfb9b62bf3

Pith citing papers

Observation a8bbc461-08d7-43e7-a15d-eac1263a4d91 · inbound

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code cites this paper.

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:14.700137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T07:59:25.198736Z digest=sha256:317858591aa3bdf218675a390d277b49edad6dac894ad280410001a770b424fa

Observation 8902ec67-7c8b-4674-894b-3fbf0d010e3c · inbound

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing cites this paper.

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.517733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T19:19:18.667967Z digest=sha256:c46811bfc9126f3b4ba8efebd12c3169af8d4a6dbd3cdcdfddd6a619ca8432d9

Observation 1fdc2dcf-d38d-430e-9dd7-f7c350880795 · inbound

Qiskit Code Migration with LLMs cites this paper.

Qiskit Code Migration with LLMs Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-06-26T16:29:35.581303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T16:24:25.357338Z digest=sha256:0ad1e7dccfacbb9c8b1776219dbc838f1f241d0ef14ff7a748e763fc47e6d03b