Pith. sign in

Paper Citation Record · LEDGER

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

As of 17 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 5 inbound Pith citation observations for arXiv:2504.16027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16027 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:15:40.443879Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:10:06.520779Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T07:22:31.204263Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb8905c4-dcd8-4010-98cd-93dbaf7f9ed4 · outbound

This paper cites Coding by Design: GPT-4 empowers Agile Model Driven Development.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Coding by Design: GPT-4 empowers Agile Model Driven Development

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.387113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.387113Z digest=sha256:32a5bf3c9a010d5629cb58630e000596fa0257c8a1ac0dfc9799138b15f71dd6

Observation 24f86a0b-968c-4614-a071-d36255375d0b · outbound

This paper cites Evaluating Large Language Models in Detecting Test Smells.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Evaluating Large Language Models in Detecting Test Smells

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.397416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.397416Z digest=sha256:9fd10916f97d4a57233787640390e8f9897ac437861eb224399ddb14b7998aac

Observation c8d06b7e-9688-49bb-bc04-b9eb0b08fdf5 · outbound

This paper cites How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.402491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.402491Z digest=sha256:982f185f0eaa05164c182607d2407c330e8bf51853785196ee774c7f63102b1d

Observation 9a07f0e5-c5e9-4509-a6be-83b625b0e0b7 · outbound

This paper cites Are sonarqube rules inducing bugs? In 2020 IEEE 27th international conference on software analysis, evolution and reengineering (SANER), pages 501–511.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Are sonarqube rules inducing bugs? In 2020 IEEE 27th international conference on software analysis, evolution and reengineering (SANER), pages 501–511

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:15:40.656078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:15:40.407496Z digest=sha256:2919ec439d5be99d1f7928a774d9542ec6ac0e719b95894d35329dfcee53d12f

Observation 791c30f9-3841-4da4-9c8d-9b99a15567a3 · outbound

This paper cites Code smells.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Code smells

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:15:40.640971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:15:40.412578Z digest=sha256:ef9bd55cd476152c9fdf9bfaa45e2a4181a3b22a3d5ecb022ee7d62e4acb1964

Observation c8a25a31-156c-4489-9809-05d30c0b66c8 · outbound

This paper cites Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.435127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.435127Z digest=sha256:cb3609f8bbe1f24b6226feaac2028aece611e2b9f49e76af24cf89f60dfba7b1

Observation fccf7bbb-24b7-4e93-8749-f80df37bc54f · outbound

This paper cites DeepSeek-V3 Technical Report.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.443879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.443879Z digest=sha256:6ce190a277b510f5a831a249b58f18abf716d41821e754df0fa38752cd2642ff

Observation 38e845e1-9557-4a24-a77e-4580133b5488 · outbound

This paper cites Code Smells in Machine Learning Systems.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Code Smells in Machine Learning Systems

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.416855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.416855Z digest=sha256:ee814a9d5a61279012dfc908c349a5daf28a89fe09d1100239e16b0bd6e3775d

Observation 3d07daf0-8c55-4afd-a766-2a805a40ee2e · outbound

This paper cites Causes, impacts, and detection approaches of code smell: a survey.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Causes, impacts, and detection approaches of code smell: a survey

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:15:40.626743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:15:40.421733Z digest=sha256:52769d0a195d70ce9a8626b6c06f33da69d859dc27d459b9a6693ebdd21465cb

Observation 80fd2e32-b144-4b00-9c33-776fb8c35c64 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:15:40.611206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:15:40.426247Z digest=sha256:0f71bce908d9c70e68c16bc21972d848548006f0f3f6f991d3905f65e318d2b9

Observation c6b4c137-0512-471f-b1b3-601e0ac00e62 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:15:40.596615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:15:40.430968Z digest=sha256:35fb9e1abc68e817c3435b102d608fe27ee006398ce4772359a4ee0702008d0e

Observation 1ba3ae49-a05f-4ab7-9734-54e5ba929900 · outbound

This paper cites ChatGPT as a Software Development Bot: A Project-based Study.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 ChatGPT as a Software Development Bot: A Project-based Study

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.392603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.392603Z digest=sha256:dd1c6bc9d3d0561cf07e8482e78024a862b8bce882a5456f8179e99a4efe9fac

Observation 3697bb99-00f1-4bae-848b-541a080f89d3 · outbound

This paper cites GPT-4 Technical Report.

Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 GPT-4 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:40.439675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:15:40.439675Z digest=sha256:43399105f37d393212ce0bbb3e051ed5bc995189c945ecd6ae4732379165acfb

Pith citing papers

Observation 5ea1dce1-7027-46f5-ac13-12847ff76092 · inbound

Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations cites this paper.

Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:10:06.520779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:10:06.520779Z digest=sha256:85476fd4bea629144199e2e3ec3f8266b4c8c1558722a71b137ed99a31a230ab

Observation 7b2aeaaf-6b14-4f49-a423-6fe8031865a7 · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:22:31.207472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:497ac063e79d3db18288b7111bd8fdb41770a8169695d0ba9be89f3ba87ce1d3

Observation 4bef57b6-c9e3-4791-bd69-b67b2f4bfe64 · inbound

Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions cites this paper.

Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T23:04:36.291592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:04:36.291592Z digest=sha256:a236b90c9dacac84351d7b61283f08d71c3c295d3a9bdacf3f6f20e9c1049b9c

Observation 1ab5abea-5ee6-4819-88d6-1acc5e1c087d · inbound

DynamicsLLM: a Dynamic Analysis-based Tool for Generating Intelligent Execution Traces Using LLMs to Detect Android Behavioural Code Smells cites this paper.

DynamicsLLM: a Dynamic Analysis-based Tool for Generating Intelligent Execution Traces Using LLMs to Detect Android Behavioural Code Smells Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:51:01.312074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:46:35.539476Z digest=sha256:f93a0820011583149c8e9b7ad3bdb96cd5d7bc5dd3d705551b91d8328ea089be

Observation c91ab388-4f3d-49ae-8a5b-ecda5e0ea9db · inbound

Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts cites this paper.

Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T11:57:56.953101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T11:57:56.953101Z digest=sha256:4f8f79f839894596eb733096db4ef3549270b4f82061c674486485e2bc084f45