Pith. sign in

Paper Citation Record · LEDGER

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.00661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00661 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:05:44.229751Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15cba260-eddc-4c60-888d-f72311fbdbe8 · outbound

This paper cites https://web.archive.org/web/20250430014530/https://www.cnbc.c om/2025/04/29/satya-nadella-says-as-much-as-30percent-of-microsoft -code-is-written-by-ai.html.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code https://web.archive.org/web/20250430014530/https://www.cnbc.c om/2025/04/29/satya-nadella-says-as-much-as-30percent-of-microsoft -code-is-written-by-ai.html

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:41.675067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:41.675067Z digest=sha256:052ea1f55fdcbdbf830f6d93d8bde7667fe190a89e5692e5ba2626928a253ef0

Observation 33e7358e-1098-4ed8-b038-f3a6dac0321a · outbound

This paper cites Microsoft cto breaks down how he sees software developer jobs evolving in the next 5 years,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Microsoft cto breaks down how he sees software developer jobs evolving in the next 5 years,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:41.751618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:41.751618Z digest=sha256:893278e6ea8b878b3e1fb2272a8d45e55ed7783b3cbd2e183f8fc3728ae4f8f5

Observation c4267c5e-f440-42dd-982a-10b61c90636d · outbound

This paper cites The impact of ai on developer productivity: Evidence from github copilot,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code The impact of ai on developer productivity: Evidence from github copilot,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:41.806283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:41.806283Z digest=sha256:d3a574ca86d16a50fa05374e05b323a68eff64637c979d4321fb67bc9fe13893

Observation 63c78d1e-7012-4c52-af23-79978bb34b1c · outbound

This paper cites The effects of github copilot on computing students’ pro- gramming effectiveness, efficiency, and processes in brownfield coding tasks,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code The effects of github copilot on computing students’ pro- gramming effectiveness, efficiency, and processes in brownfield coding tasks,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:41.894341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:41.894341Z digest=sha256:d6e198628347085e7c4f0af6e597366abe836ed49328c9085e445c59bfd0b3be

Observation 40e73a8e-6593-433d-90ab-521a6f8a380c · outbound

This paper cites Measuring coding challenge competence with apps,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Measuring coding challenge competence with apps,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:41.981644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:41.981644Z digest=sha256:66183ea544554ae5cfa543d0579ba90e92c5da7db5cb23a86c2b39e10f6bb211

Observation 02c554f0-847a-4efa-9653-266f8cb0d0eb · outbound

This paper cites Beyond functional correctness: An empirical evaluation of large language models for text- to-code generation,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Beyond functional correctness: An empirical evaluation of large language models for text- to-code generation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.076119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.076119Z digest=sha256:2bb6f98df1e68ec7bd164e6f3c7fdb42ca44f5af07e2d01d70518e23d8da5903

Observation e540e0a8-7a88-4b51-b22f-80d0b6a3667a · outbound

This paper cites Exploring and evaluating hallucinations in llm-powered code generation,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Exploring and evaluating hallucinations in llm-powered code generation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.179583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.179583Z digest=sha256:8daf9e56ad7fdceec5468bfe74a41b311b0a05cf32fd6caade7bf1d5267118e3

Observation c79a2b6a-8d52-4197-adf3-520f7031fc04 · outbound

This paper cites Towards understanding the characteristics of code generation errors made by large language models,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Towards understanding the characteristics of code generation errors made by large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.273701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.273701Z digest=sha256:e7dc7b5b3a8f10a5574d6783d63b383477208ab8ad9947cd8ed6fd160854a087

Observation d4d30bed-c8c7-4c0c-b68d-37316fe2a135 · outbound

This paper cites An empirical study of code generation errors made by large language models,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code An empirical study of code generation errors made by large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.375284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.375284Z digest=sha256:01cb4ff88f55af4126939f2b9e8257bf2c64f231e8b0930b674d97d6ac21d504

Observation 08904398-f20a-4d56-b733-bd39957c8e55 · outbound

This paper cites Bugs in large language models generated code: An empirical study,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Bugs in large language models generated code: An empirical study,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.455980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.455980Z digest=sha256:a3e0d0cd87fdfc3b814bbbfe07bb989b2de6ec2322f711a9a9a0dd978177c6ae

Observation 5c7eb81b-0e6f-4aaa-b5d0-da8c0a08f5f2 · outbound

This paper cites No need to lift a finger anymore? assessing the quality of code generation by chatgpt,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code No need to lift a finger anymore? assessing the quality of code generation by chatgpt,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.569627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.569627Z digest=sha256:163ebcf9db2e5316b00753ea05bd4854b6d36faad7bee664d91cb1266c28cf7c

Observation c89a02d9-deb8-4ece-a8f0-61a8e2ff2d53 · outbound

This paper cites A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.673499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.673499Z digest=sha256:2914cbe770ddfe50d01c900fecf206d09e43152fbd4190f0857812b803201e38

Observation 9c76c0c5-00db-482f-a258-b5f56c21ae81 · outbound

This paper cites Assessing and analyzing the correctness of github copilot’s code suggestions,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Assessing and analyzing the correctness of github copilot’s code suggestions,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.770880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.770880Z digest=sha256:254e446c8ad0257d7e00df7a3c91c0aef83e9a35160c471e95ac3030a3bfb6c9

Observation 16bb1223-c45c-414d-a8a3-2027d4653bb9 · outbound

This paper cites Gpt-4 technical report,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Gpt-4 technical report,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:42.868042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:42.868042Z digest=sha256:ca136996277e2e9148d4e20a62cbe8413248396715ec807c48884c1f56950d26

Observation 75978369-e80b-4df0-ab41-b283685d77f5 · outbound

This paper cites DeepSeek-V3 Technical Report.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.031051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.031051Z digest=sha256:364d8c34fe875a444d28b38e01db560d006b8ddd6a977fc28f04b8fcd0f0e253

Observation 25b21f5f-1578-42f9-9ef3-0283fbcfdcd1 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.179400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.179400Z digest=sha256:36eac0905b3a7761995e0a7a8f4c8e254e84feac0388041e5fac3cb6cf1390de

Observation d542d80a-b233-43b3-a295-3b6a23f6cc68 · outbound

This paper cites Qwen2.5-coder technical report,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Qwen2.5-coder technical report,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.291943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.291943Z digest=sha256:74c9df548e6925c7a3ec85cb07f4d81fb283ede7bcb690fd7b7c01e832fc7a49

Observation 59edae2c-ad76-45a4-ac5d-214238c1054a · outbound

This paper cites Incoder: A generative model for code infilling and synthesis,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Incoder: A generative model for code infilling and synthesis,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.426461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.426461Z digest=sha256:3bf4e44f9073d8d71b58334a533eea749298908044efc8536466402e35d07a1b

Observation e87f8893-dafa-4dd2-bbd2-9411519eb854 · outbound

This paper cites Evaluating large language models trained on code,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Evaluating large language models trained on code,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.563626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.563626Z digest=sha256:713de31e7e7d228d3e5108d794a10bfecb56202aa3e93a28f72a6118cb5f8f46

Observation b9dd5a41-8d70-4fd9-a21f-f294b5fd0d4c · outbound

This paper cites A large scale study of programming languages and code quality in github,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code A large scale study of programming languages and code quality in github,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.649024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.649024Z digest=sha256:89eda73b45269f91297aea28ee59f0b2c30cd4e8abd672bcd61f3a78e35e014e

Observation d83f5a41-8d31-4bc2-9944-9ac88e88ac1e · outbound

This paper cites Orthogonal defect classification-a concept for in- process measurements,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Orthogonal defect classification-a concept for in- process measurements,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.749086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.749086Z digest=sha256:a9877cd26b8822e14589d202fbefc583191769bd5ff0d5ba241d31718fc5d744

Observation eddd760d-3cf2-4af1-929f-f1b3127f923a · outbound

This paper cites Program synthesis with large language models,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Program synthesis with large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.807074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.807074Z digest=sha256:e28013f74c396a40713168e2911e36be190c82c7cd11b99d40d745a9bac7a93b

Observation 9ea8050f-a33b-410e-9c72-78f987353dff · outbound

This paper cites Codexglue: A machine learning benchmark dataset for code understanding and generation,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Codexglue: A machine learning benchmark dataset for code understanding and generation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:43.878756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:43.878756Z digest=sha256:2faec7f6ae2bb09d0e04e544f7ee0b29326ba1b908d3cf77caffd1bb4df3a9fb

Observation 0f48bb42-a000-4a5c-a932-9d11f02dcaab · outbound

This paper cites PROBE: Benchmarking Code Generation in Large Language Models.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code PROBE: Benchmarking Code Generation in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:44.021875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:44.021875Z digest=sha256:5d5e182e2abd345a658ebb6a6fa389a4f5a5f0dbeaaf7937b992bbb2208c769a

Observation ac0aedb3-fe5c-4fcb-be57-6082c2817cae · outbound

This paper cites Codenet: A large-scale ai for code dataset for learning a diversity of coding tasks,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Codenet: A large-scale ai for code dataset for learning a diversity of coding tasks,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:44.092272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:44.092272Z digest=sha256:b1fbd5298da6278466c8c3e05e7ffebbe9935ac573fd65ecf67f3322dca126a6

Observation 3e0c5cb3-7c1e-488f-8062-83ea873859b7 · outbound

This paper cites The measurement of observer agreement for categorical data,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code The measurement of observer agreement for categorical data,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:44.132381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:44.132381Z digest=sha256:7427d2f043abb6ce48f055f78fa6bf3728340dc1029cc15e8a7dbe321edf1dca

Observation 4704d800-7f0b-4228-91b7-d3baf708b730 · outbound

This paper cites Polyglot: An extensible framework to benchmark code translation with llms,.

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code Polyglot: An extensible framework to benchmark code translation with llms,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T01:05:44.229751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:05:44.229751Z digest=sha256:89f5bb0e1f88c54b2b1938ed5f8dca909aeb78d3e8ebf7d71aefdb5c6957a044

Pith citing papers

No inbound Pith citation observations are available.