Pith. sign in

Paper Citation Record · LEDGER

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.06980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06980 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:57:13.152676Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd190d38-f3f5-40f7-9983-5ec715a5a2e7 · outbound

This paper cites https://deepmind.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation https://deepmind

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.934083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:12.986869Z digest=sha256:5c7c47bdf591ac823f867f86600df7859d4910e2b9ab3d738405dd784b66eaa2

Observation e420d59f-78dc-4941-9f59-94a7e5a1999c · outbound

This paper cites https://platform.openai.com/docs/models/o1/.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation https://platform.openai.com/docs/models/o1/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.925186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:12.990649Z digest=sha256:a0e4befcad3937ff0bc7048ea4efe8e6d981e7f3ec749abe449d8683937e5235

Observation 0dd13243-0a6e-4ccb-b74b-5f82f94bdd86 · outbound

This paper cites Let the llms talk: Simu- lating human-to-human conversational qa via zero-shot llm-to-llm interactions.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Let the llms talk: Simu- lating human-to-human conversational qa via zero-shot llm-to-llm interactions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.916708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:12.994078Z digest=sha256:36409ed1551f8c3ef60873aa2732b264b2ac93c57ad63982e4d2a656a91d15d9

Observation dddb8680-2d31-4975-acb9-d35b4364e831 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:12.997242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:12.997242Z digest=sha256:7871d46ff5c197d8a6beb26649a62aa720b701f501d70a52c23d16abc9672ec5

Observation 090910c8-f8c0-45e2-b788-b55841329f11 · outbound

This paper cites Is github’s copilot as bad as humans at introducing vulnerabilities in code? Empirical Software Engineering , 28(6):129, 2023.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is github’s copilot as bad as humans at introducing vulnerabilities in code? Empirical Software Engineering , 28(6):129, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.908021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.000811Z digest=sha256:62daced6d7466d269bf7f07c56a13abb336b0ed1c2a50e7bdfb364aec27f5531

Observation 05f384a1-10a6-4683-be98-6259d406079d · outbound

This paper cites Language Models are Few-Shot Learners.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.003900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.003900Z digest=sha256:ac4d13f59332c424e9a9465c1d611e81af1424e2603034d197f37c46396f6124

Observation 045be678-fa73-44e7-a757-e3b6f3933bdc · outbound

This paper cites CodeT: Code Generation with Generated Tests.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeT: Code Generation with Generated Tests

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.007465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.007465Z digest=sha256:53e7849c84d04bd187a0586502b2ec11094bd534111d1aa457634cac0505388f

Observation 148742a7-2c10-4e69-8977-1e0c1515a2ed · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.010764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.010764Z digest=sha256:d4e8b42f41decce09dbdd48cdc78a9d413979a0b0e5299bc3b176997dfb97db6

Observation 225d0a00-59af-4c61-91b4-d43e484106a3 · outbound

This paper cites LocAgent: Graph-Guided LLM Agents for Code Localization.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation LocAgent: Graph-Guided LLM Agents for Code Localization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.014157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.014157Z digest=sha256:b0313f6af9d2fcde329889c2a74523bb4e9df4eb598b80829e7fc5509d1d3c89

Observation 80cf230c-d9ba-41d1-90f5-b62e0adc921f · outbound

This paper cites TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.017270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.017270Z digest=sha256:b425e252ec0a1151366210fe9e281510072ed9e3937a5fbdb7549319be3b1b22

Observation 32968c5c-115a-42f3-b597-5eced2799e7b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.020754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.020754Z digest=sha256:e9f06c585961f01125e42ddc195bc741c2d38efd4ffc3cbcb6f9d7dc60fc4e65

Observation 9e300668-1bd3-483e-8cd5-14f4e1a3a8cf · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Reasoning with Language Model is Planning with World Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.023840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.023840Z digest=sha256:cf09efef5b680eff6a436b5655a8118d6bb2886b3197c783a36b29a688b96262

Observation 0051cb81-9ce3-476a-8660-fb9c4255a904 · outbound

This paper cites CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.026859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.026859Z digest=sha256:34fb0fdad3ee24f0013f0ac0beed2ba81f4a7e05d770a984254619444170ec5a

Observation ffb95e5f-90fc-4f40-9179-e16ed179174e · outbound

This paper cites Adaptive mixtures of local experts.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Adaptive mixtures of local experts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.899144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.029804Z digest=sha256:1eaa86d9c9fde20e0ab7b76d58180a96b3899a77b71f1ef690d43bf37a438e48

Observation ba85e74d-42fd-45f5-948b-472c83fb2b4a · outbound

This paper cites Devanbu, and Emily Morgan.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Devanbu, and Emily Morgan

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.032645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.032645Z digest=sha256:4f6eeb11e4b75f6e02884c1a2bc1742d1a500c41ca4a3fb1bff73444e87a01c5

Observation 9698e8eb-768d-4d64-a7ba-e16e2bd72a79 · outbound

This paper cites Self- planning code generation with large language models,.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Self- planning code generation with large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.890614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.035407Z digest=sha256:e13789d3240884293b3467b71c11438d804335e4d761f1eaa013a98df21d065d

Observation c72dfe49-e09c-4735-9f4c-533e027ca233 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.041485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.041485Z digest=sha256:0611989ccb4310c3482b568bc98da8fafacb72011741fe31d74caf86f4f9cd15

Observation d9976e03-e007-4c23-ab53-7cb684635acf · outbound

This paper cites Grace: Discriminator- guided chain-of-thought reasoning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Grace: Discriminator- guided chain-of-thought reasoning

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:57:13.499467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.044893Z digest=sha256:c5787b8b7f6ab9cc6b594812d19b6ccc4061597d190c4ad9a57986a6a8de70e9

Observation cc63a596-dd09-4fde-9c68-dca2c4fcba1e · outbound

This paper cites Large language models are zero-shot reasoners.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Large language models are zero-shot reasoners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.881690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.047794Z digest=sha256:d37cd0051ed690c68819346fa3bb7e493ee52abbbfa84f8fdc75b1b71ec54876

Observation c3beeed5-0e39-4cf6-b8c5-fd66db7ff908 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.050671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.050671Z digest=sha256:f5ef3ad717532f9e76c99b09c42d24c62a9be920eb98b5e6aa2dbe9c7394b82b

Observation 65298404-a8eb-4a09-a9d3-95c87983cf7a · outbound

This paper cites CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.053921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.053921Z digest=sha256:cbc3f3d49c4a60091c1ec5a2b541577afe07c51f8795e31fcb452e91e96659fe

Observation 0695e282-24b3-438e-9c2b-e81d33b5cdbe · outbound

This paper cites Structured chain-of-thought prompting for code generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Structured chain-of-thought prompting for code generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.872380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.056940Z digest=sha256:c159d3da42bdea1801100edfa7937b5e919057a704fc80ae603b63347f10bb10

Observation 882c4230-c241-4e77-bd71-ce8f3d35ecb4 · outbound

This paper cites CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.060665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.060665Z digest=sha256:f44cab1aaa8738c9d4077b6c0266c8cb182d5f0af075e2562b1c6055e0a6289b

Observation 19e30428-4e58-4c6d-a4bd-a0a367e5d86d · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.064699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.064699Z digest=sha256:969b53ddd73d06f982bf697087cf24e7d557d070361b25e4880c20a41dbccb35

Observation 9bfbd08e-aa7c-45d3-8557-4f96bc069445 · outbound

This paper cites Fastfixer: An efficient and effective approach for repair- ing programming assignments.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Fastfixer: An efficient and effective approach for repair- ing programming assignments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.863794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.068804Z digest=sha256:7b46bc80c67b964a4a98b6026a2975ab5623e5f6e20a61b67c19556ef025f0ac

Observation 6c66ca81-1f1b-49cd-8d84-7eaa522da989 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.854116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.071864Z digest=sha256:7f1161ede61eb90fac12cc3368853d988bcdd4da8142c592c571f41d4c17e884

Observation 8e83878d-5eec-411e-9108-24a62f1728de · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.843107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.074817Z digest=sha256:df0f7333c966ac53c83a87abfd0a97d9388150e0d76d4c318860756e99df930a

Observation b52eba2d-4f23-4672-88a9-095b01315893 · outbound

This paper cites Refining chatgpt-generated code: Characterizing and mitigating code quality issues.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Refining chatgpt-generated code: Characterizing and mitigating code quality issues

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.834059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.077529Z digest=sha256:c0e9cb4d5794ee353c1e8b1b54a65efae395703babb1d1413abeb71cf180d1ed

Observation 1257cabc-3675-4503-869c-657aec8b956a · outbound

This paper cites No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.080535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.080535Z digest=sha256:cbf14c0af10caee44ad560f5cf17be57ecf6c71d41ae27513837d6717cbc81d2

Observation 799725ee-bf4a-4800-aec3-7672cdaf2d4e · outbound

This paper cites Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.083644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.083644Z digest=sha256:8e053037d3d730bdb7ab81e8a0883b9d5ae808da8e5602c2fa4d5136d2b82490

Observation 625000bc-dd32-486c-9730-b64f0772c4aa · outbound

This paper cites Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.086901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.086901Z digest=sha256:531272b7b09499c226c08252e99b67801759169b5e18855472c38da42c6ea718

Observation cdb89cb4-0d81-4e13-aba0-cd6f81df5a2b · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.089866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.089866Z digest=sha256:944e2ac4ae665644bb3947342895bb328523189bf669a6a6693692043e1615ba

Observation 1ab23470-2b95-42f2-8ac8-d0fe97ba9b41 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.092944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.092944Z digest=sha256:c1c0f58622fd4cc6a97aae539747a93a5f7afd6b998a8d97242c34b247324140

Observation dedb6c62-79e1-4381-af05-d974071d4622 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Code Llama: Open Foundation Models for Code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.096040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.096040Z digest=sha256:fcfb205b572f88fceea0e02fd73d38ea2c98a1f62fe666e843f76dcbe20a02b5

Observation db44968b-3743-45c5-aaa2-7eae9b221c1d · outbound

This paper cites Amazon codewhisperer, 2024.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Amazon codewhisperer, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.824385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.099281Z digest=sha256:b753a6a0c13d812aea42d5f4a91d16c8fcdf2640ccc37676a5fb44eb741cdd20

Observation 74954589-49b8-41f1-8aa4-c9aae5fc6359 · outbound

This paper cites Calibration and Correctness of Language Models for Code.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Calibration and Correctness of Language Models for Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.102050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.102050Z digest=sha256:23cb8f81fbc7be7dacdc7c879ef275b8b5e5869b61f653ae0f32b24f2d533202

Observation 0f73a6b6-88ab-49be-9bff-5e2aa94ad9c6 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.105318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.105318Z digest=sha256:697dc1f9eae1f166af81075f011b06743b3066579af764699bb3a448baaf9e27

Observation c2cd16ae-52fb-4962-b53a-3e929e395b55 · outbound

This paper cites Bugs in Large Language Models Generated Code: An Empirical Study.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Bugs in Large Language Models Generated Code: An Empirical Study

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.108562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.108562Z digest=sha256:352995bed2c74216004d9fedb9655fd49abf3da5db46a82c48eee6310b202185

Observation 38217dfd-d3c4-471e-8674-c43996863060 · outbound

This paper cites Code repair with llms gives an exploration-exploitation tradeoff.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Code repair with llms gives an exploration-exploitation tradeoff

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.815606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.111596Z digest=sha256:d170e80267ce2cf8a7b4827d08c394430af0b83c821caa388e14c535b88e18f8

Observation 742f76d4-a924-4161-b7e2-6746303a1dee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.114357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.114357Z digest=sha256:2cee9ecc19de4fbd4e786091b0edeb690306e61c5d943cc7a6f1d8823b64bc0e

Observation b6faa49c-bc50-40aa-a7a2-05acbcfa9af2 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.117461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.117461Z digest=sha256:d5ee0cd7029c5996c7143a5568d66d42f680c6d52fbf8154fcfbe718155cee2e

Observation fa0f77b2-7ff5-4ae6-929f-ed9c8c3642ef · outbound

This paper cites Expectation vs.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Expectation vs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.801421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.120602Z digest=sha256:b80a61692700766671a5393bdd0cdf5132937d151e0f897337ff4c9879e168ca

Observation 44cb0885-084c-434e-b6cc-0b8d01df9bad · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-thought prompting elicits reasoning in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.123452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.123452Z digest=sha256:ac52d34042cbb8650330d7b75ba733d5bd1e67e8b809cfa86da20d117d0f5327

Observation b9478c76-6c3f-4f23-b454-ca30fe6e3452 · outbound

This paper cites Using github copilot to solve simple programming problems.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Using github copilot to solve simple programming problems

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.786801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.126255Z digest=sha256:20bf277c8216d387abd51c40272e4b28e2d583e02d8ead909c55e1ff7f282428

Observation 8ec48f14-0c06-4991-a1aa-9c8606220ce2 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Agentless: Demystifying LLM-based Software Engineering Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.129076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.129076Z digest=sha256:ba9df57f9a7053764f766eaa4c1f9c9fe14b946bb3f448b56b91a12302386408

Observation 21c5a2ce-f01a-44d3-961e-dbda073ef8af · outbound

This paper cites Demystifying llm-based software en- gineering agents.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Demystifying llm-based software en- gineering agents

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.777752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.132297Z digest=sha256:5642fb457d7cda92c1529067319a183329742056cacbcb6bb9353ecc12032eb5

Observation 189cd40d-6060-4559-9255-5475497315fb · outbound

This paper cites Self- evaluation guided beam search for reasoning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Self- evaluation guided beam search for reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.769096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.134962Z digest=sha256:5e725709462908646d69cb0263452a3ba84cefa3e42d06cdd1c72ddf94c84011

Observation 6216f725-8dfc-4c42-a584-667df91f637a · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.137811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.137811Z digest=sha256:47c04af94bf88bf1b6b54db134044f7b755235add235327205a876d614b3f19c

Observation 4486a841-2aad-405e-b4ff-bfb4a2205713 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Tree of thoughts: Deliberate problem solving with large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.760345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.140899Z digest=sha256:a48938adbc948431678eef57afdb080e807f7cc4e826df62f5eb9341560e5144

Observation 0213a368-cc4d-48a4-a603-1554521f1d9c · outbound

This paper cites Framework for evalu- ating code generation ability of large language models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Framework for evalu- ating code generation ability of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.751610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T18:57:13.143896Z digest=sha256:42368d19fde7cdf9941af157ba3465289486f4b6e104a738b107a69ba7409e13

Observation 8dbffacd-d2e2-48b2-93b9-8ecc1bef1d57 · outbound

This paper cites Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.146831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.146831Z digest=sha256:d113d42ff9d7812b7d1fe96c1e5772c15b8236b8d7cef11237dd90b263aa516d

Observation 27c36329-ccc9-4fb1-a4f2-2a3602d51c47 · outbound

This paper cites Learning-based widget matching for migrating gui test cases.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Learning-based widget matching for migrating gui test cases

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.149843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.149843Z digest=sha256:b73e151c16f8ebf17b76b5f0d37d3360fe41b5f32d3c2b8e9e66454fc5a7b58d

Observation cd830f19-23cf-4218-b2d1-0d95568360ef · outbound

This paper cites ToolChain*: Efficient Action Space Navigation in Large Language Models with A* Search.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation ToolChain*: Efficient Action Space Navigation in Large Language Models with A* Search

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.152676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.152676Z digest=sha256:94519315a6528d96a70b92f5cfc30cb5dfd3ddd727da71f2328a73ddc87fb24b

Observation 3998ee99-e134-430a-a740-63a4965287f0 · outbound

This paper cites an unresolved cited work.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.038475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.038475Z digest=sha256:6484fd95259e7d1d41f1d87415c71a527882cd60d02f700b9b007f27f3079c14

Pith citing papers

No inbound Pith citation observations are available.