Pith. sign in

Paper Citation Record · LEDGER

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

As of 18 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2505.19502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19502 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:49.895278Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:33.789468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:32:34.591005Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8c3daba-9935-4ca6-9220-153383d74bfc · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.793868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.793868Z digest=sha256:35b8b32ba16ed86a1205e3752a2865232fdacfb20cc4d2b50df7c9fb0307e12a

Observation 37ac8ab1-e125-492e-a656-74e769683d7f · outbound

This paper cites Efficient and green large language models for software engineering: Vision and the road ahead,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient and green large language models for software engineering: Vision and the road ahead,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.852067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.852067Z digest=sha256:c60d435ac8ef0eb17fb2d0c8aedcf5e9f21237176d8cdbbeb591b97f9fb81fce

Observation 7059d352-e049-4c97-b5ff-c4b3d8c384a3 · outbound

This paper cites Large language models for software engi- neering: A systematic literature review,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Large language models for software engi- neering: A systematic literature review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.977776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.977776Z digest=sha256:c0664acedd98212cb492562cb992673d5a236ee423d50abb1c75958c649ef75d

Observation d5db457a-4915-4066-a1b2-91a769ee34f4 · outbound

This paper cites Exploring the capabilities of llms for code change related tasks,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploring the capabilities of llms for code change related tasks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:56.142825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:44.075982Z digest=sha256:964ef87cbb6a9ec82c9b6167ba2bdc75eafdeffdf38d2453ab54d3774613bcbe

Observation 436ddac7-8fc6-453a-bd73-f5cb6ff0165b · outbound

This paper cites Chain- of-thought in neural code generation: From and for lightweight language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain- of-thought in neural code generation: From and for lightweight language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.190067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.190067Z digest=sha256:bf6933c7954f2f47d8991fb26244bca6afeb3ca252eaa200edd8b6da2de85d62

Observation edc0ac39-d979-410b-899e-af50a5ee8016 · outbound

This paper cites An empirical study of retrieval-augmented code generation: Challenges and opportunities,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation An empirical study of retrieval-augmented code generation: Challenges and opportunities,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.323662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.323662Z digest=sha256:a31fbee2097be86e925b8d8a765cf93266375da23c0c175cd8488f3c4760df5e

Observation b8e62aa2-67fc-4b4c-8ff7-ee8e4befcdde · outbound

This paper cites A review on code generation with llms: Application and evaluation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A review on code generation with llms: Application and evaluation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.419952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.419952Z digest=sha256:017f147d0eb80a1eb19c1bfed56d43f89231a5d051e4223dba5ff2d91c225b85

Observation a016940c-4a33-4e26-af3d-37407ca86c7d · outbound

This paper cites HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.535532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.535532Z digest=sha256:5d0050385dd46e36f97201810c005bc9630f6bb7953baae00f01c260e295aa1a

Observation da753d1c-7523-4255-8b2f-db694ff4245c · outbound

This paper cites Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.722335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:44.695755Z digest=sha256:32c0f9c0e8c7a1a1edba08442b63d495356291ea5388e2c7577eb647e6aa0632

Observation 6f840e5e-74ea-4099-b2ff-15a0f23c8e5e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Bleu: a method for automatic evaluation of machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.800778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.800778Z digest=sha256:ae57dffcabd9c8cd3f63955343a21ab04508ee199cc7792d39304b8b80b37915

Observation be427ef5-801d-43d2-b29f-3dd2615fcca8 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Rouge: A package for automatic evaluation of summaries,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.882525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.882525Z digest=sha256:3d403fcc39e111ecd91454290f7ad89d6c43134c111695f6bb004aab110cb6db

Observation 61f6370f-1789-4085-b2ff-43d2915ce993 · outbound

This paper cites chrf: character n-gram f-score for automatic mt evalu- ation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation chrf: character n-gram f-score for automatic mt evalu- ation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.399227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:45.020391Z digest=sha256:d883627fad4ae375683ba392b50860672584b6a360b8d83aa90f56c549375861

Observation 788e8390-42a0-43fb-bbc9-d3f73c7c27f7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.176339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.176339Z digest=sha256:777e576d3ed456d200be8cfb2c5ea91747b2ad63cf3f65e02157ffc9fcdaff51

Observation 7bd5d4a7-0764-482d-93dc-49dec8202507 · outbound

This paper cites Exploit- gen: Template-augmented exploit code generation based on codebert,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploit- gen: Template-augmented exploit code generation based on codebert,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.097618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:45.322765Z digest=sha256:96747808a05da756fb5c2583ba98da9e18851a8cb579a7612e61ed8bbcfba8b1

Observation 328742c4-9d6d-4754-9765-2a5d317b0ba1 · outbound

This paper cites Are nlp metrics suitable for evaluating gen- erated code?.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Are nlp metrics suitable for evaluating gen- erated code?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.811838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:45.485559Z digest=sha256:372a66653153f4a6093c6547972c5814538f2f70516562aaf6a2fd1103a0ac9e

Observation af2248f9-f5c9-4c45-84c7-7ea2822b126d · outbound

This paper cites On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:51.337928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:45.617534Z digest=sha256:62f076615efeb2ac6e584c871c34939c46ee6e79ba816fe4a8fabd8390160fed

Observation 7f780868-1084-45f3-a25b-b9518ed9b730 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.728953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.728953Z digest=sha256:9cd32bee06770cb085d1da49fc017cc8fb8cbe01e0929a073dfad4c3a8751f5f

Observation 606da85a-ab5e-4433-87a0-1a1720486b5f · outbound

This paper cites A Survey on LLM-as-a-Judge.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A Survey on LLM-as-a-Judge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.866412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.866412Z digest=sha256:d20d2ced6e5b167c54e1f6bdfed9811066498424fdeb7e019163d3d0a16893a1

Observation 6591cce8-2b4c-45b4-92b5-db04020bcf7d · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.986897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.986897Z digest=sha256:09b2bad035187dc3d755954a8b9c0ed4b40fc43698973c3d100a0032aec5a05d

Observation 954b232b-51e2-48e3-9a83-80db1bc437ab · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.158301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.158301Z digest=sha256:493910e932916aa2ca76b35412930c75511662446a51cf6d6e296232432f7f78

Observation 2dc63c72-91a8-42d9-ac61-f8f830f53b5b · outbound

This paper cites Fight fire with fire: How much can we trust chatgpt on source code-related tasks?.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Fight fire with fire: How much can we trust chatgpt on source code-related tasks?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.479297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:46.267717Z digest=sha256:8e284727113b9dd9d173b213ce43c92163506f973a85155cc23c53c8c96b4b57

Observation df1fc234-6b03-4450-a8f2-74ceb248c922 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.412755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.412755Z digest=sha256:96d694d1833f16615c707e31c797b2554e49a94e3925adb985738b8586de6dfa

Observation 5eb22316-f96c-4534-9ccc-4ab52109222b · outbound

This paper cites Pissa: Principal singular values and singular vectors adaptation of large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Pissa: Principal singular values and singular vectors adaptation of large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.106941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:46.534245Z digest=sha256:ac14948d240832a02a81949e9b7bf6d695b251235e50c21e933d8b8573af3f48

Observation fb82432f-3dcd-4202-a6f1-cedd7489ad02 · outbound

This paper cites DeepSeek-V3 Technical Report.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.672703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.672703Z digest=sha256:6bd3932637cce2f64d397620e562e949bdc6141c8af13153c492dd1a4c558813

Observation 19d5527f-75f7-4917-b9f3-278883ebead8 · outbound

This paper cites Preference leakage: A contamination problem in llm-as-a-judge,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Preference leakage: A contamination problem in llm-as-a-judge,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.831139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.831139Z digest=sha256:964e2dfc04665dcc182a3d59acf9032109c6226ea8aebfeb52436e2ebe70b2e5

Observation 44ad428f-1c5c-4197-ba0d-3a659e98494a · outbound

This paper cites Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.843328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:46.961223Z digest=sha256:a0ec1cef66a87e644b204484341996c39b3238ba61abbdf5e3b88c41292682fb

Observation 97dc3270-991f-4aca-8444-5ffd1dab3267 · outbound

This paper cites Crystalbleu: precisely and efficiently mea- suring the similarity of code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Crystalbleu: precisely and efficiently mea- suring the similarity of code,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.498388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:47.156377Z digest=sha256:a1bbe8376eb8ee218970d1952c46224312f05d3cd6dc83d9ce84aba7f156842c

Observation a48577a0-d159-428b-897e-195126bdf97b · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:47.284456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:47.284456Z digest=sha256:8bb8bd9ed6989c61259ae797a1da84d6cee1579aed608aceb031362cc7134b23

Observation a87f8fb3-0a76-475d-95e9-3ec920ebc06c · outbound

This paper cites Codebertscore: Eval- uating code generation with pretrained models of code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codebertscore: Eval- uating code generation with pretrained models of code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.273794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:47.431493Z digest=sha256:d35bf919b58f5d8b0c10658df9242fcee571b8cd4e33faa8f4aba3bf93fa0bce

Observation a1f8e276-a258-4a87-8e3a-99cd99b7399d · outbound

This paper cites Codescore: Eval- uating code generation by learning code execution,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codescore: Eval- uating code generation by learning code execution,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.904879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:47.555464Z digest=sha256:95fd82ad253303b300110a91f916a95d2d7611864ca7c0d9d2aedbf297c35779

Observation b97d9c13-1d3a-4464-9b9d-2d0527eba567 · outbound

This paper cites CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:50.455945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:47.689046Z digest=sha256:346a7629a6af09167a66dcd0b3e164a7e320cae10ab941a49962004d3bae0920

Observation 2aaf3ad2-026c-4451-b6ed-2ba896733148 · outbound

This paper cites Ice-score: Instructing large language models to evaluate code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Ice-score: Instructing large language models to evaluate code,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.590439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:47.812514Z digest=sha256:ae32f76be764b43c19cd33e50425236bdfb0d62b7eb6669451d54035418eec15

Observation 831fc490-587b-43e4-9ad2-14fa7a542494 · outbound

This paper cites Codejudge: Evaluating code generation with large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codejudge: Evaluating code generation with large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.336829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:47.980501Z digest=sha256:5c464bc3d92409112ffbce038e52cf69f696ba02fd9afabe66209d86e53d5d0b

Observation d1b05140-ba45-4990-8fa0-9b199a50692d · outbound

This paper cites Benchmarks and metrics for evalua- tions of code generation: A critical review,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Benchmarks and metrics for evalua- tions of code generation: A critical review,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.066356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:48.119267Z digest=sha256:1fbe5c7a3bcfcb252f206501bcf69adf5cd6210edf781a2699c8484fbe828c28

Observation 20add06a-1154-416b-8b05-f3a185bb4f3b · outbound

This paper cites From Code to Courtroom: LLMs as the New Software Judges.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From Code to Courtroom: LLMs as the New Software Judges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.242132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.242132Z digest=sha256:84bb6fc9ed17ff6bfc4421698af20e602fc8bb55ca46fa842b7a76ead894512a

Observation 42377c96-b83f-44c4-8a1d-359778743253 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:51.761783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:18:48.442996Z digest=sha256:9c9bd7d0def7cb8397a529f71453adb2856c9ac8d5e64b3dbec0c2d53a53e394

Observation c35ee74f-45a5-416f-9a9e-a9027fa1014a · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.551215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.551215Z digest=sha256:8267f1110889186765547e8a28f0374b653e194101381aedca77b55ec7411308

Observation d5e8fd5b-c57c-445c-88d3-ed7d2cb88b75 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Qwen2.5-Coder Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.666706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.666706Z digest=sha256:207977a98a739d6efdc44afecfe10009c27f95b34b5e9d791d760404e7c5c7e9

Observation 5b65b982-8887-435a-ac28-b5f595e03134 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.828289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.828289Z digest=sha256:347f1923da252d9296ae00f8e540dc13ea93307e61b7bde2024fd193ea08faa7

Observation d316dd0b-02f7-4b68-99aa-12cdd7cda6d6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain-of-thought prompting elicits reasoning in large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.977432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.977432Z digest=sha256:ce9df16a45203850e6827c16895b561483cf1fc08b9df3a166e600d756b392a9

Observation 8a620128-32a8-4efc-9f36-9fb7a9a7ba48 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient memory management for large language model serving with pagedattention,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.169400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.169400Z digest=sha256:9cd1b1d97463bba817cb26079b8e5a11bd2f13767dc4e43c695b70553b8c8d1d

Observation 3a2a8f9d-ef7c-44b5-8dbe-1e0898697da3 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.297876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.297876Z digest=sha256:f8414dd3827bb1002fe09679d88590c745ba68c57d97056a26222abce36ca7d0

Observation d4780c4b-4d30-4d97-8fbc-306e2706f1f9 · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.422123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.422123Z digest=sha256:cb4eb8ac6591aa13eecd8c085f00cb9c24ec39818aa956f157cbc8b8065386ee

Observation f38a23b2-d917-42f4-92ce-7d3fdfb8d31c · outbound

This paper cites Less is More: Towards Green Code Large Language Models via Unified Structural Pruning.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Less is More: Towards Green Code Large Language Models via Unified Structural Pruning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.570812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.570812Z digest=sha256:33f5e333b2ed9781fa7898847ae4524899efe8a20772becc7670ebb703d877df

Observation 4877cb16-649c-4715-ab71-f88e482f3690 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.765210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.765210Z digest=sha256:653b5e7941b46ac8f5191af3a462a52f2332cf73c1bebccc9cf73e8b25439d25

Observation 6d7b34fb-a8b3-4a53-b612-d2274275d5a2 · outbound

This paper cites A coefficient of agreement for nominal scales,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A coefficient of agreement for nominal scales,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.895278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.895278Z digest=sha256:977ea8ef18b141280e71722cd5f2edfa210d9c85deaec470a50b6d3fbb2cff47

Pith citing papers

Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:32:34.645100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.789468Z digest=sha256:a5b26dfb0cfa2cec45b8e99a794693a76e9dd2664f89644a0a34f35e57fe3da6