Pith. sign in

Paper Citation Record · LEDGER

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2505.19502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19502 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:18:49.895278Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:33.789468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:32:34.591005Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8c3daba-9935-4ca6-9220-153383d74bfc · outbound

This paper cites Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Llm-based multi-agent systems for software engineering: Literature review, vision and the road ahead,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.793868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.793868Z digest=sha256:4df47cd4294e3aadebd4089fa8881abde25d9b7ad51e47c64002a8caa43ab29d

Observation 37ac8ab1-e125-492e-a656-74e769683d7f · outbound

This paper cites Efficient and green large language models for software engineering: Vision and the road ahead,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient and green large language models for software engineering: Vision and the road ahead,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.852067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.852067Z digest=sha256:fac32f72a635743fddb390129d8a42b5e48d49b86e6f5cf9eb72faea5c854037

Observation 7059d352-e049-4c97-b5ff-c4b3d8c384a3 · outbound

This paper cites Large language models for software engi- neering: A systematic literature review,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Large language models for software engi- neering: A systematic literature review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.977776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.977776Z digest=sha256:b8672abfa9cbeaafc0a56c53b593219d61eb41bd7c77c40971b66fcb2d09a9ef

Observation d5db457a-4915-4066-a1b2-91a769ee34f4 · outbound

This paper cites Exploring the capabilities of llms for code change related tasks,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploring the capabilities of llms for code change related tasks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:56.142825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:44.075982Z digest=sha256:9ed19c1900ef3e45cf3c94ae736f4efb9d105e05f81f7b1745f432334e6aa655

Observation 436ddac7-8fc6-453a-bd73-f5cb6ff0165b · outbound

This paper cites Chain- of-thought in neural code generation: From and for lightweight language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain- of-thought in neural code generation: From and for lightweight language models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.190067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.190067Z digest=sha256:95c3ef4e253e8a194840c17ee1dfeea2b29d9224a716f5f84e31e1918c54a2f2

Observation edc0ac39-d979-410b-899e-af50a5ee8016 · outbound

This paper cites An empirical study of retrieval-augmented code generation: Challenges and opportunities,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation An empirical study of retrieval-augmented code generation: Challenges and opportunities,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.323662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.323662Z digest=sha256:808202a888b40b16bf5119727b14265623d1d3f895ef1ef044f7dd704e260cf3

Observation b8e62aa2-67fc-4b4c-8ff7-ee8e4befcdde · outbound

This paper cites A review on code generation with llms: Application and evaluation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A review on code generation with llms: Application and evaluation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.419952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.419952Z digest=sha256:10b215a84efa39adcb7d11b3dd34361aa0064e7d23ca5f88609796f257fbad9f

Observation a016940c-4a33-4e26-af3d-37407ca86c7d · outbound

This paper cites HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.535532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.535532Z digest=sha256:0b8631051fbee60a46c7d3484285d84545c3ff584442678cc6badfe8457df14e

Observation da753d1c-7523-4255-8b2f-db694ff4245c · outbound

This paper cites Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Assessing and improving syntactic adversarial robustness of pre-trained models for code translation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.722335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:44.695755Z digest=sha256:e45e2859766228ffffaf44e6f63204ae1e472c92ffcc35c2242721b23b0ad477

Observation 6f840e5e-74ea-4099-b2ff-15a0f23c8e5e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Bleu: a method for automatic evaluation of machine translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.800778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.800778Z digest=sha256:154ff198077e5f00872474a2e613d8c5dbbff6039a030d328f438f67b22fd2c3

Observation be427ef5-801d-43d2-b29f-3dd2615fcca8 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Rouge: A package for automatic evaluation of summaries,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:44.882525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:44.882525Z digest=sha256:40f31498af0ca71c940d3493079b57590c2c0863832f615938c0c3817197c7d8

Observation 61f6370f-1789-4085-b2ff-43d2915ce993 · outbound

This paper cites chrf: character n-gram f-score for automatic mt evalu- ation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation chrf: character n-gram f-score for automatic mt evalu- ation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.399227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:45.020391Z digest=sha256:0c239ccb090a8546dad05080422f99b4508c1ced59bc26b61f9afdf08569efb7

Observation 788e8390-42a0-43fb-bbc9-d3f73c7c27f7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.176339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.176339Z digest=sha256:6dadc4ca2e84a93e02cb21af1622973f5e78fad660571c132b22aca70ae5a4d0

Observation 7bd5d4a7-0764-482d-93dc-49dec8202507 · outbound

This paper cites Exploit- gen: Template-augmented exploit code generation based on codebert,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Exploit- gen: Template-augmented exploit code generation based on codebert,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:55.097618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:45.322765Z digest=sha256:7004a220ef50f750612c183db5a8b27b8b21e6b1ba1d486a64c722443b185425

Observation 328742c4-9d6d-4754-9765-2a5d317b0ba1 · outbound

This paper cites Are nlp metrics suitable for evaluating gen- erated code?.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Are nlp metrics suitable for evaluating gen- erated code?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.811838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:45.485559Z digest=sha256:919f637a9dfbc2e6bee83dbf176c73d55ed147fa220b26f5b0f9f774aa8312cc

Observation af2248f9-f5c9-4c45-84c7-7ea2822b126d · outbound

This paper cites On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:51.337928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:45.617534Z digest=sha256:90fd143da74d8621f6849e1a2b1597bc1805c949df01397c3aa173e8da993d9a

Observation 7f780868-1084-45f3-a25b-b9518ed9b730 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.728953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.728953Z digest=sha256:e6e0c6e3c27ef79cbe645eb8e5123e11b545132b39494e4220e4b5f75f32fa39

Observation 606da85a-ab5e-4433-87a0-1a1720486b5f · outbound

This paper cites A Survey on LLM-as-a-Judge.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A Survey on LLM-as-a-Judge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.866412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.866412Z digest=sha256:e0ba2e58acf7a31c18b12d9f33cc99c9af345c2df54986df393b0ff720e01af3

Observation 6591cce8-2b4c-45b4-92b5-db04020bcf7d · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From generation to judg- ment: Opportunities and challenges of llm-as-a-judge,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:45.986897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:45.986897Z digest=sha256:05181a0aed68116ec8a82bcb5873c9164347ee47e9fa61485768cd4915f8324f

Observation 954b232b-51e2-48e3-9a83-80db1bc437ab · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.158301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.158301Z digest=sha256:02b9b2da0bdd43885737ab365a51ea8737f7abcaa879b8376d01a77dce6b5ecc

Observation 2dc63c72-91a8-42d9-ac61-f8f830f53b5b · outbound

This paper cites Fight fire with fire: How much can we trust chatgpt on source code-related tasks?.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Fight fire with fire: How much can we trust chatgpt on source code-related tasks?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.479297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:46.267717Z digest=sha256:e225264497d8a9396f4e719fb5b9952d9eed1f0be667b6029bd963e73c0aba9d

Observation df1fc234-6b03-4450-a8f2-74ceb248c922 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.412755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.412755Z digest=sha256:b260a53505fcc0ada067f3508aa6dba13062333e75143bb1132cfa5227cfd8a3

Observation 5eb22316-f96c-4534-9ccc-4ab52109222b · outbound

This paper cites Pissa: Principal singular values and singular vectors adaptation of large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Pissa: Principal singular values and singular vectors adaptation of large language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:54.106941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:46.534245Z digest=sha256:b242d0852c2ede9df625a7b1eb3573011c45abb454dd9fc1bce746b294caf4b9

Observation fb82432f-3dcd-4202-a6f1-cedd7489ad02 · outbound

This paper cites DeepSeek-V3 Technical Report.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.672703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.672703Z digest=sha256:c5d8340444852a4d5f5f226807575ee574c43deb2638c6a5d464ec375a9926d8

Observation 19d5527f-75f7-4917-b9f3-278883ebead8 · outbound

This paper cites Preference leakage: A contamination problem in llm-as-a-judge,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Preference leakage: A contamination problem in llm-as-a-judge,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:46.831139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:46.831139Z digest=sha256:164649328d29e1893822a590b33bad4f9c77475bae002ba749c0f2659d39905c

Observation 44ad428f-1c5c-4197-ba0d-3a659e98494a · outbound

This paper cites Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Who evaluates the evaluators? on automatic metrics for assessing ai-based offensive code generators,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.843328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:46.961223Z digest=sha256:b309db96c3366eb778ab6d058be480a3a37ede318b27e623301e3615f4db3d46

Observation 97dc3270-991f-4aca-8444-5ffd1dab3267 · outbound

This paper cites Crystalbleu: precisely and efficiently mea- suring the similarity of code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Crystalbleu: precisely and efficiently mea- suring the similarity of code,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.498388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:47.156377Z digest=sha256:3fcd78fda15c8cf032d1cd1f1b502ee0cfae2a5074a052c0a9e1ec4f3c38d047

Observation a48577a0-d159-428b-897e-195126bdf97b · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:47.284456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:47.284456Z digest=sha256:26e2199fe069d752f6394143cb9685ec9064dd4c481c7b3e94c617f58b504ada

Observation a87f8fb3-0a76-475d-95e9-3ec920ebc06c · outbound

This paper cites Codebertscore: Eval- uating code generation with pretrained models of code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codebertscore: Eval- uating code generation with pretrained models of code,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:53.273794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:47.431493Z digest=sha256:3a6c7f52388c5111388d6e92723296adc44534d5b4df6a3bd91b9a1cc43b6eb9

Observation a1f8e276-a258-4a87-8e3a-99cd99b7399d · outbound

This paper cites Codescore: Eval- uating code generation by learning code execution,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codescore: Eval- uating code generation by learning code execution,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.904879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:47.555464Z digest=sha256:2a365b054c59f4e168b607f276c3847feaa8edf511c7f6ce65f678ce1862b02c

Observation b97d9c13-1d3a-4464-9b9d-2d0527eba567 · outbound

This paper cites CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:18:50.455945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:47.689046Z digest=sha256:514c558a093f71ecb37927198b43d617f075b3a012b243b086d9d390f8b43e80

Observation 2aaf3ad2-026c-4451-b6ed-2ba896733148 · outbound

This paper cites Ice-score: Instructing large language models to evaluate code,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Ice-score: Instructing large language models to evaluate code,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.590439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:47.812514Z digest=sha256:2cbfebaa81d35f4f9b3d18b91d4dcc44a7d79515e757e3fa9a5aba7941d6b3e3

Observation 831fc490-587b-43e4-9ad2-14fa7a542494 · outbound

This paper cites Codejudge: Evaluating code generation with large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Codejudge: Evaluating code generation with large language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.336829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:47.980501Z digest=sha256:f4517d00a30b9e5feb7d46f8513fd47d3accd1cbb4afdf9ab46bfcfac8a4d33b

Observation d1b05140-ba45-4990-8fa0-9b199a50692d · outbound

This paper cites Benchmarks and metrics for evalua- tions of code generation: A critical review,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Benchmarks and metrics for evalua- tions of code generation: A critical review,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:52.066356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:48.119267Z digest=sha256:eeb85412affecf2b8c2ad6d93cdc2d1a80a87256651306ca666a5a503d669604

Observation 20add06a-1154-416b-8b05-f3a185bb4f3b · outbound

This paper cites From Code to Courtroom: LLMs as the New Software Judges.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation From Code to Courtroom: LLMs as the New Software Judges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.242132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.242132Z digest=sha256:b88c6001d4c496955178aa25c3fe2092ea079afc1520a2574114fcec10d42abc

Observation 42377c96-b83f-44c4-8a1d-359778743253 · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:18:51.761783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:18:48.442996Z digest=sha256:97dca5034f2c241cc74773ad93cfd4c4d248093ae9c465a8f64b7d803b035b73

Observation c35ee74f-45a5-416f-9a9e-a9027fa1014a · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.551215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.551215Z digest=sha256:a61e4ceadc131613c1c0e52c899ba1cde1ad6166bda94459d40e0384198cca73

Observation d5e8fd5b-c57c-445c-88d3-ed7d2cb88b75 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Qwen2.5-Coder Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.666706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.666706Z digest=sha256:f87ebe8cc666d7b8671bc6ec33e28dad38ebb515c8709b28f6123eb9409a6ee8

Observation 5b65b982-8887-435a-ac28-b5f595e03134 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.828289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.828289Z digest=sha256:91300a6c66376c3151d21853e5a274ed950b6dc1b0417a79b954f00cca5ea8c1

Observation d316dd0b-02f7-4b68-99aa-12cdd7cda6d6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Chain-of-thought prompting elicits reasoning in large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:48.977432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:48.977432Z digest=sha256:d34269bd4f429d05d0128627310b605c2590f4c35e2519584a6a2e6cb45ac7d8

Observation 8a620128-32a8-4efc-9f36-9fb7a9a7ba48 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Efficient memory management for large language model serving with pagedattention,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.169400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.169400Z digest=sha256:45c8ea9c2703ff1d0ce30ce1327abb405384fef9660e58f0e0fcb42ac58986fa

Observation 3a2a8f9d-ef7c-44b5-8dbe-1e0898697da3 · outbound

This paper cites KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.297876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.297876Z digest=sha256:f92fef3501dfc7b925c1fa317c3fd6cf8d92bf3916320a7913589a26c2ab1cd8

Observation d4780c4b-4d30-4d97-8fbc-306e2706f1f9 · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.422123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.422123Z digest=sha256:296abe80c9a486c0a99b7e2ee05bdc19c1bf12756c5bac5258f68ff8b8ead4ae

Observation f38a23b2-d917-42f4-92ce-7d3fdfb8d31c · outbound

This paper cites Less is More: Towards Green Code Large Language Models via Unified Structural Pruning.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Less is More: Towards Green Code Large Language Models via Unified Structural Pruning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.570812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.570812Z digest=sha256:85260fe17c3922c9bfbcaeee94672bcfe16220b8fa76973571659ec845af4d7e

Observation 4877cb16-649c-4715-ab71-f88e482f3690 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.765210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.765210Z digest=sha256:f52db7167051a9d5607d9d35274abdcd8422dc96955b8201e398c73b93e8feb1

Observation 6d7b34fb-a8b3-4a53-b612-d2274275d5a2 · outbound

This paper cites A coefficient of agreement for nominal scales,.

CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation A coefficient of agreement for nominal scales,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:49.895278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:49.895278Z digest=sha256:f0a25feab669fecc870ac6cf9d2e2c691a41b9358625bcc5c9b5918b10de7bbe

Pith citing papers

Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:32:34.645100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:32:33.789468Z digest=sha256:03c569b215267744df3d6e110a60ef8f8e88fb7399f58f97503559670dc8876f