Pith. sign in

Paper Citation Record · LEDGER

CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2408.10718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.10718 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:50:50.690353Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T17:35:44.171248Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5c19431-5940-4394-a6e7-465c85ba4efa · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.174084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:8bea5ae0053116cc3a82a2c9f82060b00424f08021e0df5a96047659baee1208

Observation 6d27ed5e-6df0-4916-8a91-08e3c4982c17 · inbound

Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge cites this paper.

Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:50:50.690353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:50:50.690353Z digest=sha256:215d03b01cd2e60bb3f55ea376d8a6dc1e7dd91c2c209df8afae2cb5b9a3291d

Observation c59a42ba-7e1d-4028-9e5f-267ad9b6eeb7 · inbound

Is Your Automated Software Engineer Trustworthy? cites this paper.

Is Your Automated Software Engineer Trustworthy? CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:06:53.978307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:06:53.978307Z digest=sha256:034a713b13c2cd499aca28fd203b85817dc0f2a77430aeb098ae377f8430bdfa

Observation 1db605c3-6c3c-4b88-be35-b78c37460583 · inbound

ReCatcher: Towards LLMs Regression Testing for Code Generation cites this paper.

ReCatcher: Towards LLMs Regression Testing for Code Generation CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:57:05.981261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:57:05.981261Z digest=sha256:7056207e753014c358edf7992d53c154faf733f98061ffc1a4c0e822d1ec8f29

Observation 195b9ce0-c22c-4dfb-a4ab-6043b7c4c515 · inbound

Enhancing Large Language Models with Retrieval Augmented Generation for Software Testing and Inspection Automation cites this paper.

Enhancing Large Language Models with Retrieval Augmented Generation for Software Testing and Inspection Automation CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:38.021334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T10:40:30.734804Z digest=sha256:e02d6034894d30109ea418d266d1938adb332aa3d42b7befee8803b32fae0d34

Observation 8d28c040-a207-4ebe-b0be-c058de436af6 · inbound

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering cites this paper.

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:32:00.279793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T07:29:03.994957Z digest=sha256:cc92ea1e54382e3e652e2a1373d03b5d3e6d98386512e5325c8e55f3615eede7

Observation 8b33dec6-9523-40c2-b587-d803d99e5a26 · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.012242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:9b529e99543375fea3c8cfac069489c4b18d70b3c6b77b77633951e8b666f3d9