Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2403.02839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.02839 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:53:28.643250Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 440556cc-00a4-4867-a40e-587797f21bc3 · inbound

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models cites this paper.

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:32:17.998336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T23:32:17.279836Z digest=sha256:35ce833aa0144cd383ff4007f059fa178912e41ac13907d0ad74fb999c7fbdad

Observation 619b6d6e-f015-48cd-ab29-8af910c92656 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:44:49.795483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:dc5fdd0431a184a0d3979f3bc25cfe8fa15c354ecddf4dc5d844ec49ec60384d

Observation 8fba2e7e-e1bc-4ebc-b016-358942c9def7 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.974779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:e71d0ebc3a11af62e6ef52ecbc527112081a3255073359a0f7affc8fbf6b7618

Observation 6ca8af0d-7ab7-4a8b-97ac-91e9263a8cb8 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.501674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:1c949a41ec9053a49b6f14b0382aeffa9a68f95838450c2b57739b67d9c784cf

Observation 6b368de2-740e-4e4d-86c4-5ba9d9cdd820 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:43.845001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:275d170ddfa7a8233ef9520eb01e9f636ed561ecd3d7f17ae9e02cdcc3801eb2

Observation 33615456-3847-4488-b199-59963e60b2e0 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.802863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:5c0b4efa0df778e0838b7cd7286e682a72ff3d057028a46798c1639aeeeed4e2

Observation 9b67cbc0-95b8-43ef-86fd-5d34564de8ec · inbound

Combining Large Language Models with Static Analyzers for Code Review Generation cites this paper.

Combining Large Language Models with Static Analyzers for Code Review Generation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:53:28.643250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:53:28.643250Z digest=sha256:19e5b7977c3ecf525f499ea548fea5ec8ec965e2a81fba058603cb7d89964e2e

Observation eddbdcaf-db0e-4888-8247-92ef29f2589b · inbound

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression cites this paper.

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.118892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.118892Z digest=sha256:0950bd89969e2d485a0a0fc46cdc9628e23c91ce719cbcbeca97e4c607962824

Observation 338cffb2-af94-4e6e-9237-ea82b87c11ee · inbound

VLM@school -- Evaluation of AI image understanding on German middle school knowledge cites this paper.

VLM@school -- Evaluation of AI image understanding on German middle school knowledge An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:06:32.364901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:06:32.364901Z digest=sha256:7959735a3afb8152d0db9de28ebc60846420e284a0144160c99598f59f786380

Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.829032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.829032Z digest=sha256:3400ebd4343d75529d7abb93ac8b1dce333ea3bed6df0def992790f0eb02737a

Observation 6630e635-cbe4-40e3-8524-e2f4199fe69d · inbound

On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization cites this paper.

On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:10:23.222553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:10:23.222553Z digest=sha256:967bc6abdfef58045eb2d1c1e3ff2f293757e3230c8c667ee1f4198df982b7c8

Observation f9e7aece-e3a1-4f62-a2cb-de4e76c4482a · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:56.912415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:56.912415Z digest=sha256:bafd0064c9c74bdfa2d5e2cf6d2d1d744004e50bcfce41608a831190b26c83de

Observation 094d9722-3e48-41b7-b780-5842449971b8 · inbound

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation cites this paper.

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:20.133943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:20.133943Z digest=sha256:dc6957ee4ed6286755536b464ac643890180dfda0dff8b9f580bdd531cd24db5

Observation 2333b4db-20bd-4583-a8f7-2d4feab28d01 · inbound

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning cites this paper.

U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.005139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:19:55.849451Z digest=sha256:c863586829f2e259181c0f7d3bad7cf69f7a1d96ec84e06cc4f01e53b303b906

Observation d2d22ba6-a5c3-4f42-a572-30e2337e13eb · inbound

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability cites this paper.

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:17:02.696947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T01:12:44.476894Z digest=sha256:29246c48dc6099a1e9d663241d7d22ce25ee1f29e0c20603268b14043a649143

Observation 0d67d5d2-beea-4717-858a-1893de695dc3 · inbound

Section-Weighted Hybrid Approach for Legal Case Retrieval cites this paper.

Section-Weighted Hybrid Approach for Legal Case Retrieval An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:56:39.300020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T08:39:28.510255Z digest=sha256:69c4522c90d2c23ed1ff7ae054761284a5d43decdbdc41f5468a32a1dce95d7d

Observation 9d426666-40ad-4b7c-9c99-16f0bd40c3e4 · inbound

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators cites this paper.

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-06-27T21:41:18.571899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T21:39:26.268337Z digest=sha256:ac375b790d355239f753deaa2d99be08bc1ae1d8e9626318b55fe16b4b60acaa

Observation bd79bc58-d3bd-4eb0-8bed-8afeccf7dde5 · inbound

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition cites this paper.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.492775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.492775Z digest=sha256:35a6b309feeb2ad076b15587edfa271fafe9c55b035c55d63e9e8b250451b2fd

Observation 919a65f8-29a8-430a-a1ec-2d0a3744c8f5 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.090862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.090862Z digest=sha256:6a6960fcdb630d29edea6e7469c6604998364beab986ae79c1da50ced02e79fc

Observation 27885a75-59e5-46f3-9d69-832c82e6d4cd · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 272

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:40.748915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:40.748915Z digest=sha256:5edd4d0c718e7da8d0661ebeb03d3283f126efe7f219d4878317f118387dd14f