Pith. sign in

Paper Citation Record · LEDGER

Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2408.13006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.13006 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:24:54.518422Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 02d0b65c-f5f1-4087-b51f-22bad0f786a2 · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.413886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.413886Z digest=sha256:b0645acdf2a3d452cb704d1c46d80034139f2602b4b924a1170ed8fc5da11dbb

Observation 93aad358-4b48-40d0-b804-ff1f51b37eb0 · inbound

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator cites this paper.

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:15:45.007500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:15:45.007500Z digest=sha256:a567e5d4bbbfecdb0664650ee578b26acec7624ac0634cbe3aba8088adea14c3

Observation 03aa5361-1225-4f49-870b-270152bb2229 · inbound

JuStRank: Benchmarking LLM Judges for System Ranking cites this paper.

JuStRank: Benchmarking LLM Judges for System Ranking Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:49.630019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:49.630019Z digest=sha256:5134d0cb3b7ae05232af99414960e878e408e83eb32536ed4737dc528efaebf0

Observation 717fec84-1dc1-4046-a27b-d65922246322 · inbound

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge cites this paper.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.401714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.401714Z digest=sha256:7cfa1e4763036d6e673f66663a06f00f621d72419b6ee4ba828bde8d15ff6716

Observation 64df046f-feb9-42f9-8849-fabc483a1f95 · inbound

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems cites this paper.

Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-09T19:10:53.900627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:10:53.900627Z digest=sha256:dddab6de6dd4302915ba020170a61ec67a3b0f354a0f7f81265ce34fa985b9a7

Observation 048f9501-3b4e-46a5-af7c-b65f2aed6af8 · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.549944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.549944Z digest=sha256:7483169238281e409ae356ddf598b470e2e71771bea2fc835a3a23e49c5bda66

Observation c736d4fb-3f19-43d8-bef9-b167f1e36ba6 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.404608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:f4a5ffd123acd76e49e59cfb9ed25640ff701eea1685a41ca315cbdf3a822718

Observation f589180a-ebc3-4b9c-bb6e-911a83749ed3 · inbound

Benchmarking Multi-National Value Alignment for Large Language Models cites this paper.

Benchmarking Multi-National Value Alignment for Large Language Models Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T12:24:54.518422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:24:54.518422Z digest=sha256:071c50142ee4215bb2a1850eb31f184ad3b8184b31947be5f756ed39425c6d85

Observation ea756bce-79bd-4c0a-826f-358842298d06 · inbound

Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs cites this paper.

Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:38:28.503121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:38:28.503121Z digest=sha256:2d02348cf99b29bde4ab7f0561b67f4ddaf6789fc90755d6e43ae1c1f441c57f

Observation 05e4fd5f-8b6c-4d20-b406-07454b960542 · inbound

EnronQA: Towards Personalized RAG over Private Documents cites this paper.

EnronQA: Towards Personalized RAG over Private Documents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:30.333086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:30.333086Z digest=sha256:5d5f2c5f3b4c9c202efe38ceaddaf66d85dc73623db335632a674323e2f7e271

Observation 57ba14f9-bb84-44fb-b442-9d6f7946c53a · inbound

Large Language Models for Predictive Analysis: How Far Are They? cites this paper.

Large Language Models for Predictive Analysis: How Far Are They? Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:25.799258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:05:25.799258Z digest=sha256:fa63c07d574cab24d508fcbea4352e03c51134669015281d02447adde1f98cb2

Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.693251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.693251Z digest=sha256:dc162af360874ec8b5bee11fd31d6328481f7f08847d2a67558c109096a4120a

Observation 7c9c9731-28a4-45d0-914a-91332507933e · inbound

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance cites this paper.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.386305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.386305Z digest=sha256:6e1000f1bb193bd3ec150a89624d4bcb6f1464e420656acead7f16a3fa70fafe

Observation ef315790-9b74-4f37-80c1-6a150431efaa · inbound

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation cites this paper.

Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.584381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:34:08.584381Z digest=sha256:e8902d1a143d0d5e3bfd6105bd4f89aaffd6c59e09424043d6d8dd397c0ddac6

Observation 60835722-30b9-4029-95ec-185286065f5e · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.601561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.601561Z digest=sha256:72d5366a752bbd3379ae2bbf585b08187f3a3634411378b9105a127091cbd75c

Observation f32569ca-49f8-4a11-bbe0-c12970b98b9b · inbound

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making cites this paper.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.934638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.934638Z digest=sha256:1172b56d79e5bcb3cc84b63e0655a5317d8e0e0cabd0bab2f7262e7d7544b690

Observation da944267-1197-4876-a579-307bef239571 · inbound

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing cites this paper.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.843092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.843092Z digest=sha256:19815e019513c6809b4136df5655c90097edd9a251700ef073852e17d65ddc16

Observation 31fef8a8-1a7f-49d7-b3b8-a8bf0d604315 · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:05.596485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:05.596485Z digest=sha256:73bda4064868d9a95706fd57675c5567b2a313e09ae22b2daab2e8b55e9a010d

Observation 1dabb476-796a-40e7-a1a1-95629d87f9c2 · inbound

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs cites this paper.

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:14.617849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:14.617849Z digest=sha256:129139b5aff4a0b168500573a6243bc112f0cb92828aa87bd65ecf9ef02b596a

Observation dd0bb244-766b-40a7-bb26-8341a1390907 · inbound

AI Propaganda factories with language models cites this paper.

AI Propaganda factories with language models Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:21:32.643649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:21:32.643649Z digest=sha256:2fd3752f4a983eda58a673e2d0bf142e4f014b3c9977f6d27b1b1769ee7718fb

Observation 7c465ca9-2c19-4f63-a87d-aee427466f19 · inbound

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations cites this paper.

IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T11:26:27.543864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:26:27.543864Z digest=sha256:299d2c31a04683e17d51a05888c79b9e11e34b7165b2436f7c3be468b3959a65

Observation aeb117fe-1c82-4dba-8867-5cee665d83b0 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:7025e62f792ae874f35f7686774cede6e04e63e474ccda81e957f45d3fdd62d4

Observation d49d3437-8437-407b-8c8d-2b048b58bf33 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.168845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.168845Z digest=sha256:2c0fc9e22cc3657e01d3acc17412675a3570b561344d07f7a7fc9ac77d8af68a

Observation d1d85efb-00b7-4abd-8b08-b003eaafd557 · inbound

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection cites this paper.

LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:45:48.017272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:37:19.360378Z digest=sha256:3cd0aa6d016543c1aae1ece94f69c34f09aed4f723410e5fec33588033f60d70

Observation cfcb7436-1bf7-4185-b89a-df84914418af · inbound

Iterative Finetuning is Mostly Idempotent cites this paper.

Iterative Finetuning is Mostly Idempotent Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:44.631356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T19:04:36.210200Z digest=sha256:f8900e365cacaea89724e5b15f5c26642f369d75bc4a82576b3ab7fc7af09e5d

Observation f2209723-09be-42a7-9623-9f00e582529d · inbound

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration cites this paper.

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:40.790933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T13:38:51.424919Z digest=sha256:786a8704e8b002d66cf7b7788b6a74750899cc4829bd7ef8def6649d4b9a4c9d

Observation 7133626b-ec70-40b2-b733-4667d82b46c1 · inbound

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators cites this paper.

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-06-27T21:41:18.486232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T21:39:26.268337Z digest=sha256:2429e1c1ff8014feb1a4f73c61e7657e3d91aa4029ab6005b63a81d4f4d8698f

Observation 8ff3a2b6-4a2a-453b-8bed-405c8d29267d · inbound

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment cites this paper.

Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T14:02:41.517301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:02:41.517301Z digest=sha256:be08391c2b491815d1340119e099e5244c287c98b96a467c682b7b2081cb82c8