Pith. sign in

Paper Citation Record · LEDGER

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 8 inbound Pith citation observations for arXiv:2504.17087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17087 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:53:25.065089Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:20:29.355664Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.227079Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0dd83741-44ae-4dca-ad01-b80f8ac6a110 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.707459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.707459Z digest=sha256:03b38412bdb0f8f08da703f95f2dc9da81c2bae7aaaee25df17de86b27defeb3

Observation 8abf0c6f-6050-487a-8758-2119a025dd38 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.840206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.840206Z digest=sha256:13eadefbc8c71fe79d4ae591ebec16d5956adea6d318bc5723be1b7986e753f3

Observation 8c1e6529-6777-4d0a-aee8-b7d7d03944d9 · outbound

This paper cites MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.848112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.848112Z digest=sha256:345badecccf18d1aea4b19e1140672bd13bd39385f9a9488127583bd0de39454

Observation 08bb1446-c4a7-4589-a742-fe0c0c382d0f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.860155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.860155Z digest=sha256:525436cea073be52c4c58b36d5ea68bb851a98acd4d15ac5f3b1906a54939a31

Observation d96debcb-a0ab-4ee9-ac07-e64c686188cc · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.867753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.867753Z digest=sha256:b7fd22e300333ceb597e2f86abb7885b3dde4e2af1723c22bc8f51dafa9b8ef7

Observation 164e792e-d0f6-48b5-9097-56408af56e71 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.871674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.871674Z digest=sha256:e42b783f8bd68edb7af4fe7a3ea1637111d7baeb0f540cd856dc1c09de7c461f

Observation 7609c747-f80d-423a-9148-12fccf3a474b · outbound

This paper cites Self-rationalization improves LLM as a fine-grained judge.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Self-rationalization improves LLM as a fine-grained judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.875340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.875340Z digest=sha256:df18fc4370807fbf1c1337f4f3acea802deb8edbe739f10c793740a28f6dace8

Observation 9f34a9d7-520f-44d4-9e48-78738a59461b · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.879314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.879314Z digest=sha256:aa22ffcfcda75d88c4cb0dc23f54f0090d7ad32aeb5c0f73d8671b437534a531

Observation 47dca15f-0f4b-40ed-8137-76226752cf07 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Aligning Large Language Models with Human: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.919360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.919360Z digest=sha256:a6f69de0d3aabc4a907e06ba584f2962f0da88e9cacd62e97b43f41e841c2ffc

Observation 0e07aea5-8e62-4d70-b3e6-3f623bb50998 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.990467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.990467Z digest=sha256:9a713c9dc562887bd9c6f522b93ab11edc9eb1cbc1343da3de4e6affd415f4e4

Observation 246fa49d-6c82-4101-a111-b0abb1c1e14f · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.046352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.046352Z digest=sha256:e1dc62c1f360733ff9c379360b4009c056f91991f132f87650cbb2a01d055ffe

Observation 852a8c22-4244-4b4e-b17f-86327ee6750c · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.049609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.049609Z digest=sha256:767c682b0500562457f8df728d954b27551bffbc1d15d2dabf6973c6002089b1

Observation 181b3b98-32bd-46b0-b91e-2b6f4e16afd7 · outbound

This paper cites Do Large Language Models Know What They Don't Know?.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Do Large Language Models Know What They Don't Know?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.053262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.053262Z digest=sha256:e2d2f929510271926aa40f5f8e63709d1473cb519ee76345b40d14f5747bac41

Observation 2e0b0f99-3c9a-46e2-b141-122e026017c2 · outbound

This paper cites MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:53:25.139148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:53:25.056988Z digest=sha256:769c090651db92bca3a7d995bfa8f7098cf25bafb51e1af52c3963bdd0038e82

Observation 06e17a48-7e15-4e51-9542-c596ecc4391e · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments BERTScore: Evaluating Text Generation with BERT

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:25.060888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:25.060888Z digest=sha256:49c4be1d136b5fde5427f8a24e6f6c3513ef8eec8ce1ceb13ded56f9ca41b0f0

Observation 277ec298-e4d6-44ca-be0b-cf43c770d623 · outbound

This paper cites Detail of different rubrics In the single-agent precision comparison section, we analyzed the impact of four different rubric configurations.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Detail of different rubrics In the single-agent precision comparison section, we analyzed the impact of four different rubric configurations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:53:25.370128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:53:25.065089Z digest=sha256:45cc822e221790de49932dade1898c2d800e9474d907d4d2b6884845d72dfd42

Observation b2e1119a-76ad-4be0-9649-d4222655c614 · outbound

This paper cites FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.855417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.855417Z digest=sha256:20b075a344abc53656e7987311c6b80966a965819a0d84d112d0b1e638c84ab4

Observation bd9b1b95-b754-403b-b922-e5608079069c · outbound

This paper cites MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in Summarization.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in Summarization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.852244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.852244Z digest=sha256:03f5d7438de2bc1d7f365d31fe02771171054763952e613801d18bf93c0f2c9f

Observation 475da4c9-edf1-407a-b872-92fe5788823d · outbound

This paper cites Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.863952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.863952Z digest=sha256:750261848d62eef6120cfe6156c879b987cf008dcc2cccb682a2a7421cf1ade6

Observation 320b609c-1579-4ee3-afd3-ad55c33d9230 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Evaluating Large Language Models: A Comprehensive Survey

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.776501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.776501Z digest=sha256:9db7cf908d23746677b671af8eb348dba6331b9acc964849f064299db90091e5

Observation 8ce2fbba-1898-4282-9b13-4a7dff31c2ac · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:53:24.844122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:53:24.844122Z digest=sha256:9622b7272fb9cf9197f5e5d2795a34c46f1fa2852cd0dbd0d714ec6696258e23

Pith citing papers

Observation b82cfc1d-e4e6-4d7d-b402-abcc60a7d50d · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.327951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:eaafc0ffb138b38310a86c173bb855b6a779629bc38235fcf242bd08a0d53f3a

Observation 06775abf-bd4c-4d41-9862-16567315a85c · inbound

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges cites this paper.

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:25.789277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:27:21.060953Z digest=sha256:1d294223dd122791268347793f2771a6004adcf577aded286095dc3e0f57a68f

Observation 5ad57e2c-14b7-4923-bd6c-9e68778c3c2f · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.440905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:758c062a46f797816f0e696162d8892c5dce75d8a4b488f845f2b957d61c6910

Observation 09773501-43c3-4a0a-a3b3-399f1e2445f3 · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:49:38.229229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:7b47bcab50b4822407e026119c67e3502d754cca1290ff69db172d33b0c266fc

Observation e28d5d64-fb4d-436f-a613-e16f9573ee0c · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:3bac8d370c1ed1a659565b6c089a4dc94b2b374c1c1d0bfe4408895a07ca42b6

Observation 2c691ba5-62e3-4848-bc41-35bc1552a65d · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:28.286096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:28.286096Z digest=sha256:396e10bc98fe2d2b7908aad5ed7625fdef3c555f280a6e328256ed0985e8eb6b

Observation 829ec610-9d70-49c4-ad2b-ffa70b67f5d5 · inbound

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges cites this paper.

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T00:32:29.254980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:32:29.254980Z digest=sha256:a289fb092822e9312653f1322e2b2496a1ca6babb382bcb70b1018046f83c737

Observation 79bcbf52-710e-4541-b49b-13b4aeca0da7 · inbound

HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems cites this paper.

HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:20:29.355664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:20:29.355664Z digest=sha256:93d2724cc773d530cf6f67ca6fbd03b12e3d0cd91bd2a3d657696ff81a29c4de