Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:53:25.065089Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 8 inbound Pith citation observations for arXiv:2504.17087.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:53:25.065089Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:20:29.355664Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:49:38.227079Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0dd83741-44ae-4dca-ad01-b80f8ac6a110 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abf0c6f-6050-487a-8758-2119a025dd38 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1e6529-6777-4d0a-aee8-b7d7d03944d9 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08bb1446-c4a7-4589-a742-fe0c0c382d0f · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96debcb-a0ab-4ee9-ac07-e64c686188cc · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments JudgeBench: A Benchmark for Evaluating LLM-based Judges
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 164e792e-d0f6-48b5-9097-56408af56e71 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7609c747-f80d-423a-9148-12fccf3a474b · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Self-rationalization improves LLM as a fine-grained judge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f34a9d7-520f-44d4-9e48-78738a59461b · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47dca15f-0f4b-40ed-8137-76226752cf07 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Aligning Large Language Models with Human: A Survey
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e07aea5-8e62-4d70-b3e6-3f623bb50998 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246fa49d-6c82-4101-a111-b0abb1c1e14f · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852a8c22-4244-4b4e-b17f-86327ee6750c · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181b3b98-32bd-46b0-b91e-2b6f4e16afd7 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Do Large Language Models Know What They Don't Know?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e0b0f99-3c9a-46e2-b141-122e026017c2 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06e17a48-7e15-4e51-9542-c596ecc4391e · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments BERTScore: Evaluating Text Generation with BERT
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 277ec298-e4d6-44ca-be0b-cf43c770d623 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Detail of different rubrics In the single-agent precision comparison section, we analyzed the impact of four different rubric configurations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b2e1119a-76ad-4be0-9649-d4222655c614 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd9b1b95-b754-403b-b922-e5608079069c · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments MQAG: Multiple-choice Question Answering and Generation for Assessing Information Consistency in Summarization
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475da4c9-edf1-407a-b872-92fe5788823d · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Exploring LLM Prompting Strategies for Joint Essay Scoring and Feedback Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320b609c-1579-4ee3-afd3-ad55c33d9230 · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Evaluating Large Language Models: A Comprehensive Survey
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce2fbba-1898-4282-9b13-4a7dff31c2ac · outbound
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments Debating with More Persuasive LLMs Leads to More Truthful Answers
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82cfc1d-e4e6-4d7d-b402-abcc60a7d50d · inbound
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06775abf-bd4c-4d41-9862-16567315a85c · inbound
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5ad57e2c-14b7-4923-bd6c-9e68778c3c2f · inbound
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09773501-43c3-4a0a-a3b3-399f1e2445f3 · inbound
Counsel: A Meta-Evaluation Dataset for Agentic Tasks Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e28d5d64-fb4d-436f-a613-e16f9573ee0c · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c691ba5-62e3-4848-bc41-35bc1552a65d · inbound
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829ec610-9d70-49c4-ad2b-ffa70b67f5d5 · inbound
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bcbf52-710e-4541-b49b-13b4aeca0da7 · inbound
HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.