Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2408.13006.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:24:54.518422Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 02d0b65c-f5f1-4087-b51f-22bad0f786a2 · inbound
Engineering AI Judge Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93aad358-4b48-40d0-b804-ff1f51b37eb0 · inbound
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03aa5361-1225-4f49-870b-270152bb2229 · inbound
JuStRank: Benchmarking LLM Judges for System Ranking Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 717fec84-1dc1-4046-a27b-d65922246322 · inbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64df046f-feb9-42f9-8849-fabc483a1f95 · inbound
Doing More with Less: A Survey on Routing Strategies for Resource Optimisation in Large Language Model-Based Systems Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 048f9501-3b4e-46a5-af7c-b65f2aed6af8 · inbound
AI Alignment at Your Discretion Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c736d4fb-3f19-43d8-bef9-b167f1e36ba6 · inbound
Towards an AI co-scientist Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f589180a-ebc3-4b9c-bb6e-911a83749ed3 · inbound
Benchmarking Multi-National Value Alignment for Large Language Models Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea756bce-79bd-4c0a-826f-358842298d06 · inbound
Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e4fd5f-8b6c-4d20-b406-07454b960542 · inbound
EnronQA: Towards Personalized RAG over Private Documents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ba14f9-bb84-44fb-b442-9d6f7946c53a · inbound
Large Language Models for Predictive Analysis: How Far Are They? Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c9c9731-28a4-45d0-914a-91332507933e · inbound
Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef315790-9b74-4f37-80c1-6a150431efaa · inbound
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60835722-30b9-4029-95ec-185286065f5e · inbound
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32569ca-49f8-4a11-bbe0-c12970b98b9b · inbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da944267-1197-4876-a579-307bef239571 · inbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fef8a8-1a7f-49d7-b3b8-a8bf0d604315 · inbound
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dabb476-796a-40e7-a1a1-95629d87f9c2 · inbound
Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0bb244-766b-40a7-bb26-8341a1390907 · inbound
AI Propaganda factories with language models Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c465ca9-2c19-4f63-a87d-aee427466f19 · inbound
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb117fe-1c82-4dba-8867-5cee665d83b0 · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49d3437-8437-407b-8c8d-2b048b58bf33 · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d85efb-00b7-4abd-8b08-b003eaafd557 · inbound
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cfcb7436-1bf7-4185-b89a-df84914418af · inbound
Iterative Finetuning is Mostly Idempotent Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f2209723-09be-42a7-9623-9f00e582529d · inbound
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7133626b-ec70-40b2-b733-4667d82b46c1 · inbound
Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ff3a2b6-4a2a-453b-8bed-405c8d29267d · inbound
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.