Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:27:57.472919Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2411.15387.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:27:57.472919Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:43.460215Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T23:12:44.210234Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2370a0a-672d-497f-8ea2-e4fd8fa89166 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa36a654-e550-43e4-9213-79c5aad0fa78 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765fb702-624b-4d61-8a61-ddbaaa0a1c59 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dcade8a-0a5c-43b3-8cef-98894326297d · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set M., Kanojia, D., de Souza, J
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4b980d68-92df-4fd5-9dfb-96c22b49e959 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b2bf236-a994-413c-a1eb-17e9eb7896a0 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 39a3ad6b-604d-4c7a-8f84-3b60f3ed1628 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd017cac-cb1f-4e0f-bfb7-b32a7ec95626 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e525a6-2169-4b4e-bf02-4281df5ba786 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Experts, errors, and context: A large-scale study of human evaluation for machine translation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 80c61444-66a1-4897-8179-48de0c385e17 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a98697c-e771-411f-9872-735787c7add2 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are llms breaking mt metrics? results of the wmt24 metrics shared task
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 162cdb0d-e881-4c64-8604-04fcfeee94d6 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Gemini: A Family of Highly Capable Multimodal Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d97bd7f-daf8-4e7e-8dd3-0d97340ee51e · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99e2b29-2927-4043-87dc-330d4cace675 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41501d35-36a2-46ca-a0ae-7dbdd4ab5b33 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f166c86-3490-4477-ab0e-aed8fbd52560 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Measuring Massive Multitask Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f58b68c-837e-4e60-a890-89530f504d13 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d47ae83-cd20-4f87-866c-0c52ca05b969 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Evaluating LLMs at Detecting Errors in LLM Responses
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62cbb40d-b465-49bf-969c-f7f871e7ca8b · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b5fb24d-e0af-416f-ba39-48c523efc5fe · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus: Inducing fine-grained evaluation capability in language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e798bd68-e6a6-4124-8cd8-31a2e27dc89b · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56011669-a2a1-470e-9f09-bd69fa38be1a · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da9dd46f-6bce-454b-b4fd-afeb106f7cc9 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e24ccdd-f903-4822-aaaf-b41255033d4f · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af837742-f587-423a-9371-0a94188de7f7 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Generative Judge for Evaluating Alignment
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d59b41-a065-4226-a157-4df6031187b6 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aac3a4ad-0c5a-45fa-8caf-873e72cc8048 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee483ebd-aa30-4ff4-8e25-2c27482c604c · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set COMET: A Neural Framework for MT Evaluation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be859e2-4de5-4729-930f-44b2ad56413d · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Finding Replicable Human Evaluations via Stable Ranking Probability
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d831d68e-f4cf-4eec-8d3f-f8c3b811329e · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set BLEURT: Learning Robust Metrics for Text Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64af8eac-5edb-4c99-9728-25da842f435b · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set A Benchmark for Learning to Translate a New Language from One Grammar Book
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82323832-097c-4040-9c37-f788075d9b54 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb60349e-0bc1-4a53-b4dc-57677226650d · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Y., Li, L., and Freitag, M
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0a4f2c-9491-4ecb-83cc-ce198d6d09b0 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Understanding In-Context Learning from Repetitions
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49bb4659-453f-4f01-a427-8747bd00d0f2 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 177f4ef1-ec04-471b-8832-8d477b046815 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set J., Wang, Z., Hwang, J
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1412af09-f169-42d1-9937-190f21bf84d6 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0500f9c6-192d-415a-89e4-b9e6981223a0 · outbound
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07ce0915-a39c-4b1a-ade1-b61e9c2304cd · inbound
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.