Pith. sign in

Paper Citation Record · LEDGER

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2411.15387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15387 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:27:57.472919Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:43.460215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:12:44.210234Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2370a0a-672d-497f-8ea2-e4fd8fa89166 · outbound

This paper cites write newline.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.295402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.295402Z digest=sha256:c332dd8bdb18d7450b5f1c85a32c2c931f4b7b2e916f6234e3f0badb80a4090b

Observation aa36a654-e550-43e4-9213-79c5aad0fa78 · outbound

This paper cites GPT-4 Technical Report.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.301443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.301443Z digest=sha256:2e2df2e0dbd83fa435310b03593a8900d6fa0458535d9fbec102fbfa1d29bec0

Observation 765fb702-624b-4d61-8a61-ddbaaa0a1c59 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.306616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.306616Z digest=sha256:870a0bbbbd2874542148f5c001e8653c4d15187f26d6276338adb4e982932da8

Observation 3dcade8a-0a5c-43b3-8cef-98894326297d · outbound

This paper cites M., Kanojia, D., de Souza, J.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set M., Kanojia, D., de Souza, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.204094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.311576Z digest=sha256:fc72810f4d852f45a2fa92ace9fa6ef5448a81f6a80e6e80ba36035ab020ec44

Observation 4b980d68-92df-4fd5-9dfb-96c22b49e959 · outbound

This paper cites an unresolved cited work.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:27:58.188254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.316049Z digest=sha256:cb7c10c86c81f94b1ede005b00ce1ffc991969913e51b7e6542f936992a51d14

Observation 2b2bf236-a994-413c-a1eb-17e9eb7896a0 · outbound

This paper cites "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:58.041465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.320397Z digest=sha256:1eb1c141fc3e3e572e07918aebfd1e321e21ce1c04981fe9bbbbc22721f53d20

Observation 39a3ad6b-604d-4c7a-8f84-3b60f3ed1628 · outbound

This paper cites Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.324998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.324998Z digest=sha256:26be62179aa496b9b20746564e07208436589fc94944850e8c11876e48dfa53f

Observation dd017cac-cb1f-4e0f-bfb7-b32a7ec95626 · outbound

This paper cites The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.330246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.330246Z digest=sha256:24ccd749dc78e373052f79057883e3f180536b2f87fa4f2f6036227b06056266

Observation 73e525a6-2169-4b4e-bf02-4281df5ba786 · outbound

This paper cites Experts, errors, and context: A large-scale study of human evaluation for machine translation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Experts, errors, and context: A large-scale study of human evaluation for machine translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.172302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.334819Z digest=sha256:3efaa8ccce67d34ac90b9feb33521308fdfd34548b6c7d740146f754ec938bb8

Observation 80c61444-66a1-4897-8179-48de0c385e17 · outbound

This paper cites Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.157513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.339072Z digest=sha256:3ddf9c40671abff8c7c2f5c997641448e6fae94ddf31684ca8cf1abd539c0931

Observation 3a98697c-e771-411f-9872-735787c7add2 · outbound

This paper cites Are llms breaking mt metrics? results of the wmt24 metrics shared task.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are llms breaking mt metrics? results of the wmt24 metrics shared task

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.142333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.343580Z digest=sha256:cbab8fbde9b2b40763d85fa12efe8f804d7342971d131fad615c970499c67d85

Observation 162cdb0d-e881-4c64-8604-04fcfeee94d6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Gemini: A Family of Highly Capable Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.347977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.347977Z digest=sha256:0ec6ee70963cbaa866e11f4a6182bc12e99e5f1cc0a2e75fece12aefd3ddc6fb

Observation 3d97bd7f-daf8-4e7e-8dd3-0d97340ee51e · outbound

This paper cites Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.352788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.352788Z digest=sha256:8225e45e3c288cdc506eb8be849b973f26885760a428c1b889c5db4b36c96f41

Observation e99e2b29-2927-4043-87dc-330d4cace675 · outbound

This paper cites Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.357217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.357217Z digest=sha256:dce9209eb2cd30cf1225e92e1326b8f3325a90d82457533d4f5ce8f7191a054e

Observation 41501d35-36a2-46ca-a0ae-7dbdd4ab5b33 · outbound

This paper cites xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.361653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.361653Z digest=sha256:9cc87ccf623672e543f90e8467a493832d099d38ddb6ec03ede5e5a8dc2cfa62

Observation 8f166c86-3490-4477-ab0e-aed8fbd52560 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.366287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.366287Z digest=sha256:d23cf8870120f073b0b0e09d940bb8330ce48393451996b765d13ea755866f67

Observation 0f58b68c-837e-4e60-a890-89530f504d13 · outbound

This paper cites MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.370889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.370889Z digest=sha256:ab84b252cd5c827f355f10f01406bb8a6b1e3b9b786a35db603d5e604719e8ae

Observation 9d47ae83-cd20-4f87-866c-0c52ca05b969 · outbound

This paper cites Evaluating LLMs at Detecting Errors in LLM Responses.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Evaluating LLMs at Detecting Errors in LLM Responses

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.375768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.375768Z digest=sha256:2680b36257fc312ce82793d4e14e9c2a76b351abf29aa1868579b7115dd160e3

Observation 62cbb40d-b465-49bf-969c-f7f871e7ca8b · outbound

This paper cites The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.380621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.380621Z digest=sha256:7640986b7e32ea413cddc359a09135c4d9fe346940666e6d678fd352cbb5aad8

Observation 7b5fb24d-e0af-416f-ba39-48c523efc5fe · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus: Inducing fine-grained evaluation capability in language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.127272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.386232Z digest=sha256:26c40ab8c2e33d1587c7bd629129df7db14c86dd468a8a36f9606dcc6c29a69e

Observation e798bd68-e6a6-4124-8cd8-31a2e27dc89b · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.390500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.390500Z digest=sha256:f477734824786ad48e47ac9e3479772f4167732a81785eede61b28270b41edb9

Observation 56011669-a2a1-470e-9f09-bd69fa38be1a · outbound

This paper cites GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.394882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.394882Z digest=sha256:73994ebd32cee66dc8c279f4bd0ff8bc3456f4ed0290d29a6540ab8f6f6033c8

Observation da9dd46f-6bce-454b-b4fd-afeb106f7cc9 · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.399246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.399246Z digest=sha256:831ac8f78ba12d2788b7f9f4ce1e8ee6eb416a2d1116d7b5a309160cd20ba030

Observation 7e24ccdd-f903-4822-aaaf-b41255033d4f · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.403790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.403790Z digest=sha256:631d316ac4c3588fe45dfb3f67f185cd72d6f5ee86f800a98e69b747e0717965

Observation af837742-f587-423a-9371-0a94188de7f7 · outbound

This paper cites Generative Judge for Evaluating Alignment.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Generative Judge for Evaluating Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.408571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.408571Z digest=sha256:3f436c5a64c07175305dcabf6f30751722fd3d27ba194a887b9d0fa456a68e0d

Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · outbound

This paper cites Holistic Evaluation of Language Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.413062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.413062Z digest=sha256:0b66a77f4e45c44eb46ad5de77f2338a00b4c936274e85700aeb00eb0138b7c9

Observation 79d59b41-a065-4226-a157-4df6031187b6 · outbound

This paper cites Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.112367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.417856Z digest=sha256:dc9d62667295ec2f43552d3e50df9046668cc5ae546170fcc8eb92c8a1ac75d1

Observation aac3a4ad-0c5a-45fa-8caf-873e72cc8048 · outbound

This paper cites Training language models to follow instructions with human feedback.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.422180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.422180Z digest=sha256:d3e439d67c95cfffe5ce295880fb8e06d769220fb06b785be307d7815c2c4332

Observation ee483ebd-aa30-4ff4-8e25-2c27482c604c · outbound

This paper cites COMET: A Neural Framework for MT Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set COMET: A Neural Framework for MT Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.426608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.426608Z digest=sha256:6c9c879e947a5ad514cda12171c0480d374eae52570942395b6d7a54fd7b8810

Observation 2be859e2-4de5-4729-930f-44b2ad56413d · outbound

This paper cites Finding Replicable Human Evaluations via Stable Ranking Probability.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Finding Replicable Human Evaluations via Stable Ranking Probability

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:57.772867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.430960Z digest=sha256:9f18e3814035cb69cdcc19ba894e50ec970bce3b883136892d9cba5884858b8a

Observation d831d68e-f4cf-4eec-8d3f-f8c3b811329e · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set BLEURT: Learning Robust Metrics for Text Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.435735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.435735Z digest=sha256:ae64286582339334c3bed635ca8caf6721e1aa114866df388de35255307454ca

Observation 64af8eac-5edb-4c99-9728-25da842f435b · outbound

This paper cites A Benchmark for Learning to Translate a New Language from One Grammar Book.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set A Benchmark for Learning to Translate a New Language from One Grammar Book

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.440359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.440359Z digest=sha256:51025a8dce30a322a0ca0f111abbb446a3d2014fa8fcd85b993b3a9afc1f04e0

Observation 82323832-097c-4040-9c37-f788075d9b54 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.444884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.444884Z digest=sha256:161d61b0be8c9ade0df52059626a8b5e34279847b50f9564de379b0e38b6f6c4

Observation bb60349e-0bc1-4a53-b4dc-57677226650d · outbound

This paper cites Y., Li, L., and Freitag, M.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Y., Li, L., and Freitag, M

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.449578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.449578Z digest=sha256:455c5780e173356275406f53e78befcf9a7277c4824f57fc85a8cdbc2583cf16

Observation 8c0a4f2c-9491-4ecb-83cc-ce198d6d09b0 · outbound

This paper cites Understanding In-Context Learning from Repetitions.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Understanding In-Context Learning from Repetitions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.454353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.454353Z digest=sha256:a76321b3597fa94ffba288070f481cc2e9115437874e09806fae2e131602aa60

Observation 49bb4659-453f-4f01-a427-8747bd00d0f2 · outbound

This paper cites Learning from others' mistakes: Finetuning machine translation models with span-level error annotations.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Learning from others' mistakes: Finetuning machine translation models with span-level error annotations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.459572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.459572Z digest=sha256:f17cb197d34473d3bd529b8899717ae845d4b7105a48a1475e30db08322ff1a0

Observation 177f4ef1-ec04-471b-8832-8d477b046815 · outbound

This paper cites J., Wang, Z., Hwang, J.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set J., Wang, Z., Hwang, J

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.463831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.463831Z digest=sha256:efd38954a23bae0962bbb159f99149d7c9b20bc536d9eb3b48bb6b56735a81be

Observation 1412af09-f169-42d1-9937-190f21bf84d6 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.468381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.468381Z digest=sha256:3cde43e21e892e5566378b3e0a5579ed02ba85edb571f00c1a73007c1ddc434c

Observation 0500f9c6-192d-415a-89e4-b9e6981223a0 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.088220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.472919Z digest=sha256:63a09859d4245f2395404c5750e271f5779a4f2b1f66ff4f1be138541b54c478

Pith citing papers

Observation 07ce0915-a39c-4b1a-ade1-b61e9c2304cd · inbound

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress cites this paper.

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:44.214904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T23:12:43.460215Z digest=sha256:28081cf960e61afca42d814084d0184faf8ebb47b28f7f358dd50f20536a4b45