Pith. sign in

Paper Citation Record · LEDGER

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2411.15387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15387 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:27:57.472919Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:43.460215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:12:44.210234Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2370a0a-672d-497f-8ea2-e4fd8fa89166 · outbound

This paper cites write newline.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.295402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.295402Z digest=sha256:962d3819fe59d46ffa5752acd924df40a3fb1955a92ff2a65a51ee6e44e00804

Observation aa36a654-e550-43e4-9213-79c5aad0fa78 · outbound

This paper cites GPT-4 Technical Report.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.301443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.301443Z digest=sha256:cca3824907242b5560fe4c53204229b9b43c07e92570ac126a236a7caacb51c2

Observation 765fb702-624b-4d61-8a61-ddbaaa0a1c59 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.306616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.306616Z digest=sha256:7926e0a3fe6232eb634a717c872af303696dadc3284eb7cd0aeaf289d41095bc

Observation 3dcade8a-0a5c-43b3-8cef-98894326297d · outbound

This paper cites M., Kanojia, D., de Souza, J.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set M., Kanojia, D., de Souza, J

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.204094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.311576Z digest=sha256:b6a01830b81f5a457dc2c2f00a1baed7fd5ff883c37c3744a9393f2e3ab818e2

Observation 4b980d68-92df-4fd5-9dfb-96c22b49e959 · outbound

This paper cites an unresolved cited work.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:27:58.188254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.316049Z digest=sha256:489f236170324006d4b7825cbef0b9c4482c02a0cc7ad7566db29e282cad0346

Observation 2b2bf236-a994-413c-a1eb-17e9eb7896a0 · outbound

This paper cites "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set "Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:58.041465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.320397Z digest=sha256:c1f9c3d814df5d59f906859eefbf8edb33c7f11d019b57e7233d46cc77b43405

Observation 39a3ad6b-604d-4c7a-8f84-3b60f3ed1628 · outbound

This paper cites Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.324998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.324998Z digest=sha256:39d1a1c1f6350cb402889a69feef5eb39d6fe1699b9c295afc1e5eba32a90948

Observation dd017cac-cb1f-4e0f-bfb7-b32a7ec95626 · outbound

This paper cites The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.330246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.330246Z digest=sha256:bbc7aa2e45182678446b3fbabf8fa25882aacc1ac132a37cf58895b6d812b3fa

Observation 73e525a6-2169-4b4e-bf02-4281df5ba786 · outbound

This paper cites Experts, errors, and context: A large-scale study of human evaluation for machine translation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Experts, errors, and context: A large-scale study of human evaluation for machine translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.172302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.334819Z digest=sha256:75624b32aecb7e8913344d716a706817340e283251ea8bd1311bc248fcf5dd2a

Observation 80c61444-66a1-4897-8179-48de0c385e17 · outbound

This paper cites Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Results of wmt23 metrics shared task: Metrics might be guilty but references are not innocent

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.157513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.339072Z digest=sha256:4174bd5fc370839ce2df45593677ed7ba517363b10664c0840a81e11464bc283

Observation 3a98697c-e771-411f-9872-735787c7add2 · outbound

This paper cites Are llms breaking mt metrics? results of the wmt24 metrics shared task.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are llms breaking mt metrics? results of the wmt24 metrics shared task

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.142333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.343580Z digest=sha256:29e329a1bb91a0a3fa332014421a820e41f1f11ad971c335b4da81349c1b501c

Observation 162cdb0d-e881-4c64-8604-04fcfeee94d6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Gemini: A Family of Highly Capable Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.347977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.347977Z digest=sha256:af9039110a8d5164c9f051c70972e521313affce8150eddce277aebe8199d715

Observation 3d97bd7f-daf8-4e7e-8dd3-0d97340ee51e · outbound

This paper cites Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.352788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.352788Z digest=sha256:7bdcbbdf0ed9cefbd307c4770214626b460325e063eda041ce21e960db241c32

Observation e99e2b29-2927-4043-87dc-330d4cace675 · outbound

This paper cites Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.357217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.357217Z digest=sha256:8111ccded5f175263973ceb8898006e2896a602e45ea691fe1b4d9fc4afc622e

Observation 41501d35-36a2-46ca-a0ae-7dbdd4ab5b33 · outbound

This paper cites xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.361653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.361653Z digest=sha256:c1eee4a17a93aa08a31c3fe58bf7743b5e4db01fcfd20c108e253650e16fafd3

Observation 8f166c86-3490-4477-ab0e-aed8fbd52560 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.366287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.366287Z digest=sha256:9de3fe1f8a0a8cca8c4a303fd5f8a7b528d483ba958080ed48bea1419cde60c4

Observation 0f58b68c-837e-4e60-a890-89530f504d13 · outbound

This paper cites MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.370889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.370889Z digest=sha256:98219232de8aeef810376c39354dffec6b1b02f67e139c6c8150e6124633fe63

Observation 9d47ae83-cd20-4f87-866c-0c52ca05b969 · outbound

This paper cites Evaluating LLMs at Detecting Errors in LLM Responses.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Evaluating LLMs at Detecting Errors in LLM Responses

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.375768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.375768Z digest=sha256:ffabb21316bea97b856f7cdfa4750ef0c5c4c78004acd97ed40634dbccf6bbd6

Observation 62cbb40d-b465-49bf-969c-f7f871e7ca8b · outbound

This paper cites The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.380621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.380621Z digest=sha256:0958729cbf8900d3d2f090ed6e949de41ababf9f0932b4700af44ca064de874a

Observation 7b5fb24d-e0af-416f-ba39-48c523efc5fe · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus: Inducing fine-grained evaluation capability in language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.127272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.386232Z digest=sha256:b5918462c66196d89af3a7271e79f74ed64f105018699b64422995c476ea3905

Observation e798bd68-e6a6-4124-8cd8-31a2e27dc89b · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.390500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.390500Z digest=sha256:086a61dcd01947e354548315e3d344a56a644ec0969c8dc4f5cc9ddf05379513

Observation 56011669-a2a1-470e-9f09-bd69fa38be1a · outbound

This paper cites GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.394882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.394882Z digest=sha256:84341653cf86950ba993a38fb567465d92a6567b65aa0997e4840633cb896950

Observation da9dd46f-6bce-454b-b4fd-afeb106f7cc9 · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.399246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.399246Z digest=sha256:2e9170eb92ba3811df5e3ee4c133c6d09f7a76b356dd6b3390871104ab3cf742

Observation 7e24ccdd-f903-4822-aaaf-b41255033d4f · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.403790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.403790Z digest=sha256:23c5fe22faf2a8ae25f3c6c5c97e8474751007486e7f8b588d1c6719f90f9955

Observation af837742-f587-423a-9371-0a94188de7f7 · outbound

This paper cites Generative Judge for Evaluating Alignment.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Generative Judge for Evaluating Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.408571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.408571Z digest=sha256:65aaa5115234760cad0aa30a104438e6bf0cf491458e989985c055458b25599b

Observation dbf6de75-1d22-4c99-acec-3c2b57599080 · outbound

This paper cites Holistic Evaluation of Language Models.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.413062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.413062Z digest=sha256:28f889cf57855fb2d5623bd606d053466f1403e4fdc115d4c88f650c674bb3ad

Observation 79d59b41-a065-4226-a157-4df6031187b6 · outbound

This paper cites Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.112367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.417856Z digest=sha256:9dd588d8293b3251af1cd10e8f0bf98f52dfc9440e6b10c1e860252035e37fd2

Observation aac3a4ad-0c5a-45fa-8caf-873e72cc8048 · outbound

This paper cites Training language models to follow instructions with human feedback.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.422180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.422180Z digest=sha256:daab49f3399e2f29d0ba45354ecda993496b45077d676ff50bda644f9b4c95fc

Observation ee483ebd-aa30-4ff4-8e25-2c27482c604c · outbound

This paper cites COMET: A Neural Framework for MT Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set COMET: A Neural Framework for MT Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.426608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.426608Z digest=sha256:6ee7c7142d7aa358b92cbf50882a1a0d4d82c7f65e02aedfb7ac675d031fe692

Observation 2be859e2-4de5-4729-930f-44b2ad56413d · outbound

This paper cites Finding Replicable Human Evaluations via Stable Ranking Probability.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Finding Replicable Human Evaluations via Stable Ranking Probability

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:57.772867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.430960Z digest=sha256:e9221514296bd6000b7892e73084644dc3a102420f9b4ca6ae3e0e7dd9b9dc5e

Observation d831d68e-f4cf-4eec-8d3f-f8c3b811329e · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set BLEURT: Learning Robust Metrics for Text Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.435735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.435735Z digest=sha256:636351eb19043e2316a8562a1cf5e2c9e426bd79a5747c4363d42308aa15fd36

Observation 64af8eac-5edb-4c99-9728-25da842f435b · outbound

This paper cites A Benchmark for Learning to Translate a New Language from One Grammar Book.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set A Benchmark for Learning to Translate a New Language from One Grammar Book

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.440359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.440359Z digest=sha256:ed58c80648616312cf2a2755d85e404a3c1383ac7fd3b25fe5cca9eada5be822

Observation 82323832-097c-4040-9c37-f788075d9b54 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.444884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.444884Z digest=sha256:0b20b22a62a272622a0b7559a4e22c4fc081c60630dcd306a7cdc4633dfe6a8d

Observation bb60349e-0bc1-4a53-b4dc-57677226650d · outbound

This paper cites Y., Li, L., and Freitag, M.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Y., Li, L., and Freitag, M

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.449578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.449578Z digest=sha256:d7147901935154f7faca2ff1fea1b6949e2665be51c5989021bf7cd1938f162f

Observation 8c0a4f2c-9491-4ecb-83cc-ce198d6d09b0 · outbound

This paper cites Understanding In-Context Learning from Repetitions.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Understanding In-Context Learning from Repetitions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.454353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.454353Z digest=sha256:11a75cf8bd5c1ccea40c9ddb26533457bc360864f7ff2fd4d0256b128df2f37e

Observation 49bb4659-453f-4f01-a427-8747bd00d0f2 · outbound

This paper cites Learning from others' mistakes: Finetuning machine translation models with span-level error annotations.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Learning from others' mistakes: Finetuning machine translation models with span-level error annotations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.459572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.459572Z digest=sha256:d0d560522881012869385c29ea16896350fbcb5b353e9dbc871ef39a3188cc0f

Observation 177f4ef1-ec04-471b-8832-8d477b046815 · outbound

This paper cites J., Wang, Z., Hwang, J.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set J., Wang, Z., Hwang, J

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.463831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.463831Z digest=sha256:432390edaf8f4c369167ca294e156edaf0d0c3172f7aa6d9733349e11899d0d9

Observation 1412af09-f169-42d1-9937-190f21bf84d6 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:57.468381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:57.468381Z digest=sha256:5d969c3fdc00cc3dd726628103d2bf67fad32429382f8d5c2f1f875d9328c106

Observation 0500f9c6-192d-415a-89e4-b9e6981223a0 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:58.088220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T14:27:57.472919Z digest=sha256:b4f46e5ed2434155260dad88e6846f236d67cf40d2fb42f513aa4b10bcd6b66b

Pith citing papers

Observation 07ce0915-a39c-4b1a-ade1-b61e9c2304cd · inbound

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress cites this paper.

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:12:44.214904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T23:12:43.460215Z digest=sha256:134324553bbf90774b63feb10a0ca3b10a842e761d3acd36a81badd374bf7abf