Pith. sign in

Paper Citation Record · LEDGER

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2501.18771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18771 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:35:45.144387Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81ee31a4-ede3-421f-9a89-6d0575aa511b · outbound

This paper cites write newline.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.059341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.059341Z digest=sha256:f0694e2ae85bcf848fadaf41723beb9077457e5f47f7fd9ad5c5c7adb6e8ead2

Observation cf260182-c407-487f-bf5a-3eab9e0696c2 · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.064264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.064264Z digest=sha256:1097cb77f55041320ba8a935996eee261698dc11e4ddbdb334e818e2680e0a91

Observation ec3434a9-29fc-4871-9321-a4230e5b4fa5 · outbound

This paper cites an unresolved cited work.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T22:35:45.454052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.068809Z digest=sha256:a902bf880fed9c8b88ef69b5ee5c6735f20240b75f87f1e2a645b9aa1c46f849

Observation b6de0e06-c0f6-4561-a774-e8e4c0b7971d · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation PaLM: Scaling Language Modeling with Pathways

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.072685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.072685Z digest=sha256:d9e3f36687346661c3db51d879aeeb5b4268dc648f8889899311c3696a19a707

Observation 09b43ad9-32c8-4c2c-9660-014c5228460b · outbound

This paper cites Results of WMT 23 metrics shared task: Metrics might be guilty but references are not innocent.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Results of WMT 23 metrics shared task: Metrics might be guilty but references are not innocent

Reference 5

Resolution
verified exact
doi, observed 2026-08-09T22:35:45.186621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.077315Z digest=sha256:56b5aede0a8536c802fc10118822c1587d84b86e1fbeeae946cdbd3085146425

Observation f2672c02-d44d-4d9a-82e1-a4eef29d5431 · outbound

This paper cites The F lores-101 evaluation benchmark for low-resource and multilingual machine translation.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation The F lores-101 evaluation benchmark for low-resource and multilingual machine translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.081309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.081309Z digest=sha256:417223e8bb672fe863f724c9627aabab3ffb6061a3881e1e395c615146e74c32

Observation af7211f0-4400-4785-bccb-b34ba2ed63fa · outbound

This paper cites Investigating Data Contamination for Pre-training Language Models.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Investigating Data Contamination for Pre-training Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.085387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.085387Z digest=sha256:2b3ecccb61103bc96c6db57f5231fe099e775ec7bcfedc5dec2112f92a831fc2

Observation 9608027c-02c0-44bb-be87-4b6eb08935a5 · outbound

This paper cites M etric X -23: The G oogle submission to the WMT 2023 metrics shared task.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation M etric X -23: The G oogle submission to the WMT 2023 metrics shared task

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.090267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.090267Z digest=sha256:c5cb629a952bb21f0d79a5fc9a609662e0e583d2c49d1458aac749286077b41d

Observation 96019eaa-ed2c-4170-bfa0-5c5526f4d9a1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Adam: A Method for Stochastic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.094046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.094046Z digest=sha256:e30b1a712619eec59515c1032133c743dbaf274a12a93105e312567ec1748294

Observation 0e148f8f-0431-41ee-9e46-df3d2de8329a · outbound

This paper cites Findings of the 2023 conference on machine translation ( WMT 23): LLM s are here but not quite there yet.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Findings of the 2023 conference on machine translation ( WMT 23): LLM s are here but not quite there yet

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.098252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.098252Z digest=sha256:f2067e357c9b72bdf60af41f8081cc702151ef30569908b8de89e5013f195daa

Observation 428c447a-3be4-4046-91d3-67cb70e39e6e · outbound

This paper cites Findings of the WMT 24 general machine translation shared task: The LLM era is here but MT is not solved yet.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Findings of the WMT 24 general machine translation shared task: The LLM era is here but MT is not solved yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.101914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.101914Z digest=sha256:23a0c054c4f9ba08f054b7812e4a32bb7455f113ba875577403b1fdff2355dfb

Observation 4b167f55-f31a-4a58-a8da-ea41249747bf · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.105659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.105659Z digest=sha256:0442d3699ef30f137ef7ca389cc76d3c797bbe0b7318f338b0c2224bd7d5fc6c

Observation 4d1e0dbd-e66c-4a39-84a0-d770bef61d11 · outbound

This paper cites Madlad-400: a multilingual and document-level large audited dataset.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Madlad-400: a multilingual and document-level large audited dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:35:45.431797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.109756Z digest=sha256:f782a544fa3a2123b1f29d2f59cf2ba3f832afad0e3a1c47d1a285d8faab059e

Observation 32751535-6613-4025-894f-e9b7eed3917c · outbound

This paper cites Data Contamination: From Memorization to Exploitation.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Data Contamination: From Memorization to Exploitation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.113380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.113380Z digest=sha256:b53890f7fe35f4f01b0044044006897168e03e0446917ec369981c368412200b

Observation c2f3c6fd-1ea6-42c7-93a5-4daaccee9f63 · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Proving Test Set Contamination in Black Box Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.117286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.117286Z digest=sha256:6076fa67cc6251ee9b91abfa850d5bdb0211d3dfe1219ffd7e90f2591a8fa550

Observation cd27d577-d485-4558-b151-4cf56783ba3c · outbound

This paper cites B leu: a method for automatic evaluation of machine translation.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation B leu: a method for automatic evaluation of machine translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.121076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.121076Z digest=sha256:4f0a61f7afab08a9c42fe168df28f32bc612258b41426599b8cb8dd420f536f7

Observation ab1e9976-e1a9-43a0-b805-cea319aded1d · outbound

This paper cites Data Contamination Report from the 2024 CONDA Shared Task.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Data Contamination Report from the 2024 CONDA Shared Task

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-09T22:35:45.237824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.124513Z digest=sha256:4e194da3a4facd70d0c253c9f0c20faa41d0749b9e0a8fa9d4ff30a6deb55206

Observation 41044f83-7225-41f4-82f9-28d7e2414548 · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Detecting Pretraining Data from Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.128351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.128351Z digest=sha256:0423758c7ae52758d9fd8dbc098b1fa52186ed3bb2eb5823bf0ef83fceb58fa3

Observation aa2be92d-3aa0-4151-83b0-90cfc9250743 · outbound

This paper cites Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.132126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.132126Z digest=sha256:5be6c7b7fbae250ff6ecc018dd0d0076567a2a780129f583d9671f3e71f92094

Observation bcbaf847-3e5f-450e-9f1a-2e77704fc02f · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.135371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.135371Z digest=sha256:b1ef1b90b5c708e0e88665a89345ba11a9cba2b9685ec2f288f2da0193254532

Observation e903e07a-34cc-4dce-8345-835635a10ca7 · outbound

This paper cites Attention is all you need.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Attention is all you need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.138456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.138456Z digest=sha256:9efc6c5cc05c9f8f36a250ccb29de75ab6d867b76aed47e39a61e854c4c97651

Observation ac7606fa-bfa5-4291-b7d0-361a84b6553a · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.141387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.141387Z digest=sha256:b0d8809600edae0c97c4fa8b7298ccfeb654e30b91530c7cd357fc61eaabc2ba

Observation 11b8b444-9008-475b-97b5-361819185eb7 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.144387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.144387Z digest=sha256:9ccf255ba8f0400603d452f7637e7b105ebb90089c62c61dc3e7be52a3b6784a

Pith citing papers

No inbound Pith citation observations are available.