Pith. sign in

Paper Citation Record · LEDGER

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2501.18771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18771 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:35:45.144387Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81ee31a4-ede3-421f-9a89-6d0575aa511b · outbound

This paper cites write newline.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.059341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.059341Z digest=sha256:bb00dc3a2e6954952cf3ee43d0b0f9dcb93dc3a899a282fe175ae441f5ce0a24

Observation cf260182-c407-487f-bf5a-3eab9e0696c2 · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.064264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.064264Z digest=sha256:f9c7d1c0ef9493c16f2aeb0706e6975a771452d3e4bd2d799ab05a887ac041c1

Observation ec3434a9-29fc-4871-9321-a4230e5b4fa5 · outbound

This paper cites an unresolved cited work.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T22:35:45.454052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.068809Z digest=sha256:71088c15e00480bb2ddcdd65a3402278ed6ca55913151dff9bdc524767040d96

Observation b6de0e06-c0f6-4561-a774-e8e4c0b7971d · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation PaLM: Scaling Language Modeling with Pathways

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.072685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.072685Z digest=sha256:26f81199990be77b6be8e42bf8bc0f4a2a35c8b62f82efe3d11a58d4b7bc6424

Observation 09b43ad9-32c8-4c2c-9660-014c5228460b · outbound

This paper cites Results of WMT 23 metrics shared task: Metrics might be guilty but references are not innocent.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Results of WMT 23 metrics shared task: Metrics might be guilty but references are not innocent

Reference 5

Resolution
verified exact
doi, observed 2026-08-09T22:35:45.186621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.077315Z digest=sha256:28e356591b94fcafb59d9cbb022675cacdb44b1d8f7a10d933ffe2996d837297

Observation f2672c02-d44d-4d9a-82e1-a4eef29d5431 · outbound

This paper cites The F lores-101 evaluation benchmark for low-resource and multilingual machine translation.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation The F lores-101 evaluation benchmark for low-resource and multilingual machine translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.081309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.081309Z digest=sha256:9869f846cfc005cce446dcdd781cc9f844bdb2ce86c8f5a38e2279e9c99f50bc

Observation af7211f0-4400-4785-bccb-b34ba2ed63fa · outbound

This paper cites Investigating Data Contamination for Pre-training Language Models.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Investigating Data Contamination for Pre-training Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.085387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.085387Z digest=sha256:7d4f7be9953556ae86c6051b35a57c6bc7ea1ff60fe0c75aba70103c97d28b29

Observation 9608027c-02c0-44bb-be87-4b6eb08935a5 · outbound

This paper cites M etric X -23: The G oogle submission to the WMT 2023 metrics shared task.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation M etric X -23: The G oogle submission to the WMT 2023 metrics shared task

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.090267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.090267Z digest=sha256:ffdeaf2f3948483d9db38f517802c270f3781c126ff4650059272127bb5d0dc0

Observation 96019eaa-ed2c-4170-bfa0-5c5526f4d9a1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Adam: A Method for Stochastic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.094046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.094046Z digest=sha256:48a9df0a762f2bb6acc4315ab1f2e5432f623bcd16886bdaaf6ef417f01d547d

Observation 0e148f8f-0431-41ee-9e46-df3d2de8329a · outbound

This paper cites Findings of the 2023 conference on machine translation ( WMT 23): LLM s are here but not quite there yet.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Findings of the 2023 conference on machine translation ( WMT 23): LLM s are here but not quite there yet

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.098252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.098252Z digest=sha256:3e47deb12c12908de73d332b7ea21c40c7e5c3e0ce6e700ebf03bd6043411284

Observation 428c447a-3be4-4046-91d3-67cb70e39e6e · outbound

This paper cites Findings of the WMT 24 general machine translation shared task: The LLM era is here but MT is not solved yet.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Findings of the WMT 24 general machine translation shared task: The LLM era is here but MT is not solved yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.101914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.101914Z digest=sha256:ed2ba5bcfd019d8175b6542e2fd2b28f35dbc3b99784e652bcfb58701a3097c0

Observation 4b167f55-f31a-4a58-a8da-ea41249747bf · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.105659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.105659Z digest=sha256:b7605d9e69c054fbaebe43848879c4d253640f1583fb3caf58a973f598fa33e3

Observation 4d1e0dbd-e66c-4a39-84a0-d770bef61d11 · outbound

This paper cites Madlad-400: a multilingual and document-level large audited dataset.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Madlad-400: a multilingual and document-level large audited dataset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T22:35:45.431797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.109756Z digest=sha256:b68006276d9d9420676a3a91fe02268d2e53cdfe02e4314886396771a4e4eb9f

Observation 32751535-6613-4025-894f-e9b7eed3917c · outbound

This paper cites Data Contamination: From Memorization to Exploitation.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Data Contamination: From Memorization to Exploitation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.113380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.113380Z digest=sha256:f81021136e9716416cadbb79132e6ec78f7a5bed1d97f07b39294ce0cf8e116c

Observation c2f3c6fd-1ea6-42c7-93a5-4daaccee9f63 · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Proving Test Set Contamination in Black Box Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.117286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.117286Z digest=sha256:3f1bca6fe1ada7304adbd2a23b0e75ea048e157157b72d54935b566356122798

Observation cd27d577-d485-4558-b151-4cf56783ba3c · outbound

This paper cites B leu: a method for automatic evaluation of machine translation.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation B leu: a method for automatic evaluation of machine translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.121076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.121076Z digest=sha256:ad0d21d59b534ccf36e882b5be6ce93ef580138aa7cc721af6be87d85abd062e

Observation ab1e9976-e1a9-43a0-b805-cea319aded1d · outbound

This paper cites Data Contamination Report from the 2024 CONDA Shared Task.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Data Contamination Report from the 2024 CONDA Shared Task

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-09T22:35:45.237824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T22:35:45.124513Z digest=sha256:304b238c7f933fd8fbe0a753c2d6342107eeecfd4cc7599b4dcc7e36624fcef2

Observation 41044f83-7225-41f4-82f9-28d7e2414548 · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Detecting Pretraining Data from Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.128351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.128351Z digest=sha256:0d43fb03eea5d1618e063642d1804742956f6f57c2e55a9afd034b525b375880

Observation aa2be92d-3aa0-4151-83b0-90cfc9250743 · outbound

This paper cites Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.132126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.132126Z digest=sha256:f0d23438ffd6f3af3d393bbb0d99fb06056e467ac27f0392d977ef367f0f340f

Observation bcbaf847-3e5f-450e-9f1a-2e77704fc02f · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.135371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.135371Z digest=sha256:0a0b487f0088a0cf4c5d250d111d097a2b01e8f6e8997ac1eebe65c5d399d09d

Observation e903e07a-34cc-4dce-8345-835635a10ca7 · outbound

This paper cites Attention is all you need.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Attention is all you need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.138456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.138456Z digest=sha256:9b2e02b01437d7b93a0f6a364a99496853b77b6772c0a9c55746045c62dbf624

Observation ac7606fa-bfa5-4291-b7d0-361a84b6553a · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.141387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.141387Z digest=sha256:74229439915b412917c0c0a0514c7acc4cefed96828181b3b02e9277a7f45586

Observation 11b8b444-9008-475b-97b5-361819185eb7 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T22:35:45.144387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:35:45.144387Z digest=sha256:67a9ef9f92388411c5d1d1d4b2aade0196bad90b947eae73391d8b1d716fb6b3

Pith citing papers

No inbound Pith citation observations are available.