Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:18:01.686274Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2505.22251.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:18:01.686274Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T22:51:15.098039Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T22:52:44.973580Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e51946b5-eedc-4742-9706-00bb3ea0841f · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Prompting large language models with speech recognition abilities,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82ca8082-c047-4b76-aa76-599ae9d6f911 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Salsa: Speedy asr-llm synchronous aggregation,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e5c9e119-9eac-45d8-a44e-e4032f22e4fb · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Delayed fusion: Integrating large language models into first- pass decoding in end-to-end speech recognition,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a97d9d19-e543-44ee-8123-a8ea3890fc73 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f794d4-eb9c-47e0-a134-779f495bdd9f · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1848cb06-81ef-41d0-ae81-f904473fe067 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition On decoder-only architec- ture for speech-to-text and large language model integration,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f2334377-8217-4782-9570-370dd76bed53 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Can Generative Large Language Models Perform ASR Error Correction?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579a16f0-cc67-44db-8fc6-063faad13915 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Contextual spelling correction with large language models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd9d0c1a-9bd8-4a7b-af49-08c078b1c8bc · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Denoising LM: Pushing the limits of error correction models for speech recognition,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20558f48-2ba2-4662-9ace-e767da79a8bc · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a5fbcd1b-5a4b-473f-92be-3d3fc27230e3 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Data contamination: From mem- orization to exploitation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d0f92251-4b5d-4cd0-9784-3197c3cb07ae · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2cb5b22d-a54c-4581-8366-baa3bdd85dfa · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Lib- rispeech: An ASR corpus based on public domain audio books,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e27156b0-dd5c-40d6-ac32-2675cfb83de5 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Common voice: A massively-multilingual speech corpus,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66807b87-4d71-456c-b571-8ab88bb88c99 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae18e7e6-64e0-4c1d-bd23-23fa1a04c57a · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Comparing discrete and continuous space llms for speech recognition,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47c931db-4bf3-46f6-aa78-2b90cefacb9d · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Connecting speech encoder and large language model for ASR,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d87d8bc6-a3bb-4405-8617-b3fdbe805c69 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0901f8d5-522a-4fc3-9a6d-41b30ebce6f4 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b04456-488c-47da-9c15-3d516c30a64f · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Efficient Streaming LLM for Speech Recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dc9d712-d060-4667-a2c6-104c24596087 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Ctc-assisted llm-based contextual asr,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93838fd6-cedd-4eff-9577-2989b3746f90 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition The bigscience ROOTS corpus: A 1.6TB composite multilingual dataset,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 813bbacb-f1be-4c2e-a6b9-3ceb4ffffae6 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition RedPajama: an open dataset for training large language models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9508cdb2-5227-4d14-89d6-c753b08e84fa · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Dolma: an open corpus of three trillion tokens for language model pretraining research,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bca93d8b-50e0-40af-ae8a-631cc3c33fcc · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac2a22a-6122-4127-b0af-09d970d02d07 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0a6792-3ef8-4a52-bd2e-d1347e81dd8e · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition The Llama 3 Herd of Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8ac505-2c60-4ce1-a392-e459fc14b16d · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition LLaMA: Open and Efficient Foundation Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1012a3e-8e98-4224-8bb6-a5961c78016d · outbound
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 963276c5-1b92-4756-ae75-7220c0c22c4b · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Benchmarking non- parametric statistical tests,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3e11ba02-3e84-4659-b8b7-06f2db9663c4 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Confidence intervals for evaluation in machine learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ca12751-7f35-42c3-9e4d-f0963c7c566a · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Pythia: A suite for analyzing large language models across training and scaling,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 987ac905-bc74-4c2f-9c94-179b75bebcc4 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition GPT-NeoX-20B: An open-source autoregressive language model,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72c03d69-2317-4751-a8c3-9cb4bce3e637 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition OPT: Open Pre-trained Transformer Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7103772-83af-47fd-a90e-5ef74418b67d · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition OLMo: Accelerating the science of language models,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 532a5e63-10f6-4de3-a0a2-77984b0eeef9 · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition The secret sharer: Evaluating and testing unintended memorization in neural networks,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21b8daf0-bacf-4925-9322-a8e58f77506c · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition Open- source conversational AI with SpeechBrain 1.0,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 69c183be-540d-4daf-9b3c-65659e666f3c · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition WavLM: Large-scale self-supervised pre-training for full stack speech processing,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50495929-d525-42b5-8230-3f8f68f3cb4a · outbound
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition SpecAugment: A simple data augmen- tation method for automatic speech recognition,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57f39226-af22-4d2e-bf6f-cf48af4400e7 · inbound
AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a74ba727-e03f-42d1-a4b3-ce8ce35bf26d · inbound
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7dcc69a9-dd97-4fa6-aa8c-2094b28938bd · inbound
Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER) Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.