Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:56.441502Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 10 inbound Pith citation observations for arXiv:2412.15194.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:56.441502Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:29.508961Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
54 of 54 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation f0890a7c-e109-4e22-94aa-ed6a2403c747 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Phi-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26b5a11c-43e8-4dfc-a1e8-91a3c016e4ad · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce5e8b0-b892-476b-9ae6-767594d0a466 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark GPT-4 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58e88894-dbfa-49fd-845c-31fd7293e841 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91300db8-8176-441f-b4af-564e235d1e33 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mmlu-pro+: Evaluating higher-order reasoning and shortcut learning in llms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b6bb381-41e9-4fe8-aafc-ab2b2269ec75 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Qwen Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6017d45-8fb9-46c1-b964-1ccb76709906 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark InternLM2 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789a33e8-a9fe-49a7-ace9-b86a526694d9 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e537f3-adcd-4bd2-8d4b-e3e5527c7f11 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gonzalez, and Ion Stoica
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 27005d57-f694-4ae8-a7cc-a79ffe4bbd7c · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4aa229-92c2-4e5d-ba14-dd8ef3b069e4 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d890cd-546a-4cb3-a802-2e80e0234596 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a3c1869-3fd5-4951-be65-8c3cf1198ccf · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bca4d33e-5b27-4f94-9232-b6fe7992985a · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Are We Done with MMLU?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405bd8cb-7834-4264-9d64-744e56d71f40 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18d512d1-bc91-4b72-9a74-b05d9b4d891c · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19bbbb6-d29e-45bf-bdab-8992c9615c43 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Changing Answer Order Can Decrease MMLU Accuracy
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8b19299-6703-4008-ae5e-a42b1d25934b · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Measuring massive multitask language understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d789d2-915a-4377-9665-dbffb3a5dfde · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 30223fc1-2011-488f-81ce-38f977fad982 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0961f84e-7a9a-420a-99a2-ebbaf7cf50a8 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mistral 7B
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4277810-f35f-445e-9f87-c5fd122b15b1 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mixtral of Experts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 880b56f8-6794-4cbe-8d17-c9d57c88c580 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3ad207-45b6-4514-83e9-8c2afaffda59 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd19fd2-1054-43db-84b9-b78a72dfade6 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4352b1cd-594d-474a-a045-fb2623415971 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6df5f9b2-41f5-463e-add9-14873c991060 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe50c6cc-c533-40e6-83c5-e5e82b2f2d73 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c243d202-091f-4772-ab7c-ead6da3e5bc2 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af01fa5c-200e-42fb-86e2-e97e5d88f467 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc63022-021b-4b00-929c-376efddc9047 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63284d49-6886-4e5c-b468-f5268d9a7b40 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7234562-76aa-40a6-8fd8-ab989e73170a · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0cde24f-07c4-4f69-8bf9-a7c7692aebfc · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Detecting Pretraining Data from Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63148f0-2746-45a7-a3dd-a1e910d835eb · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gemma 2: Improving Open Language Models at a Practical Size
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f70f6830-4f06-41e5-9bd6-e163fe394dd9 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b175d5b-d8af-443c-82f1-1ad156104a58 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6153dfd6-d71e-4df5-9f13-6244ed606ea2 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017b45bd-a3b3-405d-bacb-17e74bc2b72c · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76973005-a47e-482f-bb3f-1ab01ddfcd28 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18a5227-3918-406d-a583-2a6aaa86fdd0 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Baichuan 2: Open Large-scale Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f10b90-0c52-45e5-b59e-f67d6f0a9ac9 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf5126f-538a-4eff-a437-4f85b734d514 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Yi: Open Foundation Models by 01.AI
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213ca611-061c-4588-a1fb-35235ef37cb8 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf77de9-01e4-4f77-bc08-92db4da8477e · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2e9f8d-ffb1-4668-ace2-09e617f8a2d7 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf53d32-dbdc-4755-b585-f5e1011b97b8 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark xLAM: A Family of Large Action Models to Empower AI Agent Systems
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f1471d4-5cf0-41a7-852a-1882ae3f10af · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efdc6df1-d7be-49ea-b549-5f6873b797ba · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3044dc93-ad54-46d0-a910-2df71408c75d · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c87fd0d-bc73-4c2e-ab97-6e0c3501b2d2 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Instruction-Following Evaluation for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16f36bc-8bae-4637-a97a-d5ff722f0c77 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a973ea3c-e569-408c-9957-a2c7e4a653d0 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark online" 'onlinestring :=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00c10ee9-dc97-4649-8e5b-12b23cd840c1 · outbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark write newline
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 90a15906-3214-4ead-8b2e-11f4b91575ee · inbound
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54cd0d1-3536-4eff-83f8-93f98202586d · inbound
ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e135f8a-94d1-4679-acea-e22753019b80 · inbound
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e689eac6-1808-4201-a231-73f69a312a71 · inbound
TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 309c7c76-6f33-4ca6-bc7e-32a15024d37e · inbound
Weak-Link Optimization for Multi-Agent Reasoning and Collaboration MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ee3fae2-f0b0-40c6-b5a5-b7954c930d0f · inbound
Provable Joint Decontamination for Benchmarking Multiple Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 180
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c8b490c7-7a77-43b4-a1d2-0ef8d2636d42 · inbound
At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cce7df17-337f-413b-8be4-643d567840e5 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e5c36a-ae8a-4a43-a059-dc7cc848603d · inbound
Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468f8bbb-9c8c-448c-accf-13bc3d193157 · inbound
Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.