Pith. sign in

Paper Citation Record · LEDGER

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 10 inbound Pith citation observations for arXiv:2412.15194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15194 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:56.441502Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:29.508961Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f0890a7c-e109-4e22-94aa-ed6a2403c747 · outbound

This paper cites Phi-4 Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.182704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.182704Z digest=sha256:77b7ad62c861c17fbb46a4ed8a4e899fcf374d2c11064454aabad94044092b7b

Observation 26b5a11c-43e8-4dfc-a1e8-91a3c016e4ad · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.188989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.188989Z digest=sha256:5c293aa46ecc689f1acb30ca8e0a8c13bcc56a2597428b75277d6d9d29128095

Observation 7ce5e8b0-b892-476b-9ae6-767594d0a466 · outbound

This paper cites GPT-4 Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.194251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.194251Z digest=sha256:21e40cf23aba815c1070945784efa00e14f1e31afdf6a2082f977c3c931a2400

Observation 58e88894-dbfa-49fd-845c-31fd7293e841 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.195843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.199571Z digest=sha256:75c7c564efd3df386cef74077c6ccac8cd62f8439c0ab62e2ceff03da3ba7472

Observation 91300db8-8176-441f-b4af-564e235d1e33 · outbound

This paper cites Mmlu-pro+: Evaluating higher-order reasoning and shortcut learning in llms.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mmlu-pro+: Evaluating higher-order reasoning and shortcut learning in llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:57.181341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.204262Z digest=sha256:c0e9e647d80656c200240d922b306e8c94e381d42261eaa7cd1d7aaa05ece15f

Observation 4b6bb381-41e9-4fe8-aafc-ab2b2269ec75 · outbound

This paper cites Qwen Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.209191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.209191Z digest=sha256:451362b9a7fbbe692477071292203ef71ea2354872468df630ee3e541d5dde67

Observation e6017d45-8fb9-46c1-b964-1ccb76709906 · outbound

This paper cites InternLM2 Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark InternLM2 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.214345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.214345Z digest=sha256:359db790add0462b2cee692de80fb372b50f9d68649a03cf295f914840d3dde1

Observation 789a33e8-a9fe-49a7-ace9-b86a526694d9 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.219910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.219910Z digest=sha256:bac72b1326fb5eff4e67781757641dd493786496fb9ec3b31682a14d5b6e5b25

Observation c3e537f3-adcd-4bd2-8d4b-e3e5527c7f11 · outbound

This paper cites Gonzalez, and Ion Stoica.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gonzalez, and Ion Stoica

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:57.157066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.224464Z digest=sha256:a6de84013f84d227abcb6503c1177f050c9b374ca910f262da8d1ef08b46e66a

Observation 27005d57-f694-4ae8-a7cc-a79ffe4bbd7c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.229210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.229210Z digest=sha256:6d9c6776e25127ddde3545bd02706b82c3f14a8c7cf9472d38c2b5049237c9c3

Observation 4e4aa229-92c2-4e5d-ba14-dd8ef3b069e4 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.234166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.234166Z digest=sha256:da1b3cadd38d1343bb73baace5140d7f4035e7f66c795c201e8cc5a8cf7682a7

Observation d9d890cd-546a-4cb3-a802-2e80e0234596 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.131661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.238498Z digest=sha256:a2086c92b619dc04f0702024c2b9a938151089a55e135eedce07dd4e467a2c0d

Observation 4a3c1869-3fd5-4951-be65-8c3cf1198ccf · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.116782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.242949Z digest=sha256:41e59fd4f1948473fe4cef1c6f7f529f2628b9f5c39ef0740064e221377dace3

Observation bca4d33e-5b27-4f94-9232-b6fe7992985a · outbound

This paper cites Are We Done with MMLU?.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Are We Done with MMLU?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.247364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.247364Z digest=sha256:ea86b45bed15fc13a0fcdc02eedad06dc9d9be054d385eb496cf571ca195161b

Observation 405bd8cb-7834-4264-9d64-744e56d71f40 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.252020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.252020Z digest=sha256:5797f30455f85f3e6ee1614435ae5ecc7614f6279b4f073765297db735bc6856

Observation 18d512d1-bc91-4b72-9a74-b05d9b4d891c · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.256785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.256785Z digest=sha256:66f2be04c2a665e9fab054d2a5985376149af86c78072e70e15c44e138548671

Observation a19bbbb6-d29e-45bf-bdab-8992c9615c43 · outbound

This paper cites Changing Answer Order Can Decrease MMLU Accuracy.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Changing Answer Order Can Decrease MMLU Accuracy

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.261554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.261554Z digest=sha256:478851cbb392751e75f48a7f8de4894b150fb524bb8d3a4df1c32fc8eff95c11

Observation f8b19299-6703-4008-ae5e-a42b1d25934b · outbound

This paper cites Measuring massive multitask language understanding.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Measuring massive multitask language understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.266772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.266772Z digest=sha256:e47e8109e0cbbdb3bcfe98bb8ab01f7cdf83f9f864a718205c7d283ece7da207

Observation 98d789d2-915a-4377-9665-dbffb3a5dfde · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.090793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.271444Z digest=sha256:b0ba731d0f92caa2910af8cc2a053901ddd9cdc1d69465fe9bdef1bef948a45e

Observation 30223fc1-2011-488f-81ce-38f977fad982 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.276076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.276076Z digest=sha256:2dd0d14cf18f97a39f1ff670ffc4ec02369005ae5c3316331ca213fe526fc346

Observation 0961f84e-7a9a-420a-99a2-ebbaf7cf50a8 · outbound

This paper cites Mistral 7B.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.281024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.281024Z digest=sha256:000f08de62b355a63689a1ac47a308e4a8e824ec54b6293a6bdf5f7c77afc879

Observation d4277810-f35f-445e-9f87-c5fd122b15b1 · outbound

This paper cites Mixtral of Experts.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.286099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.286099Z digest=sha256:07767ec436b8002e6698b0916ec0fa136b3a85adc728ec5161fb1225f637e389

Observation 880b56f8-6794-4cbe-8d17-c9d57c88c580 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.291029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.291029Z digest=sha256:d669529fbbf36647aaab33311879c5c0a2ba27cfeb18a6e149ca4a1d950cf1ff

Observation ee3ad207-45b6-4514-83e9-8c2afaffda59 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.295602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.295602Z digest=sha256:6d56c9641839775bd01cee7de669c042356e44d8b638351e21f1f4c4667a999e

Observation 2bd19fd2-1054-43db-84b9-b78a72dfade6 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.056345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.300303Z digest=sha256:46cfb5ab3587eb35aa00ebe3ac300b4079950194e369ee42d204dd2cafa49ebb

Observation 4352b1cd-594d-474a-a045-fb2623415971 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.040129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.305291Z digest=sha256:2be6dabfe6c6f6c4554570010dc698d3988b91452d92d470a98914af07d62451

Observation 6df5f9b2-41f5-463e-add9-14873c991060 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.310575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.310575Z digest=sha256:684be12fd1bc2ba66d34ef9e1b6c8b91e0e07a8b6ed7b5909df1d445032ac960

Observation fe50c6cc-c533-40e6-83c5-e5e82b2f2d73 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.014963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.315740Z digest=sha256:f1f4be32f5b6be411e770018458d77b41361037ea120b54f883ad74202f62839

Observation c243d202-091f-4772-ab7c-ead6da3e5bc2 · outbound

This paper cites Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.320629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.320629Z digest=sha256:6be6cc3b2c16fb654d98b96ed4dd713ef5d4e03c7094f43135f7157c4f3999a9

Observation af01fa5c-200e-42fb-86e2-e97e5d88f467 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.325452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.325452Z digest=sha256:be550231715015ca26ce5a4abd245661994df6a40ef17d24bda720ddc56a6e34

Observation fbc63022-021b-4b00-929c-376efddc9047 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.330309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.330309Z digest=sha256:dd976a15ca8bd61acfcf0ae730693684abe05a8966af012e584ff39d94888874

Observation 63284d49-6886-4e5c-b468-f5268d9a7b40 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.335120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.335120Z digest=sha256:00be5f7d6d76137ae81fe235eb42f3d247083000581f80ae97834f473a577c98

Observation c7234562-76aa-40a6-8fd8-ab989e73170a · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.339386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.339386Z digest=sha256:da58e606246cc7a31f3b51e4ed7dc26cd865b8216fe1f39ce7a21beb74b70517

Observation a0cde24f-07c4-4f69-8bf9-a7c7692aebfc · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Detecting Pretraining Data from Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.344809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.344809Z digest=sha256:157868354bd267be9ec1d5a9312c60fbd0ced6dd5aea7b898c8c1ca5efb653e1

Observation d63148f0-2746-45a7-a3dd-a1e910d835eb · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gemma 2: Improving Open Language Models at a Practical Size

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.349638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.349638Z digest=sha256:3c0ef3f2a57e8abeec75c2d662b4758eb5b0bd4b5f1a6bcbbcd6715b60a7455f

Observation f70f6830-4f06-41e5-9bd6-e163fe394dd9 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.354281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.354281Z digest=sha256:3400eaddd31dea22506283a711f88d35b7bb5718e384a319e553d5b2b8b9acf1

Observation 5b175d5b-d8af-443c-82f1-1ad156104a58 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.358716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.358716Z digest=sha256:1e90f04f3545680ab548b2320528e7f265c1f4b5263c316c0e7858d6984fbbf2

Observation 6153dfd6-d71e-4df5-9f13-6244ed606ea2 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.363525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.363525Z digest=sha256:51ac7b36ab1220e0a39ef3e48015ad6080c236a13a83ccaf69c1403db919d524

Observation 017b45bd-a3b3-405d-bacb-17e74bc2b72c · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.368110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.368110Z digest=sha256:10c8e3a249ec3ef85c3db00e0196a5d3e58a0dd6be7518310fed95b0bea879c8

Observation 76973005-a47e-482f-bb3f-1ab01ddfcd28 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.373073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.373073Z digest=sha256:c16694e6d33d01515338b21a89adf87c2231c65dc2ed4a877af5624a2682b607

Observation e18a5227-3918-406d-a583-2a6aaa86fdd0 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Baichuan 2: Open Large-scale Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.377964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.377964Z digest=sha256:2d0d57b55cddc425e55b81a37d5a84ac552f5a51cf20a5ba484d2f7a4de09519

Observation 08f10b90-0c52-45e5-b59e-f67d6f0a9ac9 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.383219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.383219Z digest=sha256:fdcb349c7d9dfc78b8cd888e3bb3608d465661778490a1c741828ff9b7e1a08a

Observation 3bf5126f-538a-4eff-a437-4f85b734d514 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Yi: Open Foundation Models by 01.AI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.388284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.388284Z digest=sha256:6aa6a5ad363be99654953e048889d2e1e6a3c4e850586e682eecb04de267f0ff

Observation 213ca611-061c-4588-a1fb-35235ef37cb8 · outbound

This paper cites WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.393090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.393090Z digest=sha256:995af9de83f1485eb269355a9c59247db89e70eecba61973cb737c58b1556075

Observation ccf77de9-01e4-4f77-bc08-92db4da8477e · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.397912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.397912Z digest=sha256:ca9d1b3edd10a5305e40d70f7ec2009e7c8021c942f3c549d7365916108ea920

Observation cd2e9f8d-ffb1-4668-ace2-09e617f8a2d7 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.402576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.402576Z digest=sha256:22a00a81f69d5b117853d38950639f3a3f31a5c267e4fa22176467eec65212ec

Observation fbf53d32-dbdc-4755-b585-f5e1011b97b8 · outbound

This paper cites xLAM: A Family of Large Action Models to Empower AI Agent Systems.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark xLAM: A Family of Large Action Models to Empower AI Agent Systems

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.407462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.407462Z digest=sha256:e222724fb5ed7454f8fbea3b3609e6392ce8276ef9c803506b168ce757364ecc

Observation 4f1471d4-5cf0-41a7-852a-1882ae3f10af · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.412708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.412708Z digest=sha256:d3b218cd6a7ce01a21e634698bf0bfaba93657e5e569caf16b327f1f08185b9b

Observation efdc6df1-d7be-49ea-b549-5f6873b797ba · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.417153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.417153Z digest=sha256:1545475554313ebf89e4031164a5a1cd1089e0118e0176a242f9ef995c82db99

Observation 3044dc93-ad54-46d0-a910-2df71408c75d · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.422014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.422014Z digest=sha256:75b3e03475d1331f5e6ef7c7f4fc095c14f66a85a2e58f081db0de760e92d0c8

Observation 5c87fd0d-bc73-4c2e-ab97-6e0c3501b2d2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Instruction-Following Evaluation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.426938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.426938Z digest=sha256:8f6f1b97b77957aa8f1a578781c5a22d41a0ac958c1f52dbeb180347eb2a9844

Observation f16f36bc-8bae-4637-a97a-d5ff722f0c77 · outbound

This paper cites Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.431608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.431608Z digest=sha256:960755bec4441c36e10ed8af7340cede66f4ae555f3bfd52167dbc3605b30de3

Observation a973ea3c-e569-408c-9957-a2c7e4a653d0 · outbound

This paper cites online" 'onlinestring :=.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.436392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.436392Z digest=sha256:543f1d1717cf5ce044d07c7ab81162a4532487d42cb09ad816433299f5917c4d

Observation 00c10ee9-dc97-4649-8e5b-12b23cd840c1 · outbound

This paper cites write newline.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark write newline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:56.933825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.441502Z digest=sha256:3c4c523bbe7de8849e8bbeb2540f59457bdcc994d5ba73520190faac3848dc80

Pith citing papers

Observation 90a15906-3214-4ead-8b2e-11f4b91575ee · inbound

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism cites this paper.

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:29.508961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:17:29.508961Z digest=sha256:82601158b921da1fe117e75cc6e3a885fedf31393f82808e4c333e67b0668412

Observation a54cd0d1-3536-4eff-83f8-93f98202586d · inbound

ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities cites this paper.

ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:26.145226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:27:26.145226Z digest=sha256:196cc5aaee48c25c33890168ec3503aab13e28667171727797d3d940500807ac

Observation 5e135f8a-94d1-4679-acea-e22753019b80 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:47.983056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:47.983056Z digest=sha256:d348ed40ad6a5b8ac3aa584297e27254eb0322853c6910840936682104a78dad

Observation e689eac6-1808-4201-a231-73f69a312a71 · inbound

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models cites this paper.

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:36.032153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T12:21:39.237267Z digest=sha256:d4afb2d44dae979730b5b8f6759272d2373d44b7a197131831150b786b221585

Observation 309c7c76-6f33-4ca6-bc7e-32a15024d37e · inbound

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration cites this paper.

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:13.183812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:53:03.420926Z digest=sha256:17e408fb7aba2f3c98b9685156492b1289e8d7984e3f8d81703f5e9eeedbe28f

Observation 3ee3fae2-f0b0-40c6-b5a5-b7954c930d0f · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 180

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.471191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:94d154cb54b5e02c81aa70fe1f007abf92ed468e6ab67ba55be049a09f716b37

Observation c8b490c7-7a77-43b4-a1d2-0ef8d2636d42 · inbound

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization cites this paper.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T01:28:50.524986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T01:27:39.812228Z digest=sha256:1f92c6343d1cfd7a91548c5949a8c34846caeba5a0c275a2b5292bdad8953752

Observation cce7df17-337f-413b-8be4-643d567840e5 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:4cdacdb60dac0192c33c96c2862c3a6d4b227f4e8786e08290a6d4480d20c2f9

Observation e9e5c36a-ae8a-4a43-a059-dc7cc848603d · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.871998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.871998Z digest=sha256:540416886f24c685b7c07623ecd270f0121c3b43bdb88a559c9e5ae570b89d32

Observation 468f8bbb-9c8c-448c-accf-13bc3d193157 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.646103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.646103Z digest=sha256:8dcf347e2ebc8c46fe6336694665abddb92c8e82be49368e5ac4b20d53c5dbfa