Pith. sign in

Paper Citation Record · LEDGER

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.24263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24263 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:29.070019Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48403ae3-c436-4c4d-8b1b-b0882126a52f · outbound

This paper cites URL: " 'urlintro :=.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.126963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.126963Z digest=sha256:5b1e01eb8519f34b78be45a98d60d51df254c04b11f6b707268b7fc26709e30d

Observation 552d78f0-3bb1-4696-8fe9-f0b5a18c4c1a · outbound

This paper cites write newline.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.285963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.285963Z digest=sha256:8209dc2184d73b7c3d2b0aea8ba44009492f5a70d41fd8948ba14fb24424561b

Observation 381ca8e0-ebc8-45b7-beb2-c466d9774410 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.410371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.410371Z digest=sha256:37a3cdf156b65e4b68747b67671b5df1914c1103235c4472ba1a6b2e06676da9

Observation 5d579d1d-a646-4908-bd15-bd401705277f · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:30.205471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:35:25.533380Z digest=sha256:7907fe6639801740a27cd73bf1e73962c94ec6bb5f1be2ab190c44d322fb1918

Observation ac6a04c4-0f96-44ff-a4d0-5d758d64bf20 · outbound

This paper cites Language Models are Few-Shot Learners.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.691779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.691779Z digest=sha256:6aff98f865aa31bb3fdb02d3ac82eae155e5692362854a7879db9900e16db611

Observation 7d2be1ac-92ae-4e22-9908-3792e513e538 · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Quantifying Memorization Across Neural Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.837301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.837301Z digest=sha256:f5cf2eaf2e78fea399b9c603fb714b30adb39c03cd1400dbdc4f24a2876ad5dc

Observation 1faed743-79ba-435e-a9d9-a5ba3cb31840 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.945153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.945153Z digest=sha256:b4ea5f25483220953c48e0919f3839a3dfc0e9e5b977926bcdb28f095e2fad39

Observation 6fb46aeb-d547-4ee8-8969-fc01656f5fca · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.047017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.047017Z digest=sha256:230791720766f5a3379cb3299004fdd71d0be7cd4216318ab3db192025a2f1f7

Observation d447cf9b-c72d-4059-a0f0-3344f8a8dc83 · outbound

This paper cites Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.123753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.123753Z digest=sha256:4423201e39a0bfaed3b3b31d2e47da799a8fc18119af092e4dacda2b23fc1aff

Observation 9b4366e9-05f8-49cd-a94a-145685687d6d · outbound

This paper cites Are We Done with MMLU?.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Are We Done with MMLU?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.208319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.208319Z digest=sha256:a3400163950933912514726de3230456220d151e5fd22f60c3cb636c2ecabadf

Observation 65d8c614-fee6-4973-98cc-704b3e9632b4 · outbound

This paper cites Time Travel in LLMs: Tracing Data Contamination in Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Time Travel in LLMs: Tracing Data Contamination in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.309692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.309692Z digest=sha256:aa1c4a7eab752cfad4d77709002678506c3d78084a6cf608b80e9655c26db0f1

Observation 4f148745-1121-4a19-8c0b-296d37a97674 · outbound

This paper cites The Llama 3 Herd of Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.414783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.414783Z digest=sha256:389148121c14d7948c9292ed1d606852b1403fcc9337a3f4ef4dd115f3ef52c9

Observation f015e85b-9744-4278-af9d-5a08ebfe5715 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.515744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.515744Z digest=sha256:d0be9d8f58fed0201672ef10cd53b2e666ea309f9eefeedb26eb77a95b9c6802

Observation f872dafe-c61b-4cd0-b962-c71c5ba91f46 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.585502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.585502Z digest=sha256:4531f39cfbd4254d4ffe938eb919cb3ce0fb3f2418374714142a60cc05643dad

Observation 9fa753c5-b81e-466b-9385-c969a3c49ecf · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:35:30.011952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:35:26.693692Z digest=sha256:e03bb7e924f924909890479b8f81f504762c71a4e9636c831a7d562ecba640a9

Observation f23ba1cc-33d7-4198-8548-4725c998a439 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.812818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.812818Z digest=sha256:c59db28848be12fd26a82b26139dbd5e0b8a0e1ec5c0f59f74a0703e8f24acf6

Observation 088ddac3-4d2c-4796-97d3-8db2db9abe7b · outbound

This paper cites Membership Inference Attacks on Machine Learning: A Survey.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Membership Inference Attacks on Machine Learning: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.911572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.911572Z digest=sha256:eb4e97a5d68a5c96c117b94bc8ceb75c0b56931c38c614edd753806f018a2698

Observation 3a55477a-f0cc-4719-8e26-4ed2979c2896 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.977795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.977795Z digest=sha256:119526328a74358580c476c75974b5d213aa93028cec96b6615c8fff17a7f007

Observation c0af4880-be2f-4f7f-b237-dd4aed103c85 · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.076291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.076291Z digest=sha256:85421e13663ef20aed5f0d42c1c47572bd3d82ac27832658370a1e35ec7fe602

Observation aca29502-3b06-4411-898b-37b2eb2c6e40 · outbound

This paper cites Lin, Jacob Hilton, and Owain Evans.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Lin, Jacob Hilton, and Owain Evans

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:35:29.821277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:35:27.190607Z digest=sha256:38d4efddd0b3d4f2b057c1f8a6c6fb6a5da2cf17da184791a3c9046db3d356b5

Observation e29cbdeb-17ed-44ca-be0f-87b2bd22759a · outbound

This paper cites DeepSeek-V3 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.321284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.321284Z digest=sha256:f860528aad2d7b66e816dc236ae92865c8c4b1ec082d664818dd3a960f218333

Observation 537db04c-fc95-4991-b246-5bc604ff91bc · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.401062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.401062Z digest=sha256:f85d16204c82fcf33753956497cecd1457bf32009bdf5944e123ffc2189646b9

Observation b4c6546f-481d-4c0f-a8a6-1c9b078a59b3 · outbound

This paper cites Training on the Benchmark Is Not All You Need.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training on the Benchmark Is Not All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.501910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.501910Z digest=sha256:fcef29dd7c027d48dab7f0d4dc10ad896ebc053e5ccf5a2a9fedcd312319af17

Observation 022bc4e4-6bca-4107-a61d-e4d2afd41e0d · outbound

This paper cites GPT-4 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.657313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.657313Z digest=sha256:02ddfe100bdcf56d5bfff7007941ffa712c6fbdc2aa2df6a5ba20569d2bffbe2

Observation 85948a63-eefd-4c53-8708-a7d5933f2fbb · outbound

This paper cites an unresolved cited work.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.763705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.763705Z digest=sha256:33ad29600695973c3f7ea746e6af451ab6b674decf7775567dc2ac9a0632b61b

Observation 457c0c24-c0e0-4a64-bfc4-7e22fc3024c4 · outbound

This paper cites Qwen2.5 Technical Report.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.876148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.876148Z digest=sha256:b36a7be89095705ad45aedf191df112036f297d282b2ccb89b74b1d44a7f8cb7

Observation c8d45bf4-cb03-4b7e-b8af-fae0dddd9161 · outbound

This paper cites Leveraging Large Language Models for Multiple Choice Question Answering.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Leveraging Large Language Models for Multiple Choice Question Answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.989088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.989088Z digest=sha256:229ebd57d376f22cc9826743b58d3a984d04f4dbb0196b999077f3ec37b2ca2a

Observation a1645894-ad5c-4459-87c3-38c1f24c1685 · outbound

This paper cites Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.145036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.145036Z digest=sha256:ed419c97f48a3cc2e53ad1847315871e2027e75caf7d311f80baa1b3892a2fed

Observation d4a2cc5c-414e-4f72-93f1-84b947cfecfd · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation SocialIQA: Commonsense Reasoning about Social Interactions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.259215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.259215Z digest=sha256:cb1524719c4bc8c3297ddf5347c8cdfc43e9baa00899d010ea1f9150075791e5

Observation 877f8036-0eb7-428d-b4b5-b2904ec2975f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.402129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.402129Z digest=sha256:73401030d5f56d3e33a6d615ea5e1ec7caf24cdf4e1026a38661d8a9a082b802

Observation 17509d7d-145e-4f20-b4c0-cafb062a0914 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma: Open Models Based on Gemini Research and Technology

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.488255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.488255Z digest=sha256:d03fbaafd9f82d70d953b6584898ba49745b4a3432487f8c767e7a47acfdadd7

Observation c169a4fc-480e-431e-9945-25576704a8e1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma 2: Improving Open Language Models at a Practical Size

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.599455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.599455Z digest=sha256:cc7cfcb483bb327812dc34e78eb1296de9887431ee1ddb4c63473bcaef8423d1

Observation d931bec7-4dd3-40f8-b778-8aef1d04f872 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.705768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.705768Z digest=sha256:6febe35a69db8f085c9112d5312df613ad3558d00357826e10a3c3dbf7fd47ce

Observation 7265672e-7d0f-4ee7-aab8-3d9c9f00974d · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.812524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.812524Z digest=sha256:72d3002b0f25355da4c3614cc6da8d9584980c65e0da38cead2a1afdcb95c53a

Observation cc895bca-1fb5-4f50-872f-432387d4c5bd · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.896218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.896218Z digest=sha256:bf7e3d4a5fb943cac52373e441b46bab408c9cf2bcefa856e3bcbfb6b297cb35

Observation cbd67ca9-a106-43a1-a3b5-d6a351c50b4f · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.982799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.982799Z digest=sha256:da825585864ce67c7cc0b0d48436f1518cb446331771d3330f44471ed56eb04d

Observation dfe21523-ee1b-44fc-b0db-954c381c2c42 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.070019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.070019Z digest=sha256:060a71598938e64e7dd13a6be99aa7d3372d76d4d4437683102ca820a4b83bb4

Pith citing papers

No inbound Pith citation observations are available.