Pith. sign in

Paper Citation Record · LEDGER

AI Benchmarks and Datasets for LLM Evaluation

As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2412.01020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01020 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:49:34.499531Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:07:56.287352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:28:44.552688Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959064f1-1cef-4252-906d-e03d12c2d360 · outbound

This paper cites https://aisafetybulgaria.c om/.

AI Benchmarks and Datasets for LLM Evaluation https://aisafetybulgaria.c om/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.494189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.275900Z digest=sha256:96f16e57850dfb64cb788d4021a6ade08edd2c5a223a6ca4a403511326011be5

Observation 768576e5-db86-4bd5-b740-1828ab195762 · outbound

This paper cites Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems.

AI Benchmarks and Datasets for LLM Evaluation Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.481294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.280630Z digest=sha256:88de3e12f895962dd32607063613fa8dd63dbb835b78caee0e0b4243b23820d7

Observation 595bf9b7-49cb-428f-8c43-23511f8d17b0 · outbound

This paper cites https://compl-ai.org/.

AI Benchmarks and Datasets for LLM Evaluation https://compl-ai.org/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.468167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.285007Z digest=sha256:3891122450f19100961b022b88ec24a79592a276fb1067475faf1d64cb9ec6fe

Observation b852ae81-e737-4721-831a-9bfe82e099b8 · outbound

This paper cites Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance.

AI Benchmarks and Datasets for LLM Evaluation Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.289369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.289369Z digest=sha256:73602ffa15397ba5206722b760dc6ea6756ab5b0486c4ecd48d5ca221934b653

Observation c5a99daa-ed8a-4c81-b982-79c205a95d25 · outbound

This paper cites Robustbench: a standardized adversarial ro bustness bench- mark.

AI Benchmarks and Datasets for LLM Evaluation Robustbench: a standardized adversarial ro bustness bench- mark

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.455802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.293794Z digest=sha256:c91e2093839f8937647099b9439874c073455bb396ee33c651d89bff8b2ff86a

Observation 75850ed8-80ce-4434-8daf-38b73a9c410e · outbound

This paper cites https://artificialintelligenceact.eu /the-act/.

AI Benchmarks and Datasets for LLM Evaluation https://artificialintelligenceact.eu /the-act/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.442610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.298264Z digest=sha256:316e4cae170261ece872b1d454f013412502ccf3d4c4a25ea7deb5719ce794a5

Observation b267cce2-1bd9-4cfc-9b5d-b9ab5095e816 · outbound

This paper cites https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai.

AI Benchmarks and Datasets for LLM Evaluation https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.429455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.302992Z digest=sha256:8d12a1add69c43f764c83d3e522fe3670d02585744e40c76cc2dd2020ca36853

Observation 1232386a-9562-4f1a-a896-34461e041741 · outbound

This paper cites Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024.

AI Benchmarks and Datasets for LLM Evaluation Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.416377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.307310Z digest=sha256:ae1a7e440304d97c3120cdc059e9a151f88a879e3d2adc209781e528c77bf81e

Observation ba014203-b1b9-423d-a334-fd5ee4d2adb7 · outbound

This paper cites Measuring massive multita sk language understanding.

AI Benchmarks and Datasets for LLM Evaluation Measuring massive multita sk language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.401262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.311537Z digest=sha256:54e283062ca379242a686f04be6f44b397771ccd376686949ff8f8ac0bb9da7d

Observation 5d3b5543-7ead-435d-8400-7e2da9421296 · outbound

This paper cites Measuri ng mathe- matical problem solving with the MATH dataset.

AI Benchmarks and Datasets for LLM Evaluation Measuri ng mathe- matical problem solving with the MATH dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.386821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.315694Z digest=sha256:2b11af12bdde159373c5efdfb35a621454137c7670e2c268c3db8df0ca1c2245

Observation 8dcf9c9a-e335-440e-92ff-49c4d3a6b235 · outbound

This paper cites Weld, and Luke Zett lemoyer.

AI Benchmarks and Datasets for LLM Evaluation Weld, and Luke Zett lemoyer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.373656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.319762Z digest=sha256:940d5f7feec2d5b6b6279f3c042f5662e76f901c85de901784ed00d3468e61e2

Observation 31d31134-2521-42bc-890c-fcb31d6d2140 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

AI Benchmarks and Datasets for LLM Evaluation OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.323908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.323908Z digest=sha256:569be9f8620f3c2a25bed48c2ae1a6703f52360ed62977b5704c84ede83d2d9e

Observation 10e6adef-20f7-49ad-8eda-32d82c7a0935 · outbound

This paper cites Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024.

AI Benchmarks and Datasets for LLM Evaluation Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.359667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.328632Z digest=sha256:f9986ca030bc1bb703ab3d6f4aac12565ed4535489721628bbe88632b0433cab

Observation a3ab6eb7-bfbb-4464-ae23-cf7c389c373f · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.346655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.332779Z digest=sha256:a969f4a78f1efd4713ced0f793541117c10f1b9747608e7a5969f86b2357a0fc

Observation b024e696-d46b-48e5-baa3-d74efc343611 · outbound

This paper cites GLoRE: Evaluating Logical Reasoning of Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation GLoRE: Evaluating Logical Reasoning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.337026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.337026Z digest=sha256:3db6876c74323bd7379d620f2bfdc3b4193f4a6b4218a0bf2f0c4e0c7778c364

Observation f9509a7b-ff82-44b9-8da8-450226793d40 · outbound

This paper cites Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning.

AI Benchmarks and Datasets for LLM Evaluation Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.333475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.341499Z digest=sha256:a541486853fe017d8bce0c47eb181de223e37e117bcf6731bc877952f87bf12c

Observation 4ab148b2-0aa1-4d3d-b87b-0a3b5c1f0b6e · outbound

This paper cites Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019.

AI Benchmarks and Datasets for LLM Evaluation Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.321026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.346322Z digest=sha256:05658622c9fb77b47867382b0a27c7fd2661810fbc7cb669c7cdbace2166edd4

Observation 0fd44be2-0d82-4afa-b5ea-4905c02a7fc3 · outbound

This paper cites Abstractive text summarization u sing sequence- to-sequence rnns and beyond.

AI Benchmarks and Datasets for LLM Evaluation Abstractive text summarization u sing sequence- to-sequence rnns and beyond

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.308111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.351220Z digest=sha256:cd26e2a01a8ed82f7ddeb5f848917e80f48c2394982bd7bded99f55e7ba1432c

Observation 73b65f13-8f29-4de2-b2bf-a0fc493bcc68 · outbound

This paper cites Adversarial NLI: A new benchmark for natura l language understanding.

AI Benchmarks and Datasets for LLM Evaluation Adversarial NLI: A new benchmark for natura l language understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.293849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.355412Z digest=sha256:481ea61c7fce7980b1ca8fa9d92e93bbe0e6126c1719699f251e4c479cc893b5

Observation caa6d431-23da-4cb6-a193-00024761feeb · outbound

This paper cites https://oecd.ai/en/ai-pr inciples.

AI Benchmarks and Datasets for LLM Evaluation https://oecd.ai/en/ai-pr inciples

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.280799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.359484Z digest=sha256:358d220b0d3fd741c0f7e3f2ac576d70932001d2ad7b4e8ea43700cdb5f9f9f6

Observation a9363cd5-ffaf-4dcd-91b5-f9b2e87961e3 · outbound

This paper cites The LAMBADA dataset: Word prediction requir- ing a broad discourse context.

AI Benchmarks and Datasets for LLM Evaluation The LAMBADA dataset: Word prediction requir- ing a broad discourse context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.266378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.363678Z digest=sha256:ccc55d2e4af0a301215185c555f2764aa09282995456af3d112b215298ecbfa7

Observation de3aa830-3499-4d37-b28c-7f1490d7eda4 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

AI Benchmarks and Datasets for LLM Evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.367867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.367867Z digest=sha256:b12c6cd0beb70c32e52cd1d0bc726ef8cc8c4d6f7001cdfd3da31d04f5bd6012

Observation 32cc2676-71cc-441d-a613-b2dc3c66fdad · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at sc ale.

AI Benchmarks and Datasets for LLM Evaluation Winogrande: An adversarial winograd schema challenge at sc ale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.251863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.372425Z digest=sha256:31ed6ac3e3837027a56f9ad549797075eb9bfd9d17ee60e34dbab30bbd69fbf0

Observation 6348edc0-763c-4d98-981e-065481892f0a · outbound

This paper cites Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F.

AI Benchmarks and Datasets for LLM Evaluation Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.239019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.376471Z digest=sha256:3462af7786d57fa98caaf7ab58c46e6750d98b626dd915a5a376b0be4dca09a0

Observation 141e0544-db56-4537-8ad7-db555dc02946 · outbound

This paper cites Sur- vey of different large language model architectures: Trends , benchmarks, and challenges.

AI Benchmarks and Datasets for LLM Evaluation Sur- vey of different large language model architectures: Trends , benchmarks, and challenges

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.224479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.380505Z digest=sha256:99c912333a45b6996d2c8f8badcfbf183d396b5f35c2d83d8c1b7661770ee02d

Observation 5c0eef8c-bb5a-4fe3-8961-57df44b03318 · outbound

This paper cites Concep tnet 5.5: An open multilingual graph of general knowledge.

AI Benchmarks and Datasets for LLM Evaluation Concep tnet 5.5: An open multilingual graph of general knowledge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.211305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.384428Z digest=sha256:b5e3d5dba14b6ec6e3756230340968a9c5102f3f8e40f98ca9c91bdb9d686cf6

Observation 8e2cb9b3-8344-4777-9005-28c997947f30 · outbound

This paper cites Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing.

AI Benchmarks and Datasets for LLM Evaluation Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.198454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.388365Z digest=sha256:d6f4ce4f6a311c0a8dd0696577e4a577ec9c91840a1d520cf853b708149030ae

Observation a20935e0-0fbc-4411-88fd-e8cd5c050af2 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

AI Benchmarks and Datasets for LLM Evaluation Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.185345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.392518Z digest=sha256:cad4a03c75c59321132e1d775502f30e22eecc0e053ba55e653ed3678d4f0dda

Observation cf6d6eb4-7de6-4ec0-b3e1-a900a43a10d8 · outbound

This paper cites A corpus for reasoning about natural language gr ounded in photographs, 2019.

AI Benchmarks and Datasets for LLM Evaluation A corpus for reasoning about natural language gr ounded in photographs, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.172027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.396498Z digest=sha256:2ead87d4b11279621a7223e4db80494dd26e856bcf2aa0b2c5e04555a6f7fb54

Observation a87cdc31-c0d4-4d97-9ffc-8a2595dc2b75 · outbound

This paper cites Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024.

AI Benchmarks and Datasets for LLM Evaluation Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.157059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.400470Z digest=sha256:e94dca3be8e6bd3858bbbbfed040586c9bc3a069829dca7bc85b8e5a2bc7b237

Observation 436f95fd-5862-4c93-964f-41ceb378d71f · outbound

This paper cites Le, Ed H.

AI Benchmarks and Datasets for LLM Evaluation Le, Ed H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.142715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.404516Z digest=sha256:e13fe25272e67926ae60fba3c24d465390d8cbf5fbc27e30fd3aceaa39172388

Observation 5870f8e5-7fd8-4410-af03-e0a4673eb9d5 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting c ommonsense knowledge.

AI Benchmarks and Datasets for LLM Evaluation Commonsenseqa: A question answering challenge targeting c ommonsense knowledge

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.128940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.408653Z digest=sha256:b20d29db819146346b988a61e46f9bc51acc9f9d893af3bc90b7b97a9d6f7187

Observation da8f0378-5315-42b6-87ca-298895d91458 · outbound

This paper cites Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls.

AI Benchmarks and Datasets for LLM Evaluation Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.412604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.412604Z digest=sha256:a2c414183f3b3c5027f143ce3e8a6385eb62d6b7d994c7ea30d815422fdc0bf8

Observation 954456a9-cbf4-4331-b187-94545c406c76 · outbound

This paper cites Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C.

AI Benchmarks and Datasets for LLM Evaluation Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.115471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.416422Z digest=sha256:427c06ac7cb4456f59bc50405b6e1cc13bbea9806bfc84dc90c8ad1c1cfaee45

Observation 25f27a24-2548-4309-b237-0cff8158c062 · outbound

This paper cites Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V.

AI Benchmarks and Datasets for LLM Evaluation Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.102325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.420387Z digest=sha256:c5b921a2882074c7355a94a10b3f4a373ad2e2130cb1547e59e7e100ca888bcd

Observation 64a38680-7320-44cf-b4b3-208c97dd72fa · outbound

This paper cites Introducing v0.5 of the AI Safety Benchmark from MLCommons.

AI Benchmarks and Datasets for LLM Evaluation Introducing v0.5 of the AI Safety Benchmark from MLCommons

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.424309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.424309Z digest=sha256:167f9d44db604ba6a3f20087a25356ec20504b08c10d2bffcaa05d3a80e4574d

Observation 8ef0f9f7-90d2-4a76-8294-9d88f86a62e0 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.089272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.428747Z digest=sha256:9990a0a320fd153edc8c2e083c5eaeecbecf1057398333ebbd46ad44f79d7565

Observation dd9703b6-3513-4e6c-b93a-19a841bcb088 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.074927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.432590Z digest=sha256:9627a740580b95739bf37818063e30ab8149c04c7b068c7307cef5b16e62c14e

Observation 9694d149-2b58-40a5-b696-3d6f82c39ce7 · outbound

This paper cites CORD-19: The COVID-19 Open Research Dataset.

AI Benchmarks and Datasets for LLM Evaluation CORD-19: The COVID-19 Open Research Dataset

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.436441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.436441Z digest=sha256:e4c6617b39bc3e19ced43c0aed0a1acbdf341b9442355d58cac6044075276a08

Observation 12ca922d-2dcb-43e9-8888-3c4b06c828be · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

AI Benchmarks and Datasets for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.440600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.440600Z digest=sha256:92987ccfa0ae881728d2f526c4e9a9124022880b1ad5cd0d312b935e90432236

Observation 758a8961-3d6a-4944-b487-63ddebaff14b · outbound

This paper cites CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models.

AI Benchmarks and Datasets for LLM Evaluation CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.061640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.444965Z digest=sha256:669a1f0aabe4f34dc448eca0c800777c0f3a020f07cc72252bd18e932a0c9d4a

Observation 8c69c960-5975-41e9-98fb-689d2e8a499f · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

AI Benchmarks and Datasets for LLM Evaluation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.448945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.448945Z digest=sha256:74e6e9c06304e601d9ed6c2386e3c8f2914c107828c496ce1236bc36c4fe3d3d

Observation 8a57a991-f3da-4a06-8e1a-a1c0b5aecbbe · outbound

This paper cites Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering.

AI Benchmarks and Datasets for LLM Evaluation Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.453224Z digest=sha256:62a9e2ff7d1cfd5a050c38202d9abf2fe086c0bbb35cd02733f64cc38e49d265

Observation 5e7dfec3-743e-430c-9517-9d9a8d40c46b · outbound

This paper cites Eval- uating the quality of hallucination benchmarks for large vi sion-language models.

AI Benchmarks and Datasets for LLM Evaluation Eval- uating the quality of hallucination benchmarks for large vi sion-language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.457155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.457155Z digest=sha256:414e04c5c6e0537adc85897d11bfa5408e7f452eff3ca703b71e5be6075ce63d

Observation 20f81d44-da0d-41b9-ab53-27ec1ca13751 · outbound

This paper cites https://z-inspection.org/.

AI Benchmarks and Datasets for LLM Evaluation https://z-inspection.org/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.032335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.461684Z digest=sha256:f4a8fe5ff492aba011185b32b1ce9ad2d636d9121bef748d2f921dde492cc198

Observation 42407999-af0d-4843-9c5c-43c813337483 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.018580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.465655Z digest=sha256:9171c03688104fe3a56bf000c005871be388abca8032c201879e23bebebd0ee3

Observation 8cc604dd-0a84-4bfb-ac5e-869ec535b2d4 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R.

AI Benchmarks and Datasets for LLM Evaluation Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.990279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.474079Z digest=sha256:29a17c9701b86cc4dc80c269f38ef37376fee1179ab6c8151bc33ba853868dca

Observation 79e80c6a-d471-4375-b5de-162f8947408f · outbound

This paper cites MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.478030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.478030Z digest=sha256:0395331798439a61a359ccd94f27d7bce53148f0e51142090c0db6d2dd5c396a

Observation 3f429688-6ded-4117-9a01-d6e88f1f15cd · outbound

This paper cites Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024.

AI Benchmarks and Datasets for LLM Evaluation Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.976392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.482834Z digest=sha256:029e2d383c0e5d30e1a7c4d1839efbc4985f42949fd74bbdada7a78a521e6ec8

Observation bb8602bc-0507-49ad-b393-0c6d3a0fdaff · outbound

This paper cites Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024.

AI Benchmarks and Datasets for LLM Evaluation Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.961429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.486810Z digest=sha256:00249a764af8f2d601e780391077c39ac9b3878e491ab7aae9087a4833505728

Observation 2207480c-5c0f-4a4c-9051-22aba14151ab · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation Instruction-Following Evaluation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.490762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.490762Z digest=sha256:727560cf6e787ca237d060529df6c283b0be2642dd98180a4613518e607c8251

Observation 56dd0cbf-1cce-40a4-965c-225238405a70 · outbound

This paper cites CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs.

AI Benchmarks and Datasets for LLM Evaluation CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.495253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.495253Z digest=sha256:2a880e7626e345c0f412ae775a1c10a15fc2995d267440a5a47dd1296ebcadfe

Observation cf60c4e1-4baf-46d8-b430-995295a8caff · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:34.946685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.499531Z digest=sha256:841b3edd62614c6bb6506d9983abd155416f2e8f24f8aabafd269c96cc4ed96c

Observation 3c1835ad-51b0-4747-97f1-e0064a538d51 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.004190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:49:34.469869Z digest=sha256:0b039013ed98509fece68b13d8a93c5a9123391660bb11108fb62de94d9d388b

Pith citing papers

Observation ddab16a9-f941-4960-9d2b-7ecee0ac701b · inbound

Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning cites this paper.

Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning AI Benchmarks and Datasets for LLM Evaluation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.554094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T04:07:56.287352Z digest=sha256:14dd451036d0c98e7922659b3b866ad6c6ee053b973aa1b31340b49f8fa525ee