Pith. sign in

Paper Citation Record · LEDGER

Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2402.09880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09880 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:06:55.090364Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:10:41.039569Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a670cd0f-1687-4bc1-96ff-f2dccbbd6bbf · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.042123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:d556e3c740f620a89f46599afa36500046de6de5724a655692fe7321fce8e649

Observation 2f38d407-5bf1-4fe6-ad94-b89fce08349a · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.396135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:38b3a6fb7e72a2b4cb2853c070e497a289fd4a2deed3a80993c2025bd8638781

Observation 58a662b0-38db-4b24-8cf1-35d1af3a5c68 · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.090364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.090364Z digest=sha256:452fef0a7a609cdb2418320f150c16da12db54666306d8d41f7357f83c499fc5

Observation 46ecd899-cc2d-42fc-91b0-27259ba69af6 · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.436864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.436864Z digest=sha256:a569353a8c1b505de8892e555aaca48588c05dd9202d4b2d1fab56e2c202c60e

Observation d00f127b-f3c5-4b99-b092-d7e1cc5deca9 · inbound

Evaluating the Sensitivity of LLMs to Prior Context cites this paper.

Evaluating the Sensitivity of LLMs to Prior Context Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.197830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.197830Z digest=sha256:9708824f42df4fd24dfbc23bd3a6b8aa6915bd4941711ffc1c1f2004f9434697

Observation ac7fea2d-36bc-4c4a-96dd-4c3d6f2e3949 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.091271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.091271Z digest=sha256:f88c76a04a6a83db43e8af988c70e1fe3259b849151d308cc85b25cbd09a5437

Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · inbound

Benchmarking the Pedagogical Knowledge of Large Language Models cites this paper.

Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.697366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.697366Z digest=sha256:c382efc1249b391a83952aec238c610f14eff58bbdf5a4b0b42d3fc4c66fd8ac

Observation 730348d7-1f43-4ae1-8aea-8f0731132ca0 · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:42.082274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:42.082274Z digest=sha256:ecdd483f7d6df4b2dd0309851565925356f92dcdbfc7a329125159bbe9b3d6e2

Observation 9c272ff5-c8e7-4443-82c4-911beb380184 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.835725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.835725Z digest=sha256:e97746bf59d3144bddb74cb493ab0c710bedd90da9a8c597709440eb4054b89f

Observation 6e79f6e4-785c-4eab-9a33-5d27681e7eac · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.466197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.466197Z digest=sha256:7753fdda9d502196aa4cc961a09018a5d5691fb9e31ffae9442da0caf7b19884

Observation 8326260e-4bf2-4873-b1c8-2522a225cb5b · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:20:34.404981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:c3d6c3ed7a95ad3d77bab9575660cb3fd2ec9595489c9a748a22ae37b04fe436

Observation e2b61dff-efae-45ec-847d-e10efba11dfd · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:51.487379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:51.487379Z digest=sha256:f1b2f1b7e57312045000659529ea653c124bc1714c23934636dcbda8fda52a45

Observation 0c9c93e9-8fc0-4bda-a6c1-b5b050d53f65 · inbound

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation cites this paper.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.631644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.631644Z digest=sha256:2b52aacd93989afd5f0cbc7e3a3b4793c5264c70f7b67ce8616da541989295bf