Pith. sign in

Paper Citation Record · LEDGER

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2412.03597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03597 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:21:36.988483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T07:33:13.390618Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6a78e576-5327-4cb2-abe6-488df5659def · inbound

Rethinking the Understanding Ability across LLMs through Mutual Information cites this paper.

Rethinking the Understanding Ability across LLMs through Mutual Information The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:36.988483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:36.988483Z digest=sha256:5250b105874b2b7df1cfe666cc0c3d27728741a0a1ee91305e991484e6055101

Observation 4bd3f3c2-d3b9-4673-bab5-53e4d143ff46 · inbound

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models cites this paper.

LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:41:55.997298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:39:02.015913Z digest=sha256:021be09512ce3b403aa4d1ac85618f056c7cb9a404ab30dadf5029f4b3cf40b6

Observation 027116f9-b917-4296-870b-6cba7ef60d16 · inbound

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia cites this paper.

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:26:24.835532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T13:25:17.313704Z digest=sha256:a87c1b29b99f43eb4842414158467ac6c3bab09edcb6196dea99b1f6149918f0

Observation 3680f679-e001-4f0e-bc07-05f693db9aae · inbound

Human-aligned AI Model Cards with Weighted Hierarchy Architecture cites this paper.

Human-aligned AI Model Cards with Weighted Hierarchy Architecture The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:11:09.673896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T09:09:11.154051Z digest=sha256:e08eaf3c0827d57443fd7d2c42e55818fe0ae3d64a10150cece113dc953d051b

Observation cd72ad2b-5c21-4586-ba94-6a6d7070fc1a · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.437860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:f793c289c33cf26ac0d3f0f8667a234f1427f63e392ece9b2b6174289bcaef2d

Observation 20ac7aa0-9ded-4b97-a1af-337206e4eebe · inbound

Training a General Purpose Automated Red Teaming Model cites this paper.

Training a General Purpose Automated Red Teaming Model The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:08.329246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:20:09.224745Z digest=sha256:a59993d334a68de0e85cf5e9bf13524156cc7cc92dca640627be763e84335a23

Observation bd84c14f-df89-4b4d-9160-73d8af3b0f08 · inbound

Latent Performance Profiling of Large Language Models cites this paper.

Latent Performance Profiling of Large Language Models The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.392079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:31:02.595386Z digest=sha256:ecbd6d7b6cbd440e56af54f3ae96b0225c1c5d3d3a22d3e2850389cfd35e147a

Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · inbound

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) cites this paper.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.307850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.307850Z digest=sha256:493b73ad03e3e1e4b99578b9b49ee817865581a1f19be01ec7b9d346f00e8092