Pith. sign in

Paper Citation Record · LEDGER

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.01781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01781 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:23.929114Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T10:47:31.141279Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fba070fe-dd71-477f-9785-f490b8588e91 · inbound

RouterBench: A Benchmark for Multi-LLM Routing System cites this paper.

RouterBench: A Benchmark for Multi-LLM Routing System When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:47:31.143539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T10:47:31.006944Z digest=sha256:1482e61f0ef8e6f48c89a60a9861bd448bcedfd6be538e6d2f9cca3d4a615af9

Observation a93ec630-105a-4a4c-a428-285fb95e7f65 · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:06.285386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:68b9ae8f9354bc28ad379b0f0ef297dc45202a56c1a0c47b5beaf342aaaab62e

Observation 9e97d5ea-bf91-479e-a756-e624b0e3e64d · inbound

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices cites this paper.

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:01:31.249702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:01:31.249702Z digest=sha256:de908a10a4fd4dd758bf531c7e926dee3e5c0f43d62b05b903f209015b4ab746

Observation ec67dca3-9d33-4a33-bbc2-7c9dbcbce281 · inbound

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic cites this paper.

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:26.884869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:37:26.884869Z digest=sha256:2d65b796b08714ead7662995d9e2c6f2b27eb0c678b9a6c9309aa182ccb3e66b

Observation 7fd1be3c-d545-4ff7-98b1-2696f4ee6316 · inbound

Too Big to Fool: Resisting Deception in Language Models cites this paper.

Too Big to Fool: Resisting Deception in Language Models When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:23.051934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:55:23.051934Z digest=sha256:eb22d9e7e985bcbf7a7bfd1b1f90715903dbca0b0986ea84f4adaffea6e7c7f0

Observation dc6b0e4d-8ae7-4e3b-b0c6-2ed46a0dbf33 · inbound

Evaluating LLM Reasoning in the Operations Research Domain with ORQA cites this paper.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.224508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.224508Z digest=sha256:98d77556c9b2d7c5c97bc6543c5d07a58407b5f2578597004d62f23481d44287

Observation cc8883a8-ff25-4446-a322-8650f92e5d4f · inbound

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks cites this paper.

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:44.788604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:22:44.788604Z digest=sha256:8d67c1c6a03ad01f80fe30c9dbe86bda6caaa9f218d92f4a0d91ae08c9b72695

Observation 6c1b2e4c-07fe-4278-963b-01609ebcdb51 · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.722646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.722646Z digest=sha256:8f500d60f5bea1224eb3f7bfdf516136636032ec0d723e16ab5922cebe54db36

Observation 6aabf32b-0f38-4f7f-9150-377d5b946014 · inbound

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cites this paper.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.541762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.541762Z digest=sha256:cff627629bc555aed542830f7561c0d18540a375b544bff45c39fc133abc7c5c

Observation 8075e7fb-d2d2-4739-9b6d-eb560a6331db · inbound

RoToR: Towards More Reliable Responses for Order-Invariant Inputs cites this paper.

RoToR: Towards More Reliable Responses for Order-Invariant Inputs When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:43.090058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:43.090058Z digest=sha256:b3a7f3c5fa3ee5527d0066288a63b8bbc396ea372a904bcfab3f9d7a5c0c3467

Observation d59d5cb2-0912-422a-a0cf-c2dece4483e3 · inbound

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition cites this paper.

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:23.929114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:33:23.929114Z digest=sha256:657eb4459901b82da9575418988eaf2687b2f28d5c5ecbe4b908fdca13cbf6b4

Observation 7d2fff4b-062f-4cdd-a428-2e2513130e46 · inbound

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism cites this paper.

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:29.300022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:17:29.300022Z digest=sha256:c6a1fb25dd2269d59e0a86342d84bdf5342b8e9695f62e759722b9af4b18bd81

Observation 3d6a5638-e8f3-4e0e-9d76-0c8b5cbfff94 · inbound

Quantifying Memory Utilization with Effective State-Size cites this paper.

Quantifying Memory Utilization with Effective State-Size When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:22.018300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:58:22.018300Z digest=sha256:a4713cde3ef08e77441d8fd923fa51dee6535a3aa197a7a9f88a9a0c9e60c0e3

Observation 67b6341e-63f4-46a0-8f7d-e48e562108f8 · inbound

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration cites this paper.

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T02:14:08.103425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:14:08.103425Z digest=sha256:b34ecd4dfe28c7d4f547520e4acb57a90720dceb939e99f862df7a5135ba095f