Pith. sign in

Paper Citation Record · LEDGER

Do Large Language Model Benchmarks Test Reliability?

As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 25 inbound Pith citation observations for arXiv:2502.03461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03461 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:45:34.619467Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:35.170514Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1b9bfd96-d5b6-4cfa-84e2-87a3cc39c5ff · outbound

This paper cites Vqa: Visual question answering.

Do Large Language Model Benchmarks Test Reliability? Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:45:34.908324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:45:34.567400Z digest=sha256:94c1056b9e3c520bba5b8caf975045645887f7de64e12db2005743a75d7b8f70

Observation db02c29d-d7fc-483a-a2a6-86b7ab1bb954 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Do Large Language Model Benchmarks Test Reliability? Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.574702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.574702Z digest=sha256:1a2f67b25343a8e1267f5d9405fa954d8217b34b15017f93326b9a28ddda3f0e

Observation 82bdb3d6-3b3f-4600-ab45-79296c2a0103 · outbound

This paper cites DeepSeek-V3 Technical Report.

Do Large Language Model Benchmarks Test Reliability? DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.586462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.586462Z digest=sha256:8ecfcc4e10f61c6c1998b324c84ad105b83c5eec1a70cf67ccecd7e8f12860fb

Observation 411b121d-201d-435c-8964-c23dd0ce06a8 · outbound

This paper cites Reasoning with language model is planning with world model.

Do Large Language Model Benchmarks Test Reliability? Reasoning with language model is planning with world model

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-09T04:45:34.819895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:45:34.591317Z digest=sha256:924b0cad71b4da85d77951bba5fac74a068a9d0b1e0e83698bce7b9a200875bd

Observation 498f2953-7801-4d3f-9dec-74e579fcd234 · outbound

This paper cites Parsing algebraic word problems into equations.

Do Large Language Model Benchmarks Test Reliability? Parsing algebraic word problems into equations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:45:34.897607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:45:34.595213Z digest=sha256:93b9c90413919abb6a4c387eb2b7492b813746ee65f34b66ee2789561da2ef9c

Observation 05fdc43a-0a68-41be-aaad-864ab2683443 · outbound

This paper cites Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks.

Do Large Language Model Benchmarks Test Reliability? Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.602333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.602333Z digest=sha256:96e9d50c1f361846ae791b621fcbc5b3e63c54b4f08c10086c5573a959190a2e

Observation dcd4c2f5-c4f7-4621-8547-9e8978bb0e2c · outbound

This paper cites ] The 15th Nepal China’s Tibet Economic and Trade Fair was held on 17-22 November 2015 in Bhrikutimandap, Kathmandu Nepal.

Do Large Language Model Benchmarks Test Reliability? ] The 15th Nepal China’s Tibet Economic and Trade Fair was held on 17-22 November 2015 in Bhrikutimandap, Kathmandu Nepal

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:45:34.877030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:45:34.619467Z digest=sha256:8ef31cbccb2ea697023a3ff6d9c06ac3450bfea528a729ffd3812f6768871a2f

Observation 1d55c155-c7b5-47a6-957c-ed5e752ebc3e · outbound

This paper cites A Chevy for $1? Car dealer chatbots show perils of AI for customer service.

Do Large Language Model Benchmarks Test Reliability? A Chevy for $1? Car dealer chatbots show perils of AI for customer service

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:45:34.887133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T04:45:34.599089Z digest=sha256:f9a22558e1184dd229b496d2288c4735d54882741c6e84cc48dc71e9111b684c

Observation 763ed826-7b74-4da8-82b9-b6bbac143dc2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Do Large Language Model Benchmarks Test Reliability? GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.610345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.610345Z digest=sha256:332504c8c9b032126de268fbb73e9cbc291c5e85cf033f1e556448c8c6011453

Observation 68821419-7ae1-4062-a5f3-d200ee6cddf1 · outbound

This paper cites TabFact: A Large-scale Dataset for Table-based Fact Verification.

Do Large Language Model Benchmarks Test Reliability? TabFact: A Large-scale Dataset for Table-based Fact Verification

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.578445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.578445Z digest=sha256:1e8cebea431e7b2c702179fb966f3280c6054535e8489e610af2c0009c14eef6

Observation 346a8a38-706f-490b-9a30-289f03360723 · outbound

This paper cites The Llama 3 Herd of Models.

Do Large Language Model Benchmarks Test Reliability? The Llama 3 Herd of Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.582280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.582280Z digest=sha256:a6580b615a00434e635c70c6bd9ab9827b13e50fa42467cfdf03975a0b8415d3

Observation 85b513ad-882e-4c7b-b755-fb7b32d523f8 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Do Large Language Model Benchmarks Test Reliability? GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.615477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.615477Z digest=sha256:2f388246a7d3c287fef55fdfa702e405119e4440658f06bc5419935badad11e9

Observation 5f633ecc-145a-4fa1-99df-696893cb17e7 · outbound

This paper cites What Will it Take to Fix Benchmarking in Natural Language Understanding?.

Do Large Language Model Benchmarks Test Reliability? What Will it Take to Fix Benchmarking in Natural Language Understanding?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.570985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.570985Z digest=sha256:c452d5d463eca68dfbaab97c995045ba73b6768f17c6b93c730eb758c582869c

Observation 4a42cf9e-0c95-4bfb-9fe4-597c8e5b6b27 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

Do Large Language Model Benchmarks Test Reliability? Are NLP Models really able to Solve Simple Math Word Problems?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.606271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.606271Z digest=sha256:db2a20c9df010071c05cb1ee508c4f88bfb794ff8758c239ad65e2ddcaa78832

Pith citing papers

Observation c946dab6-15b8-4f3d-8712-c8ef1f28d4f8 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Do Large Language Model Benchmarks Test Reliability?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:35.170514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:35.170514Z digest=sha256:f8451ada035b6433967a7915d4861854fe9fffc411e4744ea181c6ca94b89530

Observation e91f3fcc-2d10-4a13-b825-ef4dc75f380b · inbound

Benchmarking Misuse Mitigation Against Covert Adversaries cites this paper.

Benchmarking Misuse Mitigation Against Covert Adversaries Do Large Language Model Benchmarks Test Reliability?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:32:14.736497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T10:29:05.104520Z digest=sha256:8dd1e6a2ae8ab432524ad2948396b8e5d0c3d83de77a420a22aff27b6f4e2604

Observation d35071e2-6a71-4045-afaa-bd7f04821575 · inbound

Domain Specific Benchmarks for Evaluating Multimodal Large Language Models cites this paper.

Domain Specific Benchmarks for Evaluating Multimodal Large Language Models Do Large Language Model Benchmarks Test Reliability?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:39:41.721221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:39:41.721221Z digest=sha256:b371bc3eb367ccec724b96825762cc6c7dca7b6a58a645930ae069546891f763

Observation a71ed187-186d-4b4f-9697-0ac273e92bb4 · inbound

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models cites this paper.

Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models Do Large Language Model Benchmarks Test Reliability?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:56.526077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:31:56.526077Z digest=sha256:6874868cee700c11d723e99657c90de20009d0f1fe8c3180917e4b166dc33bf4

Observation 63b7f608-d52f-4b6d-b041-f429a29b044e · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation Do Large Language Model Benchmarks Test Reliability?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:47.173923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:47.173923Z digest=sha256:39167b2bb9b2a5339a382b28d53ca70ba658d1b8143ddd4e4477c76380e3a37c

Observation 47cb0821-33b3-4e5b-8e37-95ce3a928b1a · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence Do Large Language Model Benchmarks Test Reliability?

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:49:28.068240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:55a39bd769858e615aa7f7b89db22ba4a734520120139c80d4dfe6517bea0db1

Observation ce22e119-46ce-4751-80fc-ad0fb8f48fd7 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Do Large Language Model Benchmarks Test Reliability?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.328205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:4846fdaeacbab747203b3fd9460c07627fa1a83f1ade0c55986913d689b7df27

Observation f6fa3ca2-49ee-4a85-9eb3-6a48992e787d · inbound

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems cites this paper.

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems Do Large Language Model Benchmarks Test Reliability?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T05:11:55.853126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:11:55.853126Z digest=sha256:d0c0d0b90c16dfd7c7ae2d009c57b23bc9e9c93ad096eb60562ba54ee93cb589

Observation 4a505671-19a1-4e9c-bf77-740d23839d0f · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Do Large Language Model Benchmarks Test Reliability?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.451638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.451638Z digest=sha256:7c5eb0f30c1dc2e044d0850d7c2598222ea6bf7848ff47082a62f9a8888c006b

Observation 7a218eed-e892-43b4-8b68-880850af79c2 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Do Large Language Model Benchmarks Test Reliability?

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.470153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:e21fc9c6e63ee1382bf439cd9ef2df62c0e193196d468202227a9d320f31caa9

Observation e1db10e7-1e87-4e1d-9f63-3a1b455b282b · inbound

Simple Policy Gradients for Reasoning with Diffusion Language Models cites this paper.

Simple Policy Gradients for Reasoning with Diffusion Language Models Do Large Language Model Benchmarks Test Reliability?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:53.608588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:53.608588Z digest=sha256:4e090737dd89b00443e63ed2adf177b5f99baf91746b6a5bd8456121336ad0c1

Observation c5e5c7af-49ef-426f-ba8a-3cc8773d3d7e · inbound

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks cites this paper.

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks Do Large Language Model Benchmarks Test Reliability?

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T08:10:10.254371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:10:10.254371Z digest=sha256:5e20a7380b6c2db9704fcfe6dd6b96c1f1a7a3e2132d65f6643d644b80b1669d

Observation fdb3e2d9-139a-4afd-91a6-040638a7f6c1 · inbound

Model soups need only one ingredient cites this paper.

Model soups need only one ingredient Do Large Language Model Benchmarks Test Reliability?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T02:51:56.829649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:51:56.829649Z digest=sha256:31240e084dcd8129f9ab739032246b17d21e12b9e6e3d8ab58be8a3d2f63e9fa

Observation ea4d81c5-9b06-4941-89d2-36dc1cf1db8f · inbound

Weight Decay Improves Language Model Plasticity cites this paper.

Weight Decay Improves Language Model Plasticity Do Large Language Model Benchmarks Test Reliability?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T00:15:37.341267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:15:37.341267Z digest=sha256:387c6b0ee299af03fda4413f59219a9fec1dff377d0a7c3bac144861cfb0c950

Observation 7b9344d9-c88c-44ae-ba95-4d3497aa72bd · inbound

Position: Evaluation of ECG Representations Must Be Fixed cites this paper.

Position: Evaluation of ECG Representations Must Be Fixed Do Large Language Model Benchmarks Test Reliability?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:47.827344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:47.827344Z digest=sha256:3f716058f30babcb21cbf1a5da4045e5539191fa9ddcc42f1593fe617e057297

Observation 152365d0-7c94-4818-81b7-45f60b924bed · inbound

LLM Reasoning Is Latent, Not the Chain of Thought cites this paper.

LLM Reasoning Is Latent, Not the Chain of Thought Do Large Language Model Benchmarks Test Reliability?

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.559520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:49:05.178087Z digest=sha256:a42bb671c7d950d01f02dacdd7d2e379fa1f9e36bb5ba1ba3f5f6b58fae1c1da

Observation 54fe561a-54e9-4d96-96ff-04e5900b9a83 · inbound

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity cites this paper.

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Do Large Language Model Benchmarks Test Reliability?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.302907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T11:45:54.181019Z digest=sha256:385b14b4cf9c22539ea346ee19ed5d632dc24420a54d15cb8311e3859a557634

Observation 24f5328d-8cc5-4a49-9c81-d91b044b41d4 · inbound

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation cites this paper.

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation Do Large Language Model Benchmarks Test Reliability?

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:06:27.313724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:02:49.168872Z digest=sha256:659056310bf9b23fdfc28bcffcb68d70d751c46825929a92b227dad07f973d6b

Observation e23209f5-b865-4f58-a4c2-071e5dceec73 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Do Large Language Model Benchmarks Test Reliability?

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-08T18:54:00.977963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:ce14bf2e4f38ba0cb7475b0efd853d1edb894ef394bdfb1445914b592a6fbb6b

Observation 7f005ee2-3673-4666-9e11-b68055fe34a7 · inbound

Auditing LLM Benchmarks with Item Response Theory cites this paper.

Auditing LLM Benchmarks with Item Response Theory Do Large Language Model Benchmarks Test Reliability?

Reference 3

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T07:33:13.327707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:32:32.771146Z digest=sha256:8d1b4b4d72e206743a94c64d9c16ee9f0c7e4c794f7cbd91cfdb711b6da0f448

Observation 67bd4a2f-6efd-449e-87e1-4b08eed3105c · inbound

Flaws in the LLM Automation Narrative cites this paper.

Flaws in the LLM Automation Narrative Do Large Language Model Benchmarks Test Reliability?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.302392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:52:36.252919Z digest=sha256:d70a6eb4e83a4de5c331717f5b8e4c0cc3463b106cfe96723f2beb19322eb220

Observation 55c2c5de-d07c-4225-835d-3ef66f2d1930 · inbound

A Sovereign, Open-Source Foundation Model for German and English cites this paper.

A Sovereign, Open-Source Foundation Model for German and English Do Large Language Model Benchmarks Test Reliability?

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-13T03:06:27.991558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:06:27.991558Z digest=sha256:af6bd9cd056ca5afd037dcec22781a33ccf96a1f719c829e4374d51ec1550c73

Observation 45c50bf4-4920-4331-9b6e-eff107ac780d · inbound

A Sovereign, Open-Source Foundation Model for German and English cites this paper.

A Sovereign, Open-Source Foundation Model for German and English Do Large Language Model Benchmarks Test Reliability?

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-14T15:13:39.458378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:13:39.458378Z digest=sha256:3c856ba9ace51473847be9d44a6304702e3f26e6895b2124d3fc54b94db4ce5d

Observation 45372cbc-a301-48cf-bc0d-2700a7a04f6f · inbound

A Sovereign, Open-Source Foundation Model for German and English cites this paper.

A Sovereign, Open-Source Foundation Model for German and English Do Large Language Model Benchmarks Test Reliability?

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T07:46:08.660600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:46:08.660600Z digest=sha256:171f7afc39ade9bd41a3ba26a0e1503faeb24076f50ad83049416db15287fb0c

Observation 6ae6023b-7bb3-4441-9584-56729f48151e · inbound

Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility cites this paper.

Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility Do Large Language Model Benchmarks Test Reliability?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:38.247606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:38.247606Z digest=sha256:bad12b8a10803a04a145d8983bbf3995ae18c8293f88c5429d3b84d59001e0e9