Pith. sign in

Paper Citation Record · LEDGER

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.07762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07762 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:22:56.260694Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf2e32d2-3bed-4994-94ac-1a44be24183d · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS 2023), vol.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Advances in Neural Information Processing Systems (NeurIPS 2023), vol

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.908288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.169882Z digest=sha256:3099dff8f46de7528c1e5b2b1de97692cf78b62333fe0a042cd902eea17b7b4c

Observation e166e748-13c0-4c34-b676-936972121e76 · outbound

This paper cites Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.174076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.174076Z digest=sha256:fe5e7a54fe4260d07e063d0ac805efdcd7468fdb8fb8c123823c7754161a683f

Observation c5a99705-a1f3-4e0a-81ed-2d0a943d3177 · outbound

This paper cites In: Inui, K., Sakti, S., Wang, H., Wong, D.F., Bhattacharyya, P., Banerjee, B., Ekbal, A., Chakraborty, T., Singh, D.P.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Inui, K., Sakti, S., Wang, H., Wong, D.F., Bhattacharyya, P., Banerjee, B., Ekbal, A., Chakraborty, T., Singh, D.P

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.899434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.178086Z digest=sha256:18c217b6386d6789a8d3bbfb62bec0cd947c5ebff010c5a7aa962b02c4eac13e

Observation 458a2532-3489-4f88-9b0d-fb780f03dd6d · outbound

This paper cites In: Interna- tional Conference on Learning Representations (ICLR 2026), Poster Presentation (2026).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Interna- tional Conference on Learning Representations (ICLR 2026), Poster Presentation (2026)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.890390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.181967Z digest=sha256:38c93da0f98cd16eecb86c48c901b5c2dc5fc35013bb42da4bd0604ce5e9aeb8

Observation 37865eca-7145-4812-9758-55f1747645c4 · outbound

This paper cites arXiv:2602.08229 (2026)https://arxiv.org/abs/ 2602.08229.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation arXiv:2602.08229 (2026)https://arxiv.org/abs/ 2602.08229

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.186240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.186240Z digest=sha256:cf56d608761994f9ac61de2c95d0412522670b73fee1f583b55a016a2ccae00f

Observation fffea009-16b4-46cf-b6c9-cbdae67828ff · outbound

This paper cites Accepted as a poster at ICLR 2026 (2026).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Accepted as a poster at ICLR 2026 (2026)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.880704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.190328Z digest=sha256:bfc46cf2f409eaba85cadae8284b7abe2093854afef5aa0cf640b100c8b6b67e

Observation 521f8251-7ece-467f-9319-4d506131640f · outbound

This paper cites an unresolved cited work.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:22:56.871622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.194609Z digest=sha256:4ea06baf24bad573f6e9a85b1d36024d09941a9ab35e847c9fa3a4ede909769e

Observation f2216670-4d8a-4fbc-8153-74e57fde2fa7 · outbound

This paper cites In: WETSEB 2026 at ICSE 2026, Rio de Janeiro, Brazil (2026) https://conf.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: WETSEB 2026 at ICSE 2026, Rio de Janeiro, Brazil (2026) https://conf

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.862051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.198338Z digest=sha256:2381b3e003a1f809cfc20fac65d1aa1d5246030b5c387092826e9a36606a97c5

Observation f80ee9dc-89c1-455a-a113-0dc004435781 · outbound

This paper cites Journal of Information Technology & Politics (2026).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Journal of Information Technology & Politics (2026)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.201851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.201851Z digest=sha256:df4705efb31dddf0d9b8f195cb68a24db7efefea8a5ac4de29362bb19847a012

Observation 4158930c-7cc7-42cf-961c-0434245eda9f · outbound

This paper cites arXiv:2601.08785 (2026) https://arxiv.org/abs/2601.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation arXiv:2601.08785 (2026) https://arxiv.org/abs/2601

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.205427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.205427Z digest=sha256:25caa03c7ed15aa86f0e5ef070c697cb7742725543a0d4b6988cebe537b655b3

Observation 7e6ca39a-d70a-43d8-8ac3-ca9fa1b3174c · outbound

This paper cites Stanford Graduate School of Business (2024).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Stanford Graduate School of Business (2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.852170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.209118Z digest=sha256:b2966043f932597b2dcd3538b03ce02b02933f9060cc905b035ae27fea5e0633

Observation 436d052d-5327-43e2-bd65-8feaa3b6b379 · outbound

This paper cites npj Artificial Intelligence 2, 7 (2026).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation npj Artificial Intelligence 2, 7 (2026)

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:22:56.842065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.212757Z digest=sha256:927e0e140de62c252185cb8b2cde1c79c7e62d7b7afa9f025bdfb239bfcb5d4f

Observation 1cc679d5-d34d-4917-827d-4cd15a324b19 · outbound

This paper cites an unresolved cited work.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Unresolved cited work

Reference 13

Resolution
verified exact
doi, observed 2026-08-11T00:22:56.293730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.216496Z digest=sha256:b51b168e4ba0c710d7885930d2af8f59fbaa0251cb26ddf420922a0fdd4bd6a8

Observation b45914c5-63e6-4b35-90ed-9f3fa130fbd7 · outbound

This paper cites The Guardian, Jan- uary 27 (2025).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation The Guardian, Jan- uary 27 (2025)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.832811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.220531Z digest=sha256:25b4e4f868ccb6398fe960161af9865bd08c49f225681654c34ba2fb30127fda

Observation afb52df9-7fd8-4de3-87bf-5d8dd3d276ca · outbound

This paper cites August 5 (2025).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation August 5 (2025)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.823165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.224152Z digest=sha256:7546bcc7762aa5fd723d8ade6d2495c45fbba177d8807090ea7e53f3d41acae8

Observation 65be31c7-3a11-47f3-8a76-54605581cb7a · outbound

This paper cites May 5 (2026).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation May 5 (2026)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.814371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.227723Z digest=sha256:0c9ebe2df35ac751ebbd7c01ca4f15d6c2f39b05dd740b23cbbe6da28a6854b0

Observation 1472eddb-abb6-4c29-adfb-5e71cea46ba9 · outbound

This paper cites https://www.llama.com/docs/ model-cards-and-prompt-formats/llama3 3/.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation https://www.llama.com/docs/ model-cards-and-prompt-formats/llama3 3/

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.803588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.231202Z digest=sha256:900ae28e21b54e8457d2ee34ef2299f692eaf8abb71058f8cf5c7a894e083cd9

Observation 66d16501-bed2-4ed1-8937-ec5fe65a0478 · outbound

This paper cites Applied Artificial Intelligence 39, 2439610 (2025) https://www.tandfonline.com/ doi/full/10.1080/08839514.2024.2439610.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Applied Artificial Intelligence 39, 2439610 (2025) https://www.tandfonline.com/ doi/full/10.1080/08839514.2024.2439610

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-11T00:22:56.519959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.234769Z digest=sha256:1f1be19993413c37925aed1a84ee7549ad1778e8589a510909b553b02192c906

Observation 890846a9-ba39-48ba-a2f2-89c55725a734 · outbound

This paper cites https://finance.yahoo.com/news/ yann-lecun-meta-fudged-little-100000402.html.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation https://finance.yahoo.com/news/ yann-lecun-meta-fudged-little-100000402.html

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.793570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.238322Z digest=sha256:07206204ae352b1cb35f5ade93f27f937d9c25cfcf3307e7555fa1bf6a165e04

Observation 0d6927ed-ed44-4574-ab1d-11e8a50d6eac · outbound

This paper cites Qwen3 Technical Report.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.241929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.241929Z digest=sha256:6cd859132eec3defa481634c5c88b06f83e6db75d193d766c78562b34ed17a57

Observation 94f8e92f-6676-4a59-990d-861b48ac4b46 · outbound

This paper cites Z.AI Blog (2026).

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Z.AI Blog (2026)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.783034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.245827Z digest=sha256:487905b0036b61af069e35d73d80a0ed7a260bea4575b49a23f4af278197297f

Observation f6019df7-8db9-45e2-83ee-2bb25c95b629 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.249482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.249482Z digest=sha256:ad70b051c95c9c3a7665b7b5f49496e9377a2047ae0156a8bb804561c01046df

Observation 6f7c5642-9d80-4550-81ba-4287d47db44a · outbound

This paper cites In: Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024), pp.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024), pp

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:22:56.253308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:22:56.253308Z digest=sha256:67a06e97001788e380cef68f34a3b1c162e8bd98e663dde09ca25c9cdb9e5dbe

Observation 41a29074-301a-4a4d-98ee-a225d9ab67e0 · outbound

This paper cites https://huggingface.co/collections/mistralai/ mistral-large-3.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation https://huggingface.co/collections/mistralai/ mistral-large-3

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:22:56.772514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.256866Z digest=sha256:965596960367b7cb06646114ffd3b8bf91bed6c5b98de3807696a900dd9fede0

Observation aa815a6b-3d72-43eb-b797-aa3134347215 · outbound

This paper cites an unresolved cited work.

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:22:56.762387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T00:22:56.260694Z digest=sha256:5422b5b3075c62992494cfb0d386f49636799eb19e983c71839fc145c9b31c9b

Pith citing papers

No inbound Pith citation observations are available.