Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:22:56.260694Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.07762.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:22:56.260694Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cf2e32d2-3bed-4994-94ac-1a44be24183d · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Advances in Neural Information Processing Systems (NeurIPS 2023), vol
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e166e748-13c0-4c34-b676-936972121e76 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a99705-a1f3-4e0a-81ed-2d0a943d3177 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Inui, K., Sakti, S., Wang, H., Wong, D.F., Bhattacharyya, P., Banerjee, B., Ekbal, A., Chakraborty, T., Singh, D.P
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 458a2532-3489-4f88-9b0d-fb780f03dd6d · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Interna- tional Conference on Learning Representations (ICLR 2026), Poster Presentation (2026)
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 37865eca-7145-4812-9758-55f1747645c4 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation arXiv:2602.08229 (2026)https://arxiv.org/abs/ 2602.08229
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fffea009-16b4-46cf-b6c9-cbdae67828ff · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Accepted as a poster at ICLR 2026 (2026)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 521f8251-7ece-467f-9319-4d506131640f · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f2216670-4d8a-4fbc-8153-74e57fde2fa7 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: WETSEB 2026 at ICSE 2026, Rio de Janeiro, Brazil (2026) https://conf
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f80ee9dc-89c1-455a-a113-0dc004435781 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Journal of Information Technology & Politics (2026)
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4158930c-7cc7-42cf-961c-0434245eda9f · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation arXiv:2601.08785 (2026) https://arxiv.org/abs/2601
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6ca39a-d70a-43d8-8ac3-ca9fa1b3174c · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Stanford Graduate School of Business (2024)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 436d052d-5327-43e2-bd65-8feaa3b6b379 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation npj Artificial Intelligence 2, 7 (2026)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1cc679d5-d34d-4917-827d-4cd15a324b19 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b45914c5-63e6-4b35-90ed-9f3fa130fbd7 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation The Guardian, Jan- uary 27 (2025)
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation afb52df9-7fd8-4de3-87bf-5d8dd3d276ca · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation August 5 (2025)
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 65be31c7-3a11-47f3-8a76-54605581cb7a · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation May 5 (2026)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1472eddb-abb6-4c29-adfb-5e71cea46ba9 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation https://www.llama.com/docs/ model-cards-and-prompt-formats/llama3 3/
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 66d16501-bed2-4ed1-8937-ec5fe65a0478 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Applied Artificial Intelligence 39, 2439610 (2025) https://www.tandfonline.com/ doi/full/10.1080/08839514.2024.2439610
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 890846a9-ba39-48ba-a2f2-89c55725a734 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation https://finance.yahoo.com/news/ yann-lecun-meta-fudged-little-100000402.html
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d6927ed-ed44-4574-ab1d-11e8a50d6eac · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94f8e92f-6676-4a59-990d-861b48ac4b46 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Z.AI Blog (2026)
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f6019df7-8db9-45e2-83ee-2bb25c95b629 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7c5642-9d80-4550-81ba-4287d47db44a · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation In: Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024), pp
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a29074-301a-4a4d-98ee-a225d9ab67e0 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation https://huggingface.co/collections/mistralai/ mistral-large-3
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa815a6b-3d72-43eb-b797-aa3134347215 · outbound
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
No inbound Pith citation observations are available.