Pith. sign in

Paper Citation Record · LEDGER

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.03492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.03492 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:08:06.512628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:14:03.236562Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 772c4ac6-2bcc-4f8e-9b10-16db216a2aa6 · inbound

Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations cites this paper.

Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:06.512628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:08:06.512628Z digest=sha256:0e6c1baa599cfdf2d2c09d04b4a9b1b26c3d59e6e286006a1af66e8e3396d9de

Observation f7e1caf3-44f0-4c07-86a0-d19a3b528224 · inbound

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs cites this paper.

From Queries to Criteria: Understanding How Astronomers Evaluate LLMs Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:37.798046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:37.798046Z digest=sha256:2db6ea8c7b0a00d7c4e96fc24b41b2604104fcd3393c8c0d6f686f6b450be86b

Observation 464e1b2e-033f-4181-b6cd-8445e4645deb · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:44.325768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:44.325768Z digest=sha256:b11d79430518a68d39376db4cc901921b105f6add0bfa4c4d54f43485d31a0ad

Observation 985e0aaf-b029-4c94-806b-29d130e9e035 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:53.503157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:53.503157Z digest=sha256:681f06b776e6575892df49b61d41ad513188767a8fc366fb18b90e82edb63f04

Observation 74b179ef-447e-4276-a3e8-e63b291a653c · inbound

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito cites this paper.

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T08:03:52.276129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:03:52.276129Z digest=sha256:30c224764056b7502ded4af0270fa1660442ecd547f0fa564c7034312fba1e76

Observation 59ea88af-9ed3-44f4-a313-28e2d64dd207 · inbound

Inspectable AI for Science: A Research Object Approach to Generative AI Governance cites this paper.

Inspectable AI for Science: A Research Object Approach to Generative AI Governance Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:20:59.433259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:06:30.197232Z digest=sha256:6f9add913657d35c9eb199bedff365c04d601ae621bc9fdab77a154e367a284d

Observation 9b3f1e80-46a1-4959-8f7a-6b08e0af7eff · inbound

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing cites this paper.

How Compliant Are GitHub Actions Workflows? A Checklist-Based Study with LLM-Assisted Auditing Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.558726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:19:18.667967Z digest=sha256:f17cf0063e3910a037e6d89447444c721e26b2706ffea874f6c1d2a89412170e

Observation 84a8b33b-3879-4803-bac8-d08a11bfb0a4 · inbound

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi cites this paper.

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.608918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:18:08.326304Z digest=sha256:1e3e5e9ca577291d2f4a7fe002347fe32fd7a9ffb4c98e2a524b2e1a0875ea32

Observation 54f3f90f-0e1b-4d45-87d8-e5942a63277f · inbound

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility cites this paper.

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.947735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T06:55:05.581205Z digest=sha256:d273474d6ac7a127ff54d49a95540287b20be725044bc24f6a841fd5186cb643

Observation a2561c61-f772-4ee2-bdac-2a9207f60677 · inbound

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility cites this paper.

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:14:03.238763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:13:00.601818Z digest=sha256:f1a700bd6f621ce7b552e408d9b171b16499df618ca627ea8afa8a468b919e63

Observation d722c6cf-7d54-4b14-88a4-d94c794e860a · inbound

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings cites this paper.

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:05.907440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:05.907440Z digest=sha256:1c2b74c75fa1c7489c68ebafdf2a3c0d586cfaedcfc7a40d19815356b36c4bd9

Observation 332ff4db-a0c9-4568-97eb-ad4a60e7fae0 · inbound

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings cites this paper.

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:20.114741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:20.114741Z digest=sha256:e2b26787ee25d528a96e4d8d724e380fc98a18a3fd9279b5a5d362b06dc1d148