Pith. sign in

Paper Citation Record · LEDGER

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights

As of 17 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.06893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06893 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78cbdc7a-d2f5-4e63-b11c-d016b53f4b04 · outbound

This paper cites Pairwise analysis of model performance.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Pairwise analysis of model performance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.600472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.247692Z digest=sha256:535d3561cf359b244f1988376e677051035c70d994ad5309c0fe979ca1dd8ece

Observation 77bc16fb-21d6-4a91-8814-da81f917fa3e · outbound

This paper cites Calculating optimal resampling for model evaluation.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Calculating optimal resampling for model evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.579888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.259478Z digest=sha256:c66b0d09370f77e564fd2f28dc8eea56a6f711db20a473c85dce4ea9d4a7896a

Observation 5d2cc3ba-3566-4d4c-a820-19a7c16aafe4 · outbound

This paper cites The AI evaluation substack, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights The AI evaluation substack, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.553817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.264917Z digest=sha256:6d6f2e62795be173954fce1bbf417c396b9e9555efc1172163f55e02053b9129

Observation e71c644c-e5ec-43b0-88c7-04992f764326 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.271238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.271238Z digest=sha256:be2c3807f859064d3763cc70295cb778d13da8813ebf49b8e833e5eb407d1440

Observation 12197ec2-6a20-49da-8bea-0eed1617a721 · outbound

This paper cites Autonomous systems evaluation standard.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Autonomous systems evaluation standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.536321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.280724Z digest=sha256:2e6742952981c0a5f867827bc3c87f0f1cb8c12d980cc6fadcb0f221530bd763

Observation 368d4930-bed7-4405-a097-373a5921c3c2 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GAIA: a benchmark for General AI Assistants

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.286229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.286229Z digest=sha256:6fa705e8d6fcf24255498890d04e2d9e94e0fdd19c2ef45a89970a7beec84e5e

Observation b3cfc9c6-33ac-45ac-924a-853631b4403a · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.292785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.292785Z digest=sha256:be3e1e8ee82dd75fa8abd3c189fd393b34ddeb03da272d7133fdcce97dff1d8a

Observation 1feabfbd-b7ff-4243-b054-eb37a744cf80 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.300173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.300173Z digest=sha256:93879137ad86e4b9cfce695c03a429e20af7af1c5dc6fb61e88adacc31c0bb25

Observation aeb4f00d-108e-4a45-bbae-dee22bea3d1c · outbound

This paper cites inspect\_ai: A framework for large language model evaluations, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_ai: A framework for large language model evaluations, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.519830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.308264Z digest=sha256:3febe51268948f33c7d7614798d2fe0ce75f2d49e562540c77b254638b49a652

Observation 02c2801e-22b1-4a59-bb3b-0df2e8d85e00 · outbound

This paper cites Research agenda.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Research agenda

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.503280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.313755Z digest=sha256:7e8f9145c765e5b4fa512fb6f46228430a1e98554cd29c583fe9102de00e340f

Observation 82e3a5a0-0d7a-4421-a559-d4231eeb1fe1 · outbound

This paper cites inspect\_evals: Collection of evals for Inspect AI , 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_evals: Collection of evals for Inspect AI , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.487244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.320409Z digest=sha256:dfaab65a17f79b9430836d3ebf9a5e2b5f6084ae449b3ff77850949defd99f31

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:6c57dba3c897da3243d484ae5c23d59afea114e8038e2eb798c59921ef5ca5fe

Observation f14fcdfe-824c-47f4-a1c7-a7fd4d137ddd · outbound

This paper cites General Scales Unlock AI Evaluation with Explanatory and Predictive Power.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.331074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.331074Z digest=sha256:063aecc3d6fe9b833483d435a18e1bf8b062146e80c12ee23be47bfa167c4929

Pith citing papers

No inbound Pith citation observations are available.