Pith. sign in

Paper Citation Record · LEDGER

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.06893.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06893 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78cbdc7a-d2f5-4e63-b11c-d016b53f4b04 · outbound

This paper cites Pairwise analysis of model performance.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Pairwise analysis of model performance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.600472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.247692Z digest=sha256:0caefce9af1163f0bb4594fddd1f071f98fe6436ab4bcf6db708c4412e172160

Observation 77bc16fb-21d6-4a91-8814-da81f917fa3e · outbound

This paper cites Calculating optimal resampling for model evaluation.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Calculating optimal resampling for model evaluation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.579888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.259478Z digest=sha256:dda870dbd150dc13e75c2a69297fcbfd309da4331a6ef0fb606cd021cdbb176d

Observation 5d2cc3ba-3566-4d4c-a820-19a7c16aafe4 · outbound

This paper cites The AI evaluation substack, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights The AI evaluation substack, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.553817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.264917Z digest=sha256:77dfe6979db1abd2d729b10a141c6e3c4f1e0c9af976f9f1dd499ad103e7ebd8

Observation e71c644c-e5ec-43b0-88c7-04992f764326 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.271238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.271238Z digest=sha256:129795fe1425ab7ea8c72138e4136f7b26c4dfb5c3ff76112151869fa1d22b01

Observation 12197ec2-6a20-49da-8bea-0eed1617a721 · outbound

This paper cites Autonomous systems evaluation standard.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Autonomous systems evaluation standard

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.536321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.280724Z digest=sha256:4b083fdb93a426257c11371fd6f8a99b45d0e2e176e13ed8788398dd507417eb

Observation 368d4930-bed7-4405-a097-373a5921c3c2 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GAIA: a benchmark for General AI Assistants

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.286229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.286229Z digest=sha256:01886f0e4de9997310f51f2a060abe4dd2e95cda2c7f39739c0d0f197fc4bf9f

Observation b3cfc9c6-33ac-45ac-924a-853631b4403a · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.292785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.292785Z digest=sha256:2d5c339e09b35209697aee0a575412cbe0cbf90e51a437d51333e094e4125b45

Observation 1feabfbd-b7ff-4243-b054-eb37a744cf80 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.300173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.300173Z digest=sha256:69a80664a5af9ad4d57de81e46ffb7f657d08d127150bbc4baa6c15d8e9c9743

Observation aeb4f00d-108e-4a45-bbae-dee22bea3d1c · outbound

This paper cites inspect\_ai: A framework for large language model evaluations, 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_ai: A framework for large language model evaluations, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.519830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.308264Z digest=sha256:4909953fdcca100fdd1582b855ef5d7068c9a1d86aeb59c0164791d5dd8e74dd

Observation 02c2801e-22b1-4a59-bb3b-0df2e8d85e00 · outbound

This paper cites Research agenda.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Research agenda

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.503280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.313755Z digest=sha256:c8afa5ee1f7d6665a3baf35b012d7f888345ca11568d7863f3b60c23aa483f21

Observation 82e3a5a0-0d7a-4421-a559-d4231eeb1fe1 · outbound

This paper cites inspect\_evals: Collection of evals for Inspect AI , 2024.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_evals: Collection of evals for Inspect AI , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:34.487244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:55:34.320409Z digest=sha256:90ba518632d56e03286a830b206eede6025e2b883bf48b7a49e900e4afd7ec71

Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.325643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.325643Z digest=sha256:4c3a1b1d916ffc92267a504d5e0c1f7e332df1b05ec048cc03f34069f18d2837

Observation f14fcdfe-824c-47f4-a1c7-a7fd4d137ddd · outbound

This paper cites General Scales Unlock AI Evaluation with Explanatory and Predictive Power.

Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:34.331074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:34.331074Z digest=sha256:f0773848dcb32d851ecaf02ecaabb6a7dfa5f2e45dcc89d1bb879869b1b5ed26

Pith citing papers

No inbound Pith citation observations are available.