Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.06893.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:34.331074Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 78cbdc7a-d2f5-4e63-b11c-d016b53f4b04 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Pairwise analysis of model performance
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77bc16fb-21d6-4a91-8814-da81f917fa3e · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Calculating optimal resampling for model evaluation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d2cc3ba-3566-4d4c-a820-19a7c16aafe4 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights The AI evaluation substack, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e71c644c-e5ec-43b0-88c7-04992f764326 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12197ec2-6a20-49da-8bea-0eed1617a721 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Autonomous systems evaluation standard
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 368d4930-bed7-4405-a097-373a5921c3c2 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GAIA: a benchmark for General AI Assistants
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cfc9c6-33ac-45ac-924a-853631b4403a · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1feabfbd-b7ff-4243-b054-eb37a744cf80 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeb4f00d-108e-4a45-bbae-dee22bea3d1c · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_ai: A framework for large language model evaluations, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02c2801e-22b1-4a59-bb3b-0df2e8d85e00 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Research agenda
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82e3a5a0-0d7a-4421-a559-d4231eeb1fe1 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights inspect\_evals: Collection of evals for Inspect AI , 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc9251f0-e653-4768-ba17-b7bf3c19b967 · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f14fcdfe-824c-47f4-a1c7-a7fd4d137ddd · outbound
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.