Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:04.462120Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2608.02444.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:04.462120Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6bc585d6-d3f4-4f24-a23c-d177e6a5f6aa · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Efficient benchmarking of AI agents,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92eeb214-c9da-4e4e-98e0-76bf282c353a · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision SWE-bench: Can language models resolve real-world GitHub issues?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79952cdf-d754-4245-a16b-8bf184e87cbf · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision SWE-bench leaderboards,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ccd03e1-e641-4d5b-ae85-d74efbe8308e · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8568ab-4219-4297-b317-7b68302d4ed2 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d1fca5-746a-4604-8d0c-8c63bee87aec · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7c8a70-366e-4211-9552-44826efe4c5d · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cc61c9a-68e1-4bcb-9572-d3e069f682ba · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b46a2fb-923c-4a78-afb6-ffb4da4e48dc · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Holistic Agent Leaderboard: The missing infrastructure for AI agent evaluation,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734ac525-f565-4975-ae5d-2ef05c0342ce · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Sequential tests of statistical hypotheses,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6431829b-6e8f-4f81-b69c-a24188277490 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Time-uniform, nonparametric, nonasymptotic confidence sequences,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a94a15f9-e1c9-4cf2-bcac-fa24abe10bc4 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Selectivenet: A deep neural network with an integrated reject option,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39e41e85-b040-4894-9449-735bd43a3d09 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Consistent estimators for learning to defer to an expert,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ffb8890-b1a4-49d8-bc03-2ecdbb92b577 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Conformal Risk Control
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7792a26e-f3aa-4191-abc4-f47f294f57c8 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Holistic evaluation of language models,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d58f2f3-9dca-4162-ac89-0903f8683b9c · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97598b0f-a068-43aa-84ad-40580d288ceb · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7499cc94-7d84-4097-bd2b-d41870d8d26f · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Efficient Evaluation of LLM Performance with Statistical Guarantees
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31fe661-ddad-45a7-86c4-106aa6de11f4 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Probability inequalities for the sum in sampling without replacement,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f59e237-2679-4419-913c-2c0e66cc0c2f · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761f7a42-9a06-4c22-945a-1902bccad3ac · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision AI Agents That Matter
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9a4f83-5ac8-44db-80e7-18816f51053f · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511290d2-3acb-4861-a129-130a80233df0 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision General Agent Evaluation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ebbd39-fd75-4c41-9c4c-640a81ef4233 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision A2Perf: Real-World Autonomous Agents Benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d92a8d-e366-44b7-87f8-f158dcc5a18a · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision ProSoftArena: Benchmarking hierarchical capabilities of multimodal agents in professional software environments,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77656e4-d561-4551-af1c-1a4ce5632115 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cda3a64-7c95-415c-9335-232e8daec7c4 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831ec9eb-dd1d-4816-ab7f-e3b29418f223 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ea0bdf-45b3-4e53-9e15-d98befaa2c74 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision 7076–7087
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3acd5e-8496-42bb-9064-eeb2690e0555 · outbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Available: https://www.swebench.com/
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.