Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:23:53.854989Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 9 inbound Pith citation observations for arXiv:2505.03814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:23:53.854989Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:56:08.347771Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T20:41:10.421116Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 070aa99e-798c-4664-91e5-9b69560b13e3 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Explaining neural scaling laws
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c12a49-c62e-4ce7-88f6-909e3b6d997c · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Concentration Inequalities: A Nonasymptotic Theory of Independence
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da05ea5-c7fa-42fd-ac4b-12c3756e685f · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs AutoEval Done Right: Using Synthetic Data for Model Evaluation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1e5364-631c-43d7-ac8e-c95529a8dccc · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs A survey on evaluation of large language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57a2fd2-2d44-4972-89fb-554ebff838a4 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77dcce3f-853a-4a34-9aa5-101d2d75c1be · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc40bd80-d021-49c4-bd5a-c2a6c9a79c8d · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Shieldagent: Shielding agents via verifiable safety policy reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9e064a-772b-4505-ba98-f388ed4cd87d · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation abbc23a7-8002-40ed-bd68-d4ad5ce481d3 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6698245a-93d7-472f-9c65-4a6ec5c51f93 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd95a8bf-d7a5-449d-8dec-281166fc77f1 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Asymptotic behavior of expected sample size in certain one sided tests
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c90d72d7-5ce4-4641-a4c0-a3296abe824a · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Stratified prediction-powered inference for hybrid language model evaluation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation db2f6605-dd37-4c0d-8f56-a89c593b473c · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs LLM-based NLG Evaluation: Current Status and Challenges
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70748a1a-4215-4ebf-af3f-15e03c45cff2 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Measuring massive multitask language understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b5d647-f022-4ae2-a9cb-35ed027f2a9e · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Measuring mathematical problem solving with the math dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed6c01d-e666-4131-9cb0-bd2f72dacc36 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs An empirical analysis of compute-optimal large language model training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee9e94c2-8869-423d-8e9e-1b6f32556786 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24b7fe8-1ee9-4605-9387-bcb6cd6497e5 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Scaling Laws for Neural Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44fc8211-d21b-48a8-846d-53fbc5f4954c · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Noisy binary search and its applications
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a6588064-5d9e-4d26-b49f-96263b89fa93 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5d40505a-3946-4f22-8c5e-6efaccecf898 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f7f9ac-27bb-4dea-9bff-fddae637e1de · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs TruthfulQA: Measuring how models mimic human falsehoods
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eab109c0-a0cd-497e-8e13-66d3fbe182f8 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bcfb9a5-a3c4-4846-b6c6-7cdfcdce33e2 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs The sample complexity of exploration in the multi-armed bandit problem
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 845ec27a-67fd-4707-b668-0da9010588d8 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2c49e0-4a20-4626-872b-0f6e2ca21575 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs tinybenchmarks: evaluating llms with fewer examples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2e2f9e1f-1368-4b63-ba3b-a4cbb5563499 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d737d72-4dfd-4b2e-8702-6aebceee47f0 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e852a4b-bcce-47fb-82c9-bf0a7669a068 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Collaborative Performance Prediction for Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15114ddb-4cc5-4af6-8c92-21501c98d8b6 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs SafetyBench: Evaluating the safety of large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64e5ca4c-c723-4346-b87f-27141c34c658 · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Adaptive concentration inequalities for sequential decision problems
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1343638e-800f-4368-9598-32641874554c · outbound
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4c05919c-aa1e-4ef6-8b37-46ce0ca263c2 · inbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0a392f-3bd4-4fc9-ab66-29e8e58cca04 · inbound
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c40ee8-e93d-4b1a-a8fd-2c4f649f8ba9 · inbound
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb21be98-65cc-4441-808f-cc85b0c17896 · inbound
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c271690c-8241-43d3-93c2-3d91eaeedb8a · inbound
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c98c1b-116c-44c4-800f-9e4655137da9 · inbound
BayesAME: Bayesian Active Model Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97598b0f-a068-43aa-84ad-40580d288ceb · inbound
ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60bbfc65-7c72-4eae-aa35-660be58f2965 · inbound
Dynamically Allocating Evaluation Effort for Model Ranking Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecdb3c17-efd6-427f-ba9b-3647ca7e036f · inbound
RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.