Pith. sign in

Paper Citation Record · LEDGER

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

As of 19 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 9 inbound Pith citation observations for arXiv:2505.03814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03814 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:23:53.854989Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:56:08.347771Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:41:10.421116Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 070aa99e-798c-4664-91e5-9b69560b13e3 · outbound

This paper cites Explaining neural scaling laws.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Explaining neural scaling laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.715142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.715142Z digest=sha256:f2dbd4a9d587bee7bf283928dac8bd14d1ff8753af564045b8e18e1c6778fc97

Observation 26c12a49-c62e-4ce7-88f6-909e3b6d997c · outbound

This paper cites Concentration Inequalities: A Nonasymptotic Theory of Independence.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Concentration Inequalities: A Nonasymptotic Theory of Independence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.720305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.720305Z digest=sha256:611f50dffc8c4560a61b892fd8ca6ab2491a79013bb329fe66abf449b3f19a22

Observation 6da05ea5-c7fa-42fd-ac4b-12c3756e685f · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.725264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.725264Z digest=sha256:392cbe75a7398d7c73f885b2f5603bfd4ed8c995dbca1ca8eeb36a0722a3ace4

Observation 3a1e5364-631c-43d7-ac8e-c95529a8dccc · outbound

This paper cites A survey on evaluation of large language models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs A survey on evaluation of large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.730441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.730441Z digest=sha256:d5841c178171678665d5b08aba538a6dd68cec4e42b5c6e4502c41994147388f

Observation b57a2fd2-2d44-4972-89fb-554ebff838a4 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.734953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.734953Z digest=sha256:8bb8aafda030760557760981561f105aae911e1cdc4ea05fc6cf12244974681b

Observation 77dcce3f-853a-4a34-9aa5-101d2d75c1be · outbound

This paper cites SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.740215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.740215Z digest=sha256:0e5b551952897d3cd3ecac08327e6106f177cee13820ad11e7fefaa58c420f71

Observation cc40bd80-d021-49c4-bd5a-c2a6c9a79c8d · outbound

This paper cites Shieldagent: Shielding agents via verifiable safety policy reasoning.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Shieldagent: Shielding agents via verifiable safety policy reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.745840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.745840Z digest=sha256:c8febd8b9d4ff5584b74c226a4174819579f34e3cffa52d3c29b114e7deeb5d6

Observation 6f9e064a-772b-4505-ba98-f388ed4cd87d · outbound

This paper cites Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.494512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.750234Z digest=sha256:7f7db88abf45f29f098ea08c680c9ec1367b736651c335e26620857a0a224550

Observation abbc23a7-8002-40ed-bd68-d4ad5ce481d3 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.479843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.754513Z digest=sha256:e22305932b8cc242926ea30cb9743ff39e25bd38ca16ba5e99c7dc8e0c0db3c7

Observation 6698245a-93d7-472f-9c65-4a6ec5c51f93 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.758967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.758967Z digest=sha256:8fe34057ef6952915ab5264775a16d0187873e3fa3574dc28d92a45b42722d48

Observation dd95a8bf-d7a5-449d-8dec-281166fc77f1 · outbound

This paper cites Asymptotic behavior of expected sample size in certain one sided tests.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Asymptotic behavior of expected sample size in certain one sided tests

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.464556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.763560Z digest=sha256:f1c123de29fd8af2ae26508a17bf378b3d5d0ea6563e13cf57da6f4f73151aa2

Observation c90d72d7-5ce4-4641-a4c0-a3296abe824a · outbound

This paper cites Stratified prediction-powered inference for hybrid language model evaluation.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Stratified prediction-powered inference for hybrid language model evaluation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.442973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.767922Z digest=sha256:5c308dfa33eb7aadcc8b1defc3ec04be864fcfa93e1ddc5fbdb22ff18a5cd3c1

Observation db2f6605-dd37-4c0d-8f56-a89c593b473c · outbound

This paper cites LLM-based NLG Evaluation: Current Status and Challenges.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs LLM-based NLG Evaluation: Current Status and Challenges

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.772212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.772212Z digest=sha256:9b6e34111767dade83d3f35ef6f09dfe93424ab6b9f7ab3c8f37e95a682cc9eb

Observation 70748a1a-4215-4ebf-af3f-15e03c45cff2 · outbound

This paper cites Measuring massive multitask language understanding.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Measuring massive multitask language understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.776770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.776770Z digest=sha256:b5520a018ae1a42e7cffbd54a9fa9ba511e3af6817cb7355b3768e390ba24143

Observation 76b5d647-f022-4ae2-a9cb-35ed027f2a9e · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Measuring mathematical problem solving with the math dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.780898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.780898Z digest=sha256:5b654e739a7c041c5e42ea14f2a86cd3e669ee90361c8cc0e982f911f4fd0213

Observation 0ed6c01d-e666-4131-9cb0-bd2f72dacc36 · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs An empirical analysis of compute-optimal large language model training

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.403175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.784899Z digest=sha256:0b74f95a375da6eba208271aa7fcf2024c8bd0d70bc3991264816f4a2a8adcd4

Observation ee9e94c2-8869-423d-8e9e-1b6f32556786 · outbound

This paper cites TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.789041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.789041Z digest=sha256:b21f3559f14c00d4b758db397046cf1119dd23eb0d2e3c616cffa462e6c66480

Observation c24b7fe8-1ee9-4605-9387-bcb6cd6497e5 · outbound

This paper cites Scaling Laws for Neural Language Models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Scaling Laws for Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.793745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.793745Z digest=sha256:bf2404869f4d26770cb3c29dc8d4cd6e4f807fca026bcfc9ebc87265a7aa12d8

Observation 44fc8211-d21b-48a8-846d-53fbc5f4954c · outbound

This paper cites Noisy binary search and its applications.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Noisy binary search and its applications

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.386167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.798063Z digest=sha256:b3c6678a960234b62050fa897823443ff95b3bd85756870168835dc5e1ef9c25

Observation a6588064-5d9e-4d26-b49f-96263b89fa93 · outbound

This paper cites an unresolved cited work.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-08-16T04:23:53.892684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.802074Z digest=sha256:84d7e42a4d8f3bdd89557350257aecccc9ccc05cf3e45bb15e02f8f57833d213

Observation 5d40505a-3946-4f22-8c5e-6efaccecf898 · outbound

This paper cites metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.806181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.806181Z digest=sha256:1c80e7db40c972966e7ca20a73f88e24caca15a5ff812cc1034293180f07983c

Observation f6f7f9ac-27bb-4dea-9bff-fddae637e1de · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs TruthfulQA: Measuring how models mimic human falsehoods

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.368557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.810155Z digest=sha256:be3d859f724ecc17943d80ea0cf621119d469742ed0bab720d4c087504df7dd5

Observation eab109c0-a0cd-497e-8e13-66d3fbe182f8 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.813865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.813865Z digest=sha256:9ee8e72733f700c1c42c15382bd2db1ab2943254e54d6d57a887890b85415f72

Observation 4bcfb9a5-a3c4-4846-b6c6-7cdfcdce33e2 · outbound

This paper cites The sample complexity of exploration in the multi-armed bandit problem.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs The sample complexity of exploration in the multi-armed bandit problem

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.350285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.817752Z digest=sha256:af3d4095e204622d21c5328f56b20850e251d582868df98885215bc02e4ee5ec

Observation 845ec27a-67fd-4707-b668-0da9010588d8 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.822075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.822075Z digest=sha256:97e0725b9333b9e37d91a202768286958695d38db694b4661be1b14b5acc13c5

Observation 4b2c49e0-4a20-4626-872b-0f6e2ca21575 · outbound

This paper cites tinybenchmarks: evaluating llms with fewer examples.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs tinybenchmarks: evaluating llms with fewer examples

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.334763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.826397Z digest=sha256:fc97518e5599e5edf74eebe23514f3589f02ef00f217fb40e8e5ffce2a98f255

Observation 2e2f9e1f-1368-4b63-ba3b-a4cbb5563499 · outbound

This paper cites MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.830673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.830673Z digest=sha256:96ae7819fa46446b51735b57c69282924909997dde30c954b1c3cc7ba29b3557

Observation 6d737d72-4dfd-4b2e-8702-6aebceee47f0 · outbound

This paper cites Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.834959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.834959Z digest=sha256:c0b7b4398e0f0bfc27b46c1c3a9964a7e763c9daf03dfcd16285f4680bcc9f29

Observation 2e852a4b-bcce-47fb-82c9-bf0a7669a068 · outbound

This paper cites Collaborative Performance Prediction for Large Language Models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Collaborative Performance Prediction for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:23:53.839711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:23:53.839711Z digest=sha256:700fc225cb7a68c4d3a3835d25910567404af2dc09859525fc0c693a4f285fda

Observation 15114ddb-4cc5-4af6-8c92-21501c98d8b6 · outbound

This paper cites SafetyBench: Evaluating the safety of large language models.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs SafetyBench: Evaluating the safety of large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.317908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.845368Z digest=sha256:8371848c04ae8dc490da08e8cd41176f4914f46c8e6031af5d8dfaf5407c1c34

Observation 64e5ca4c-c723-4346-b87f-27141c34c658 · outbound

This paper cites Adaptive concentration inequalities for sequential decision problems.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Adaptive concentration inequalities for sequential decision problems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.302012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.850763Z digest=sha256:d29e32db46cec8e6e7e92faa2410b930f1b5cc8b73f4afb84d3e540d24134b35

Observation 1343638e-800f-4368-9598-32641874554c · outbound

This paper cites Promptbench: Towards evaluating the robustness of large language models on adversarial prompts.

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs Promptbench: Towards evaluating the robustness of large language models on adversarial prompts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:23:54.284190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:23:53.854989Z digest=sha256:f097e4408e16046605316934dfdcd1f8e55c126bb95e7fb6d88a17a2bda0741a

Pith citing papers

Observation 4c05919c-aa1e-4ef6-8b37-46ce0ca263c2 · inbound

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models cites this paper.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:13.720220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:13.720220Z digest=sha256:e8d226cc2570bec533d5a1dfe7d57134d6d0cfc7d77bb65a54c89d2df1ad6845

Observation 8a0a392f-3bd4-4fc9-ab66-29e8e58cca04 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:56.098545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:56.098545Z digest=sha256:978ae9a8f5bf97bfad9ee63795420a070e4b38aa2211881aa1b34776a6bf1a0d

Observation b3c40ee8-e93d-4b1a-a8fd-2c4f649f8ba9 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:10.426607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:b5efecc2d11541efed7d2e0ef2eb9f796d10531e87eaab8bd3bd077812b79801

Observation bb21be98-65cc-4441-808f-cc85b0c17896 · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.933340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:13:16.896486Z digest=sha256:e5a5301d27eac122282caa8f397176f98f6bf22cb25f4d4eec67774ff629fb3b

Observation c271690c-8241-43d3-93c2-3d91eaeedb8a · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T14:49:15.527713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:49:15.527713Z digest=sha256:523b4128cf178038082c92870f7447e7cc5272237d2c14a49309b84ea4a69fcc

Observation 05c98c1b-116c-44c4-800f-9e4655137da9 · inbound

BayesAME: Bayesian Active Model Evaluation cites this paper.

BayesAME: Bayesian Active Model Evaluation Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-30T13:33:13.002509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T13:33:13.002509Z digest=sha256:43621ee8653aa57db597dbe0ef9419f366ada36b78f9d17011bf4fdc227cdd64

Observation 97598b0f-a068-43aa-84ad-40580d288ceb · inbound

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision cites this paper.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:03.239193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:03.239193Z digest=sha256:f6e7a335432f3efc8bec6e38adba975c0c07184ad1d78bdb031ed6f2bd9ff06f

Observation 60bbfc65-7c72-4eae-aa35-660be58f2965 · inbound

Dynamically Allocating Evaluation Effort for Model Ranking cites this paper.

Dynamically Allocating Evaluation Effort for Model Ranking Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T14:56:08.347771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:56:08.347771Z digest=sha256:8865cf3879780096809a6a810ab8d4981f5f3d23fcecd48218d19a066ce06bd8

Observation ecdb3c17-efd6-427f-ba9b-3647ca7e036f · inbound

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough cites this paper.

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T00:37:48.158499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:37:48.158499Z digest=sha256:cb87f073fd054be1cc48a4689da9bb8e69c8715bb45dd3d99578379e88da01f4