Pith. sign in

REVIEW 3 cited by

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.05602 v3 pith:XZTV2VEW submitted 2025-05-08 cs.AI stat.AP

classification cs.AIstat.AP
keywords hibayesbayesianadvancedevaluationhierarchicaldataevaluationsframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As Large Language Models (LLMs) and other AI systems evolve, robustly estimating their capabilities from inherently stochastic outputs while systematically quantifying uncertainty in these estimates becomes increasingly important. Further, advanced AI evaluations often have a nested hierarchical structure, exhibit high levels of complexity, and come with high costs in testing the most advanced AI systems. To address these challenges, we introduce HiBayES, a generalizable Hierarchical Bayesian modeling framework for AI Evaluation Statistics. HiBayES supports robust inferences in classical question-answer benchmarks and advanced agentic evaluations, particularly in low-data scenarios (e.g., < 20 data points per evaluation). Built on Generalized Linear Models (GLMs), Bayesian data analysis, and formal model comparison, HiBayES provides principled uncertainty quantification and robust parameter estimation. This paper offers a comprehensive introduction to HiBayES, including illustrative examples, comparisons to conventional statistical methods, and practical guidance for implementing multilevel Bayesian GLMs. Additionally, we provide a HiBayES software package [4] (Beta version) for out-of-the-box implementation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Splitting an attack across more coordinating agents lowers per-commit monitor suspicion, making distributed attacks harder to detect than single-agent attacks.

  2. Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

    stat.ML 2026-07 conditional novelty 5.0 of 10

    A hierarchical Bayesian HMM is introduced to capture sequential dependence in LLM task outcomes, and experiments suggest that ignoring this dependence yields overconfident reliability estimates.

  3. Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

    cs.CR 2025-10 conditional novelty 5.0 of 10

    A Bayesian model that groups similar LLM test prompts into clusters gives better predictive scores than a no-clustering baseline but does not prove that it truly corrects prompt dependence.

Pith tools