Pith. sign in

Paper Citation Record · LEDGER

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

As of 21 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 5 inbound Pith citation observations for arXiv:2505.05602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05602 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:00.821737Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:19:47.155013Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T13:26:16.578885Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b992279-ef0f-4ea2-8cb9-195731264452 · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku Anthropic.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics The Claude 3 Model Family: Opus, Sonnet, Haiku Anthropic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.480096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.676900Z digest=sha256:2b2fa9c5908b294cd474bb58e32c4aa29df3a60861aef39a5a226a60b7593340

Observation 1a954fc1-b84a-49dd-8ba5-d53d4b028d2f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Evaluating Large Language Models Trained on Code

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.682344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.682344Z digest=sha256:fd58432a536eba60abf30bb4cda5971e82a0b9cf709369361353a5b92bcf09d6

Observation 27384ec8-b8a2-46d1-9c7a-2eefb30f2a05 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.687781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.687781Z digest=sha256:7a2f9f5f9a50098492455a1293305988fa9f1102875a533cb2e4d14d54c16fb4

Observation 1b5264fc-b950-47c3-96e3-f8d7c9a38a55 · outbound

This paper cites HiBayES: A Python package for analysing data from Inspect logs using statistical modeling techniques.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics HiBayES: A Python package for analysing data from Inspect logs using statistical modeling techniques

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.463513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.692755Z digest=sha256:dfef68c75c7b09e33f38ce0fe0ca45be0a72da364a505b6da150648f16b4ef89

Observation d673f58a-b21b-487d-b6aa-667b7e4c7ca2 · outbound

This paper cites Bayesian Data Analysis.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Bayesian Data Analysis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.447681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.697420Z digest=sha256:5fd2ca618621f69a5312a1dff180ec65e612e35731c4f6c6f3032e345f5e096e

Observation 80936299-fd0c-470d-8416-983aa8da0f56 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-world GitHub Issues? https://github.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics SWE-bench: Can Language Models Resolve Real-world GitHub Issues? https://github

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.432468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.702596Z digest=sha256:102a5acdbe63991460a3f2ad232694657b6778373fe9e785c83e07e1d867e425

Observation 5dea24d3-4a55-4bd2-b448-77190f63fd58 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Measuring Massive Multitask Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.708335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.708335Z digest=sha256:10b8b5bd3818ddb8e0827d5d25bcc3a650b5175d2858bc3a785389cd75d33f8b

Observation c2c02a5d-792d-45fc-9b95-1cd52e1737e6 · outbound

This paper cites The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.713179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.713179Z digest=sha256:8c076c29715723a605cfe259b3fff6b137da39e89ec06c2be5c38963bdd7d87a

Observation 11046e32-51a9-4f0d-ba75-9aa8aaa33a2d · outbound

This paper cites Analyzing OpenAI’s o1-preview in Comparison with Claude 3 Opus and GPT-4 on ARC AGI Evaluation.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Analyzing OpenAI’s o1-preview in Comparison with Claude 3 Opus and GPT-4 on ARC AGI Evaluation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.415575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.718200Z digest=sha256:05a7374732b7e2504c2b3f6cca779f34e268ca8e376d507edd963345987a7595

Observation b4066978-fed2-489a-875a-770c57d75235 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Measuring AI Ability to Complete Long Software Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.723242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.723242Z digest=sha256:f58522fe523388d314cefae8c9cda263360715077059c36f7e7bc7c3ac508a34

Observation acd725f4-78e9-4ebf-9e32-e3cc7679d014 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.728561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.728561Z digest=sha256:ee5dfc034713fc1abd193fc78819512a4d20fe71940724c64b7756456b912d5b

Observation 52cc77eb-7348-4177-9cdf-f0465efe70a9 · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.733761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.733761Z digest=sha256:3108140548f7fff92414a0b789c82bcdaf87af64b6f382cb92b028e519295d0e

Observation 71bb4ed7-d858-437c-b4b4-a8eece49a1ff · outbound

This paper cites Holistic Evaluation of Language Models.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Holistic Evaluation of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.739467Z digest=sha256:8ec47e1b9deafb27fb08c2bf2297d1776c24f72c6f60e2fd695d40946dc0616b

Observation 1fb628ed-4f58-47a9-b4b9-e1ccdc356764 · outbound

This paper cites Statistical Rethinking.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Statistical Rethinking

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-15T23:06:01.181796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.744516Z digest=sha256:b56ba305c05dae32414eb2e1a7d0dfb817d97bb6e2d916b98f5aea0843326da1

Observation 96fa30b8-495a-4ed5-bf6c-4a5a33a080ed · outbound

This paper cites Statistical Rethinking 2023.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Statistical Rethinking 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.397270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.749395Z digest=sha256:37bcafe5ce4977de2ef4e6c3cdc47ae9b20de4d06fc6cd95f8435df67e7c0752

Observation 12fe11be-7608-47b9-aa90-6e3139fe0050 · outbound

This paper cites METR’s GPT-4.5 pre-deployment evaluations.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics METR’s GPT-4.5 pre-deployment evaluations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.381825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.754016Z digest=sha256:6ff791eb2bfacd85970bbee3ba4874b11e4be5362cbd12b8a3a1f158f28e8e91

Observation 9836eb00-9250-43df-bd34-e45c04006f28 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics GAIA: a benchmark for General AI Assistants

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.758834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.758834Z digest=sha256:989eca008acdf8e4bccdee749a7f448236afdd7f364360c4c4ff063e5b5e965b

Observation 374b41c5-137e-4c26-9246-1802793bcb96 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.764019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.764019Z digest=sha256:836ede00b722a5295a38618ddf2c336ea6375e787886a0ac00c4ca419073aa1a

Observation 3cdcd32d-ecd6-4ae8-8eb7-3255ff978764 · outbound

This paper cites Generalized Linear Models.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Generalized Linear Models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.365363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.769405Z digest=sha256:3274cc04474420366b4f5170962f6d6f10cacb6c396ba0d43db505274235cd38

Observation e0634780-3ced-43ec-9b6b-c21d531641dc · outbound

This paper cites GPT-4o System Card.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.774258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.774258Z digest=sha256:28e2f6b2648405b22cbd2c963d5e779552c22e33e2bced3500d64540bbb96c1a

Observation cc59f489-bb5d-45c5-90c9-49d8e8261ecf · outbound

This paper cites Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.779556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.779556Z digest=sha256:8bece603e942b70c8969e9630a053bdf3f406e81d93039a7c8602b7157baad81

Observation fa7e32e5-50b3-442d-b2ff-8deaa137af63 · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics HCAST: Human-Calibrated Autonomy Software Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.784797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.784797Z digest=sha256:86e71e26e927afdee161ecb7c44ddb2c0a40735a8c3eb02dff317ec8ae58fba0

Observation 21467e2f-7281-40f2-90ba-213d476f0093 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Language Models are Multilingual Chain-of-Thought Reasoners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.789738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.789738Z digest=sha256:aa3acf0dc716562af904459929ab15a00b1b3c5de449acf5cb4e2bad554c89a2

Observation 01837506-8aa0-4401-8ea4-3dec0ab627f1 · outbound

This paper cites Clio: Privacy-Preserving Insights into Real-World AI Use.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Clio: Privacy-Preserving Insights into Real-World AI Use

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.794746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.794746Z digest=sha256:99bd2590c9fe34789a81038ae419d9d110ad4f133dfa7350ba77277e85838342

Observation eaffacaa-0ea1-4dbb-a666-120a126d8421 · outbound

This paper cites inspect_ai: AI-assisted inspection policy development.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics inspect_ai: AI-assisted inspection policy development

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.349649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.800737Z digest=sha256:ceea3d62a373af9a7a9c5af40a161e1e4ce0e5bbf8090683fcbe13ac6d1b39f2

Observation 765e2c86-5715-43ce-b82c-a47b7c936a8a · outbound

This paper cites Asymptotic Equivalence of Bayes Cross Validation and Widely Applicable Information Criterion in Singular Learning Theory.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics Asymptotic Equivalence of Bayes Cross Validation and Widely Applicable Information Criterion in Singular Learning Theory

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:01.332331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T23:06:00.806930Z digest=sha256:e5c71b398e7f0341b0f810be5475d7b4741d5d019c0a7310dc9c505d8a3ffacb

Observation 3c78b97b-aa03-42e4-a582-56dfac4df817 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.812017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.812017Z digest=sha256:15c7924a67143d04db1fef0e640670d764cab58b53112a5c90beba0b8b76459b

Observation 512fc14f-2ee8-4ff8-9bdd-c5f840592840 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics ReAct: Synergizing Reasoning and Acting in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.817052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.817052Z digest=sha256:3281027ecdb54c52464ab023aa6bf8a75b3cea829367057c0a424407419e5789

Observation e75e17fe-3ce5-4b1f-9c40-1b8d65f3dfce · outbound

This paper cites HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.821737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.821737Z digest=sha256:bc6450fefb150434e74cde4b34a0f5ed74015d4ca97b808c925041fa7707e37f

Pith citing papers

Observation 2801fa99-1e31-4692-aea1-f4a2389eaa15 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:47.155013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:47.155013Z digest=sha256:364654155ac89bacb80318596a0f20a282f4e76e7592691c2d1043fcb02a5de4

Observation 1da28fec-a90b-44bd-84b5-e19a96647d9b · inbound

People readily follow personal advice from AI but it does not improve their well-being cites this paper.

People readily follow personal advice from AI but it does not improve their well-being HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:12:05.923970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T21:11:04.614448Z digest=sha256:4d89acbda629e1d0a3072dd3895c46ec2478a2d547551cc66964910e48853853

Observation b12517b5-34a7-45e6-b613-6f37e875ce02 · inbound

CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation cites this paper.

CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T03:28:10.067536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T18:10:55.215746Z digest=sha256:14efe4c023ba050c8eaa2fd7e718ad65f87571cdb438f4a31d84da6f3daa6e80

Observation ba9cb7f2-d722-4af3-91d8-159fc88940f0 · inbound

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors cites this paper.

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T13:26:16.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T13:21:01.483489Z digest=sha256:49bc779fb929dba11cf40eac70138fceaf35e8a8443d8d3ae43e61eb4cd1b58d

Observation 60843e07-56a9-446a-bbe6-974cbf04c93a · inbound

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model cites this paper.

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T04:06:15.154820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:06:15.154820Z digest=sha256:22973674adde0e28170157ae7ab9f92ae6903dc8a71bc87407285a87a7b639fa