Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:10.894434Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2505.15240.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:10.894434Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T06:49:20.019956Z
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eadde58f-0925-492b-a928-73b06b038893 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 467dbc23-a0d5-4c5f-9049-756a0e40b7c0 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Categorical data analysis
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db7db87a-aea4-42ba-bf0e-65311de07201 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge A computationally intensive ranking system for paired comparison data
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96127373-87f6-45f5-a8eb-cbeca7724a9f · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank analysis of incomplete block designs: I
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d1a2ed-7104-4683-83e3-433242852d34 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5936fee-505a-43d0-9255-cb5dda048f1c · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d8571d4-de8d-4bcf-a2d5-ef4679f4ee63 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank: from pairwise approach to listwise approach
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 356dde46-a9d5-46f6-888c-3f5bdc7e283d · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient bayesian inference for generalized bradley--terry models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c945fa-246a-4147-8f9f-64fecc9a50e3 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Models for paired comparison data: A review with emphasis on dependent data
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e7ab943-58d5-4e45-8989-d6ed586af534 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Humans or llms as the judge? a study on judgement biases, 2024 a
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d7ef445-02df-496d-bb81-bf337597a535 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Premise Order Matters in Reasoning with Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0015be6-7a2c-4c43-a796-579b9d683950 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Of human criteria and automatic metrics: A benchmark of the evaluation of story generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9799cb4-970b-4e35-ba8e-c3918c478614 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Can Large Language Models Be an Alternative to Human Evaluations?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6691a0-6d51-4edd-9f5e-8926df2563c6 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Scaling instruction-finetuned language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a6aa0b7-0ce2-484c-9fec-a6d176015723 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Ranking by pairwise comparisons for swiss-system tournaments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6069fd9a-d81c-47ed-9d81-93c7aaf9f6cc · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The method of paired comparisons, volume 12
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 446665cf-8dd5-4de3-8efc-e0ff5235425c · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The Llama 3 Herd of Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2db2e4d-2a6d-443b-90f4-870843a00bba · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank aggregation methods for the web
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da8ea7c2-855e-4f5a-bf8a-017ff2f09038 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Summeval: Re-evaluating summarization evaluation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8042578b-947a-4264-b9c7-a3660f7d2c4f · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPTScore: Evaluate as You Desire
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b01b668-9df0-496a-961b-60ca7706f7f0 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill™: a bayesian skill rating system
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b56986d4-cd59-4ed1-8c13-8dfcba911fe0 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6ff5acb-ce24-4ec7-8b84-622957033dce · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339a8c9d-66c9-4d53-acbf-8cf41c0a0d9b · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b61eed-2a0f-4020-a8a6-cbf2e2bca907 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank for information retrieval
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fffee44e-23f9-4fd5-a59a-e7f9814163e9 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge G -eval: NLG evaluation using gpt-4 with better human alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6340000d-5367-4002-8518-90d0d56a1d97 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15275cc5-d69f-428f-9750-5ce9be79c1f0 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4dc69aba-b45f-4250-bafe-7e67660805b5 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2da9566a-ddd2-4c3b-b617-7e6400494f5a · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d705c09-3f70-4972-8cf6-05be56ecf021 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stated choice methods: analysis and applications
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604b0b71-aa6f-440e-9a13-713aa2d9803b · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The structure of random utility models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5bcca16-bf0f-4c06-8e03-87463ceed9af · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill 2: An improved bayesian skill rating system
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ed99d0f-dab3-457f-8066-f050f4cb1d0b · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient computation of rankings from pairwise comparisons
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3c346b9-2726-4ab0-a5ca-f87e319bb147 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Training language models to follow instructions with human feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a898b7d7-b1e6-4c88-990c-ea67df9c3fb8 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1e43ffc-2cc6-4775-9367-77bafad316bf · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45b21d5-04e9-4df4-9199-7dcac28bfe06 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Qwen2.5: A party of foundation models, September 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1abfbcf-c988-4af7-b6c5-960383d42a60 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Finetuning LLMs for Comparative Assessment Tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b2d98d3-156a-4685-a01b-f136671e54a9 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stanford alpaca: An instruction-following llama model, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef6c73d-62fc-4021-b73b-af03aada0b01 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Is ChatGPT a Good NLG Evaluator? A Preliminary Study
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62c44d6-1bf2-4314-b9dd-834bdb2ca0b6 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large language models are not fair evaluators, 2023 b
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d010cadd-bc90-4b67-98f0-71747f35ac9e · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Primacy effect of C hat GPT
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab765c02-662e-40be-9d11-00c87efb2024 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96b7956-0621-4486-b72c-160692212958 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 240c239b-42aa-4f59-a97f-d2e88c222feb · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af9ba3e4-36ba-4be3-85fa-bb043de809d4 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3954a661-deb0-4836-96b3-456b129590b5 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Lima: Less is more for alignment
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b426243-0836-4127-a569-ebc5eb8d8e14 · outbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judgelm: Fine-tuned large language models are scalable judges
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27443c4a-b1ef-430f-89dd-8f762f430d72 · inbound
A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.