Pith. sign in

Paper Citation Record · LEDGER

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

As of 20 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2505.15240.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15240 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:10.894434Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:53:02.187519Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T18:53:02.683433Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact4
  • verified fuzzy21
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eadde58f-0925-492b-a928-73b06b038893 · outbound

This paper cites GPT-4 Technical Report.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:07.772078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:07.772078Z digest=sha256:33370a264676d403cb2e6c2aee0873dff81822711e8b97eb0e4fffb963ecabca

Observation 467dbc23-a0d5-4c5f-9049-756a0e40b7c0 · outbound

This paper cites Categorical data analysis.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Categorical data analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.953900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:07.872856Z digest=sha256:cb878f1f33ed3fab60301c2c49bede683adabcb52c73b9c108882c14f40055d5

Observation db7db87a-aea4-42ba-bf0e-65311de07201 · outbound

This paper cites A computationally intensive ranking system for paired comparison data.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge A computationally intensive ranking system for paired comparison data

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.838672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:07.925438Z digest=sha256:3afaeb053fbf62d77d0c21a09c470acbc62e3d55b8d3b3a5d79a580cd6e576d9

Observation 96127373-87f6-45f5-a8eb-cbeca7724a9f · outbound

This paper cites Rank analysis of incomplete block designs: I.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank analysis of incomplete block designs: I

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:07.965543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:07.965543Z digest=sha256:72251733afd56149ffdd1e54a52db8933ef8b0b4adb4cd6fba87fabb91501ee6

Observation 35d1a2ed-7104-4683-83e3-433242852d34 · outbound

This paper cites Language models are few-shot learners.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.029766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.029766Z digest=sha256:bd54903b2b3848c32622fad9c78b1adc68bb59ac246ee9bcbf8f4eb958a77b6e

Observation e5936fee-505a-43d0-9255-cb5dda048f1c · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.130142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.130142Z digest=sha256:a218e514794417a2f5d709d78db24c6808cb73749714f36c7573c6366c6d567f

Observation 5d8571d4-de8d-4bcf-a2d5-ef4679f4ee63 · outbound

This paper cites Learning to rank: from pairwise approach to listwise approach.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank: from pairwise approach to listwise approach

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.179818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.179818Z digest=sha256:fb120933ab257618b096a7d37cd020291c4e237dbcc347aada06d113424b99a1

Observation 356dde46-a9d5-46f6-888c-3f5bdc7e283d · outbound

This paper cites Efficient bayesian inference for generalized bradley--terry models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient bayesian inference for generalized bradley--terry models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.252459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.252459Z digest=sha256:a1811fcfbd909d5cd42e761ed218a350603eef188ccdd039a5ac66effccdd831

Observation 45c945fa-246a-4147-8f9f-64fecc9a50e3 · outbound

This paper cites Models for paired comparison data: A review with emphasis on dependent data.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Models for paired comparison data: A review with emphasis on dependent data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.688879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.346278Z digest=sha256:df94c602818f4c6eac2f43e8a998a6ae6d98d5ca86750b2f59e050aa719ff462

Observation 1e7ab943-58d5-4e45-8989-d6ed586af534 · outbound

This paper cites Humans or llms as the judge? a study on judgement biases, 2024 a.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Humans or llms as the judge? a study on judgement biases, 2024 a

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.587348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.409355Z digest=sha256:8070ec059a9dd940b34ab041d68e74737281e2230c6f687448ea81a43b38019b

Observation 0d7ef445-02df-496d-bb81-bf337597a535 · outbound

This paper cites Premise Order Matters in Reasoning with Large Language Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Premise Order Matters in Reasoning with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.450887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.450887Z digest=sha256:5cfb3860dda1eed38e0cf9a6c221fb6d10d00ebb26cf6600b8d7e9f24d70d078

Observation d0015be6-7a2c-4c43-a796-579b9d683950 · outbound

This paper cites Of human criteria and automatic metrics: A benchmark of the evaluation of story generation.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Of human criteria and automatic metrics: A benchmark of the evaluation of story generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.488265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.502773Z digest=sha256:31a53643154cac5e9e369ca2b0b0f751cc93c0d66bc8588814525b34019f6f97

Observation f9799cb4-970b-4e35-ba8e-c3918c478614 · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.594167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.594167Z digest=sha256:1e11a1da3965389b7ad9afb5d68508159ff6a78e8659f4bb6f26bb4c1788385e

Observation 2d6691a0-6d51-4edd-9f5e-8926df2563c6 · outbound

This paper cites Scaling instruction-finetuned language models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Scaling instruction-finetuned language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.397921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.640171Z digest=sha256:e7a4f7cdc54813c04352ab348609a40b844b7ef67fd20475b78600614e450b1d

Observation 1a6aa0b7-0ce2-484c-9fec-a6d176015723 · outbound

This paper cites Ranking by pairwise comparisons for swiss-system tournaments.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Ranking by pairwise comparisons for swiss-system tournaments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.315039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.679816Z digest=sha256:65dd8be4c0819ed88c75d2f52e0ff923ab1ad8e0decba09d197ea33eaee476f0

Observation 6069fd9a-d81c-47ed-9d81-93c7aaf9f6cc · outbound

This paper cites The method of paired comparisons, volume 12.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The method of paired comparisons, volume 12

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.212792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.741755Z digest=sha256:331a29fa6fd82972023c337d45b9a6743aad39018ebd899f12a4674d9e86d6bb

Observation 446665cf-8dd5-4de3-8efc-e0ff5235425c · outbound

This paper cites The Llama 3 Herd of Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.826293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.826293Z digest=sha256:cbe525e44ce5b6486bcc0f47b0745889dc6e85b4e7509521d29c563984d14448

Observation f2db2e4d-2a6d-443b-90f4-870843a00bba · outbound

This paper cites Rank aggregation methods for the web.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Rank aggregation methods for the web

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:14.106964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.860405Z digest=sha256:91ff761f4da12ed53417dfc7dde9733d23f9dd068ef912d199ec7febccfb2c6a

Observation da8ea7c2-855e-4f5a-bf8a-017ff2f09038 · outbound

This paper cites Summeval: Re-evaluating summarization evaluation.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Summeval: Re-evaluating summarization evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.961343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:08.964922Z digest=sha256:54e43fbf1de42ef96151f183dfee1b61514d05f9bdf6114160f87dc6f850d1d4

Observation 8042578b-947a-4264-b9c7-a3660f7d2c4f · outbound

This paper cites GPTScore: Evaluate as You Desire.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge GPTScore: Evaluate as You Desire

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.027206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.027206Z digest=sha256:25090b8990576844ed1f4d4f1a634418cc7bcff602e6920b96602f22e91841bd

Observation 8b01b668-9df0-496a-961b-60ca7706f7f0 · outbound

This paper cites Trueskill™: a bayesian skill rating system.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill™: a bayesian skill rating system

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.835903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.064810Z digest=sha256:83cf6361e3bcd96bbf0b164ba4fe4ba75a662ff201b248f88e9d7328df283490

Observation b56986d4-cd59-4ed1-8c13-8dfcba911fe0 · outbound

This paper cites an unresolved cited work.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:13.709267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.174549Z digest=sha256:08e11b777829fa5ed97ab81d913585343f10cb0834a89a1e15e1c7b56e2cd5c0

Observation e6ff5acb-ce24-4ec7-8b84-622957033dce · outbound

This paper cites Large Language Models Are State-of-the-Art Evaluators of Translation Quality.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models Are State-of-the-Art Evaluators of Translation Quality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.226377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.226377Z digest=sha256:2b0d137413a4980b3a88a5e118f9e2baf7f5eca42c3d243bd3dcd88f264793c9

Observation 339a8c9d-66c9-4d53-acbf-8cf41c0a0d9b · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.268446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.268446Z digest=sha256:b252de8902924c24d500c86e88d90d947ddc27fecd5014bb2a976a52d79b0558

Observation 73b61eed-2a0f-4020-a8a6-cbf2e2bca907 · outbound

This paper cites Learning to rank for information retrieval.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Learning to rank for information retrieval

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.310157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.310157Z digest=sha256:8243fe3f90bfd7ac9a281c6311a855ac8b1b8e144f5c660370884d1eaf9dfb29

Observation fffee44e-23f9-4fd5-a59a-e7f9814163e9 · outbound

This paper cites G -eval: NLG evaluation using gpt-4 with better human alignment.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge G -eval: NLG evaluation using gpt-4 with better human alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.389780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.389780Z digest=sha256:d988f1fee3dafef6669092e4a7f0405d15f910b901484ecd6056198c0fb9e71b

Observation 6340000d-5367-4002-8518-90d0d56a1d97 · outbound

This paper cites Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Aligning with human judgement: The role of pairwise preference in large language model evaluators, 2024 b

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.546621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.452271Z digest=sha256:d1456de43a5a73687f24aa7912ed10864a44739e0aeff19b8d56dd207136abd1

Observation 15275cc5-d69f-428f-9750-5ce9be79c1f0 · outbound

This paper cites Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.788655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.490840Z digest=sha256:ec42e2e9ed740c0c1a2624671c85c078d62759494d10f444ea9d35a27ae1fa56

Observation 4dc69aba-b45f-4250-bafe-7e67660805b5 · outbound

This paper cites LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:13.241221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.555067Z digest=sha256:467ef4eefe284745b1c113f90a1a7f031492c907c6d7e1f62d304f224895f5f3

Observation 2da9566a-ddd2-4c3b-b617-7e6400494f5a · outbound

This paper cites Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.676956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.676956Z digest=sha256:d672d2ac5c555a780e50de88a9dbe06c14cadc745e93efadc0c261929f700de7

Observation 8d705c09-3f70-4972-8cf6-05be56ecf021 · outbound

This paper cites Stated choice methods: analysis and applications.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stated choice methods: analysis and applications

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.992603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.717192Z digest=sha256:ac60ff708419453c3e9e527ce044f9424a5ce6def911a83eca98592295edc9a5

Observation 604b0b71-aa6f-440e-9a13-713aa2d9803b · outbound

This paper cites The structure of random utility models.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge The structure of random utility models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.905404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.761167Z digest=sha256:397f398d8d37ca53653af0dc277b9c33030fa9aaa42c96f67e9fae9a92e73cd1

Observation e5bcca16-bf0f-4c06-8e03-87463ceed9af · outbound

This paper cites Trueskill 2: An improved bayesian skill rating system.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Trueskill 2: An improved bayesian skill rating system

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.795135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.855334Z digest=sha256:a8ddb42e456ed88ae2fb66a47c4beb1685f74a24e3204055aadf64bcb5bd39b7

Observation 1ed99d0f-dab3-457f-8066-f050f4cb1d0b · outbound

This paper cites Efficient computation of rankings from pairwise comparisons.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Efficient computation of rankings from pairwise comparisons

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.683651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:09.912909Z digest=sha256:9c2699f62858f963cda130a67ed528ab77e960515a59d2890d9e823b78f8db7b

Observation f3c346b9-2726-4ab0-a5ca-f87e319bb147 · outbound

This paper cites Training language models to follow instructions with human feedback.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Training language models to follow instructions with human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:09.957908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:09.957908Z digest=sha256:ad2b5634412e83b300a7a88443b47145e97a652ddad814bd51c9e0dc7820b7c8

Observation a898b7d7-b1e6-4c88-990c-ea67df9c3fb8 · outbound

This paper cites PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.613917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.001331Z digest=sha256:bb2856d072074fd551ab70fa9b7c5d6138cff918851e70cbcd4ea1b84ad1d8c5

Observation e1e43ffc-2cc6-4775-9367-77bafad316bf · outbound

This paper cites Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.098915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.098915Z digest=sha256:4a1d3600b00d384be38912dbbf2a710615ddb366630b8a82e3ca16c37d28b1ea

Observation a45b21d5-04e9-4df4-9199-7dcac28bfe06 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Qwen2.5: A party of foundation models, September 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.160968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.160968Z digest=sha256:8398d0f71d3e8b2247e7eba9f6ec0b1caf43015b59241f5509b0553692aaaf05

Observation f1abfbcf-c988-4af7-b6c5-960383d42a60 · outbound

This paper cites Finetuning LLMs for Comparative Assessment Tasks.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Finetuning LLMs for Comparative Assessment Tasks

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:11.376856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.202374Z digest=sha256:91451967eb1d76f4e2cd605ed722ca21aa938fb43b38d83e5f473a3091238ccc

Observation 1b2d98d3-156a-4685-a01b-f136671e54a9 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Stanford alpaca: An instruction-following llama model, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.260157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.260157Z digest=sha256:d5034d71b6ef9d3cedb6573a11e79a45936b9078c36c9b23248434d7f6afbf80

Observation 9ef6c73d-62fc-4021-b73b-af03aada0b01 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.352664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.352664Z digest=sha256:51b8145e11598c5c7945b58ee320ff1ea642a949e716dd3b2a8c051a0a2f3111

Observation e62c44d6-1bf2-4314-b9dd-834bdb2ca0b6 · outbound

This paper cites Large language models are not fair evaluators, 2023 b.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Large language models are not fair evaluators, 2023 b

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.506759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.387688Z digest=sha256:c8a92fba1f1fb1228932517818c76d106b9401ba1c409bfbc859c9a6b00c9467

Observation d010cadd-bc90-4b67-98f0-71747f35ac9e · outbound

This paper cites Primacy effect of C hat GPT.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Primacy effect of C hat GPT

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.453435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.453435Z digest=sha256:98d5cf1a95517bde113079baac0bb07d71dd0e849098eea41e86e4ca2f176558

Observation ab765c02-662e-40be-9d11-00c87efb2024 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.548443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.548443Z digest=sha256:6fb6efedeb319f6770ae0d60ae3ddb72fc42eede2997bb7b6e669d488cd19232

Observation e96b7956-0621-4486-b72c-160692212958 · outbound

This paper cites an unresolved cited work.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Unresolved cited work

Reference 45

Resolution
verified exact
doi, observed 2026-08-07T15:28:11.117158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.610236Z digest=sha256:33f8177fa65f3b4eaee133acc3582768d847e2e32048151afb1ec3a8a56bb421

Observation 240c239b-42aa-4f59-a97f-d2e88c222feb · outbound

This paper cites Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Die berechnung der turnier-ergebnisse als ein maximumproblem der wahrscheinlichkeitsrechnung

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.399371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.723001Z digest=sha256:dd24878eadedd0b5e7ba429ac47add2625eca61fdd58886891fa766364b0134b

Observation af9ba3e4-36ba-4be3-85fa-bb043de809d4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.779957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.779957Z digest=sha256:8090fe4787917d289938db16ee77bc568b447b43d07ad1eda6cb92a2587ef448

Observation 3954a661-deb0-4836-96b3-456b129590b5 · outbound

This paper cites Lima: Less is more for alignment.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Lima: Less is more for alignment

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.275766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.828110Z digest=sha256:23a33f29ca0b1faac74253cb29ffef614be48dcd8aa883ec1933edc34a4f95fd

Observation 9b426243-0836-4127-a569-ebc5eb8d8e14 · outbound

This paper cites Judgelm: Fine-tuned large language models are scalable judges.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Judgelm: Fine-tuned large language models are scalable judges

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:28:12.145963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T15:28:10.894434Z digest=sha256:10704134c34842a054b1a41f6b36c3bcd4828b8a385a265ed0ab82d5d719ac99

Pith citing papers

Observation 27443c4a-b1ef-430f-89dd-8f762f430d72 · inbound

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth cites this paper.

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:49:20.019956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:49:20.019956Z digest=sha256:7ec6f631c959658d9a5640bd184e823a8564471f77225b1a080f80864fb95272

Observation 8fec4e1d-53fa-47cd-94f0-2b042fbe1a96 · inbound

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation cites this paper.

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T18:53:02.689460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T18:53:02.187519Z digest=sha256:0d704993f8a8c85100d2e01d1b5f08912dc1a9426bf661be3f9398baf544af38