Pith. sign in

Paper Citation Record · LEDGER

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2607.21632.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21632 v1

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:57.063622Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 174ecd00-7ed2-48dc-8d1a-9a1ee4ae157b · outbound

This paper cites Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.472986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.472986Z digest=sha256:d7a9fcedb519cb3504f31ef1fe8b305503d20b76d8ed8f658d0ac1ddb8de6433

Observation 3863f4a8-1c37-4280-bf95-f1cff6be3e64 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.566993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.566993Z digest=sha256:b37dd545070291d4756f15e2d6040149b25256c634bea86a8c2a6230c29a5c84

Observation d4ba2c6a-5722-4e51-9c59-dc4c612f51d3 · outbound

This paper cites an unresolved cited work.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.657906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.657906Z digest=sha256:ec0dba093f86db2b0198b2f8661ab3ce56ed50961f3a8bec77b5d8c5987ecc33

Observation 14609784-9758-4790-ba70-7261c3ee26a6 · outbound

This paper cites 2025 , url=.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models 2025 , url=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.725094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.725094Z digest=sha256:24894e8eeb3f767a68d53296825b5c303d9657019961439e7ae750eabe048890

Observation 9baf0bb4-3b8e-441c-af2f-19c91a6a3c4a · outbound

This paper cites ResearchGate , month=.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models ResearchGate , month=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.808407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.808407Z digest=sha256:a060af8f1732a9d9941e625ddcd3721a89c2d9fdde878aedd79fd6c9624c9d11

Observation 7aeec3ce-04c7-45d7-b70f-9361993fc159 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.897891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.897891Z digest=sha256:33715e69f1483fda205204aaf5955e2c2958408f0dc37baac55f919018d4932b

Observation 9108a4be-c59b-483b-a509-0c42c87b153b · outbound

This paper cites arXiv:2510.09738 , year =.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models arXiv:2510.09738 , year =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.984523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.984523Z digest=sha256:157160423796e6dcfc5bd7ac5d50b1cbf234737fa829b3db274e71860a270ee2

Observation 060b097a-2e9d-4b3c-bd3d-008938d198e3 · outbound

This paper cites PLOS , url=.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models PLOS , url=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:57.063622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:57.063622Z digest=sha256:cb87a5de2b383d9e721afd68084036bdd15cdf84adcd42b7e61f4b36d3046411

Pith citing papers

No inbound Pith citation observations are available.