Pith. sign in

Paper Citation Record · LEDGER

Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2403.16950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.16950 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:14:57.387370Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a413f4a9-afcb-478d-8dcb-db1cbab5a1b4 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.365470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:dfded340f3264b407877d0711c68ba4bfdb1fdf59cc9a53c91f466cb0fcabba4

Observation 77254223-cb60-4048-bf6d-6f6806b002ac · inbound

AI Alignment at Your Discretion cites this paper.

AI Alignment at Your Discretion Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T16:14:57.387370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:14:57.387370Z digest=sha256:cfccb4c6cd85478524113a7933529774c8216f8a32fa594d77e85e86f26752af

Observation 045f3ca9-c561-4f3f-a762-ee3b14c03279 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:35.554333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:35.554333Z digest=sha256:0e090505180a8ea0a52566139349f41f60384743e77db4d7f94e0fb119f11c88

Observation daa50822-7eaf-40f1-ba4a-31023c04a764 · inbound

Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases cites this paper.

Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:45.759286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:45.759286Z digest=sha256:d814166f69880464015126174d210eb2f53ee64bfb29ca369b6d47aa38bc0771

Observation 02c6a2a9-05d9-4fdf-b7ce-ef5d54c8ffc8 · inbound

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models cites this paper.

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:07:14.262841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T10:05:22.562737Z digest=sha256:b223ba2132e3cb9ce681e764c68c4e9e7a580a144d6f3feaf5021bb48db1fe78

Observation 45712385-d5db-491a-9802-e76798fa58b0 · inbound

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models cites this paper.

Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:01.132316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:01.132316Z digest=sha256:55b15e4d0bfcd2556a8a07efccfa7c946f5a15dd07cbae49f0577d092c465cd7

Observation 24071394-692b-4ee8-9aaf-90635da0d43c · inbound

Deep Researcher with Test-Time Diffusion cites this paper.

Deep Researcher with Test-Time Diffusion Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:21.497194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:24:21.497194Z digest=sha256:3afadd74e247e8737b4bb7ca47c4b29fe189e6cdd558f6961f7d668bd12f277d

Observation 9b4a0313-598a-45b9-a452-6ac597410690 · inbound

Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization cites this paper.

Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T16:13:49.902821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:13:49.902821Z digest=sha256:47f8f3f440ed504917841585a2e4dee62a43b9725a70220fb695f9fa11acda17

Observation 23d6e99e-59bb-4bdc-9aae-80085bd46aba · inbound

Semantic Data Processing with Holistic Data Understanding cites this paper.

Semantic Data Processing with Holistic Data Understanding Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:08:09.534763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:07:49.756349Z digest=sha256:d405d4c02517ab0a04693ff4502e4cfd56ee078dbe1152b9818d704d7a170b5d

Observation 333773b1-c3c1-434b-95d2-a90fbc04dafe · inbound

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild cites this paper.

Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:27:48.133595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:26:38.634540Z digest=sha256:ec9d0c439b47bc97ffee9457ffad34a5ea8b51c6f43abe091c6f3ad3907f8744

Observation a1076fd7-f506-448d-837a-fac402bdfa98 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.214388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:c6b985f99f5feb3135196b55934b38bcef4199a5790c0fa86c58812dfa64698b

Observation 4ad14bc4-2452-4807-a925-15e9573003cc · inbound

GRASP: Deterministic argument ranking in interaction graphs cites this paper.

GRASP: Deterministic argument ranking in interaction graphs Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:14.909207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:57:55.198779Z digest=sha256:04d05899b3372dd7d80849c385ed9f9e41a9e44c87d010c47f5451b8358cfbef

Observation a917ec0f-6780-4430-8a9c-2f8683a63e5d · inbound

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges cites this paper.

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:36:47.723043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T05:58:59.870335Z digest=sha256:030c473dd0b6fc64e81bdd9ac687bc50272b8c0c3ba37388d7a1c1a7b67a35bc

Observation 99177898-7b7c-45c6-9c21-ca4dc5678be1 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:47.116638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T07:49:36.816100Z digest=sha256:85fa834d59bfb079804a6fc9d70b81227424db012d2277a6bf9d7c560f2c5ecf

Observation 895e2b5d-14bb-41bc-aa26-1af41d739c06 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:39.390666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T10:17:33.176525Z digest=sha256:51f6e8384bf45ccc447d9240464574604c25ba53d0fdbc5b22943d3fbd33c94c

Observation 62ed0cd6-59b0-4254-ba16-c3d9dd522f82 · inbound

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank cites this paper.

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-26T03:58:57.117163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T03:56:28.271760Z digest=sha256:0db858e91c4dfa216a6b7f0c59f6b9660a3f6f27f6adc01c408f6f4266815006

Observation 4c09fb6f-9b4e-43ac-973a-10cf1f96a2c3 · inbound

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 cites this paper.

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:23:00.942005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T02:17:26.872854Z digest=sha256:7dd7e896d2bb5e396f1d78d4bb13272325bb298323091126ef9c1ceb11d683ef

Observation 323dacc8-4392-4f0d-b5ad-4560fa5a6002 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.163786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.163786Z digest=sha256:fa61d3ffa52947c426a908fcbb3e70e06be82bf9a080274541bae781a5799db5