Pith. sign in

REVIEW 7 cited by

Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.16950 v5 pith:VSWSIS5G submitted 2024-03-25 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmsevaluationevaluatorspairwisehumanlanguagepairspreference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated promising capabilities as automatic evaluators in assessing the quality of generated natural language. However, LLMs still exhibit biases in evaluation and often struggle to generate coherent evaluations that align with human assessments. In this work, we first conduct a systematic study of the misalignment between LLM evaluators and human evaluation, revealing that existing calibration methods aimed at mitigating biases of LLMs are insufficient for effectively aligning LLM evaluators. Inspired by the use of preference data in RLHF, we formulate the evaluation as a ranking problem and introduce Pairwise-preference Search (PAIRS), an uncertainty-guided search-based rank aggregation method that employs LLMs to conduct pairwise comparisons locally and efficiently ranks candidate texts globally. PAIRS achieves state-of-the-art performance on representative evaluation tasks in long-form generations and demonstrates significant improvements over direct scoring. Furthermore, we provide insights into the role of pairwise preference in quantifying the transitivity of LLMs and demonstrate how PAIRS benefits from calibration using debiased pairwise evaluations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. AI Alignment at Your Discretion

    cs.AI 2025-02 conditional novelty 7.0 of 10

    The paper formalizes alignment discretion and shows empirically that annotators and models exercise substantial, often arbitrary, and mutually divergent discretion when applying alignment principles.

  2. Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models

    cs.AI 2025-06 unverdicted novelty 6.0 of 10

    LLMs show a quality-dependent position bias, favoring the first option for high-quality choices and later options for low-quality ones, and higher-temperature sampling can reveal the underlying preference.

  3. Deep Researcher with Test-Time Diffusion

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A test-time 'denoising' loop that repeatedly revises a draft report using fresh web retrieval, plus a component-wise self-evolution step, beats existing deep research agents on several benchmarks.

  4. Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases

    cs.CY 2025-06 conditional novelty 5.0 of 10

    In a physician-rated crowdsourced study, 76% of LLM responses to everyday health queries were valid, with GPT-4o highest (85%) and Llama3-8b lowest (50%); RAG did not consistently improve responses.

  5. (Towards) Scalable Reliable Automated Evaluation with Large Language Models

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.

  6. Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Meta-learned test-time adaptation, focused on the shakiest parts of a video, improves the stability and quality of full-frame neural video stabilizers.

  7. OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

    cs.CY 2025-05 conditional novelty 4.0 of 10

    The paper advocates protecting and leveraging OpenReview's peer review corpus as a community asset for LLM-based review assistance, benchmarks, and alignment.

Pith tools