Pith. sign in

REVIEW 3 cited by

Human-Centered Design Recommendations for LLM-as-a-Judge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03479 v1 pith:TZOR2QRO submitted 2024-07-03 cs.HC

classification cs.HC
keywords humandesignllm-as-a-judgellmscriteriaeffectiveevaluationoutputs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional reference-based metrics, such as BLEU and ROUGE, are less effective for assessing outputs from Large Language Models (LLMs) that produce highly creative or superior-quality text, or in situations where reference outputs are unavailable. While human evaluation remains an option, it is costly and difficult to scale. Recent work using LLMs as evaluators (LLM-as-a-judge) is promising, but trust and reliability remain a significant concern. Integrating human input is crucial to ensure criteria used to evaluate are aligned with the human's intent, and evaluations are robust and consistent. This paper presents a user study of a design exploration called EvaluLLM, that enables users to leverage LLMs as customizable judges, promoting human involvement to balance trust and cost-saving potential with caution. Through interviews with eight domain experts, we identified the need for assistance in developing effective evaluation criteria aligning the LLM-as-a-judge with practitioners' preferences and expectations. We offer findings and design recommendations to optimize human-assisted LLM-as-judge systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks

    cs.HC 2025-07 conditional novelty 6.0 of 10

    Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...

  2. Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across five LLMs and ten no-consensus datasets, neutrality drops sharply when models act as pairwise judges, pointwise judges, or debaters compared to when they generate answers with an explicit neutral option.

  3. Engineering AI Judge Systems

    cs.SE 2024-11 conditional novelty 5.0 of 10

    A constitution-based, search-driven development framework for AI judge systems improves judged accuracy by up to 6.2% on commit message generation, with about 58% of general principles reused across five languages.

Pith tools