Pith. sign in

REVIEW 10 cited by

Can LLM be a Personalized Judge?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11657 v1 pith:456TNFUF submitted 2024-06-17 cs.CL cs.CY

classification cs.CLcs.CY
keywords llm-as-a-personalized-judgeevaluationhumanuseragreementhigh-certaintyjudgellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensuring that large language models (LLMs) reflect diverse user values and preferences is crucial as their user bases expand globally. It is therefore encouraging to see the growing interest in LLM personalization within the research community. However, current works often rely on the LLM-as-a-Judge approach for evaluation without thoroughly examining its validity. In this paper, we investigate the reliability of LLM-as-a-Personalized-Judge, asking LLMs to judge user preferences based on personas. Our findings suggest that directly applying LLM-as-a-Personalized-Judge is less reliable than previously assumed, showing low and inconsistent agreement with human ground truth. The personas typically used are often overly simplistic, resulting in low predictive power. To address these issues, we introduce verbal uncertainty estimation into the LLM-as-a-Personalized-Judge pipeline, allowing the model to express low confidence on uncertain judgments. This adjustment leads to much higher agreement (above 80%) on high-certainty samples for binary tasks. Through human evaluation, we find that the LLM-as-a-Personalized-Judge achieves comparable performance to third-party humans evaluation and even surpasses human performance on high-certainty samples. Our work indicates that certainty-enhanced LLM-as-a-Personalized-Judge offers a promising direction for developing more reliable and scalable methods for evaluating LLM personalization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Informing AI Policy Assessment using Large-Scale Simulation of Interventions

    cs.CY 2026-04 conditional novelty 6.5 of 10

    A genetic algorithm optimizes weighted combinations of LLM-perceived harm mitigation, expert costs, and participatory scores over stakeholder-action pairs to surface viable AI policy packages for media harms.

  2. Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A dwell-time and breaking-news filtering framework plus a new benchmark improves personalized headline generation over prior models.

  3. SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    SynthesizeMe induces synthetic user personas from a few pairwise preferences and uses them to improve personalized LLM judging accuracy by a few points on a new benchmark.

  4. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  5. Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across five LLMs and ten no-consensus datasets, neutrality drops sharply when models act as pairwise judges, pointwise judges, or debaters compared to when they generate answers with an explicit neutral option.

  6. Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Dialog-act and maxim-aware prompting improves LLM judge accuracy on multi-turn preference data by up to 8 points, with further gains from jury-style voting.

  7. Can You Trick the Grader? Adversarial Persuasion of LLM Judges

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from lengt...

  8. Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models

    cs.SE 2025-08 conditional novelty 5.0 of 10

    A keyframe-plus-GPT-4o pipeline retrieves the single most representative frame for a reported gameplay bug, with F1@1 of 0.79 and Accuracy@1 of 0.89 on industrial bug-report videos.

  9. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  10. Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Language models show negligible persona-based differences on MMLU benchmarks but large, income-relevant differences when asked for salary negotiation advice.

Pith tools