REVIEW 10 cited by
Can LLM be a Personalized Judge?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Ensuring that large language models (LLMs) reflect diverse user values and preferences is crucial as their user bases expand globally. It is therefore encouraging to see the growing interest in LLM personalization within the research community. However, current works often rely on the LLM-as-a-Judge approach for evaluation without thoroughly examining its validity. In this paper, we investigate the reliability of LLM-as-a-Personalized-Judge, asking LLMs to judge user preferences based on personas. Our findings suggest that directly applying LLM-as-a-Personalized-Judge is less reliable than previously assumed, showing low and inconsistent agreement with human ground truth. The personas typically used are often overly simplistic, resulting in low predictive power. To address these issues, we introduce verbal uncertainty estimation into the LLM-as-a-Personalized-Judge pipeline, allowing the model to express low confidence on uncertain judgments. This adjustment leads to much higher agreement (above 80%) on high-certainty samples for binary tasks. Through human evaluation, we find that the LLM-as-a-Personalized-Judge achieves comparable performance to third-party humans evaluation and even surpasses human performance on high-certainty samples. Our work indicates that certainty-enhanced LLM-as-a-Personalized-Judge offers a promising direction for developing more reliable and scalable methods for evaluating LLM personalization.
Forward citations
Cited by 10 Pith papers
-
Informing AI Policy Assessment using Large-Scale Simulation of Interventions
A genetic algorithm optimizes weighted combinations of LLM-perceived harm mitigation, expert costs, and participatory scores over stakeholder-action pairs to surface viable AI policy packages for media harms.
-
Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
A dwell-time and breaking-news filtering framework plus a new benchmark improves personalized headline generation over prior models.
-
SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
SynthesizeMe induces synthetic user personas from a few pairwise preferences and uses them to improve personalized LLM judging accuracy by a few points on a new benchmark.
-
Localizing Persona Representations in LLMs
Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.
-
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
Across five LLMs and ten no-consensus datasets, neutrality drops sharply when models act as pairwise judges, pointwise judges, or debaters compared to when they generate answers with an explicit neutral option.
-
Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
Dialog-act and maxim-aware prompting improves LLM judge accuracy on multi-turn preference data by up to 8 points, with further gains from jury-style voting.
-
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from lengt...
-
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
A keyframe-plus-GPT-4o pipeline retrieves the single most representative frame for a reported gameplay bug, with F1@1 of 0.79 and Accuracy@1 of 0.89 on industrial bug-report videos.
-
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.
-
Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
Language models show negligible persona-based differences on MMLU benchmarks but large, income-relevant differences when asked for salary negotiation advice.
Discussion (0). Continue with ORCID to comment.