Pith. sign in

REVIEW 3 cited by

PersonalLLM: Tailoring LLMs to Individual Preferences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20296 v2 pith:VIDVC5MT submitted 2024-09-30 cs.LG cs.CL

classification cs.LGcs.CL
keywords preferencesllmspersonalllmuserdatadatasetparticularusers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As LLMs become capable of complex tasks, there is growing potential for personalized interactions tailored to the subtle and idiosyncratic preferences of the user. We present a public benchmark, PersonalLLM, focusing on adapting LLMs to provide maximal benefits for a particular user. Departing from existing alignment benchmarks that implicitly assume uniform preferences, we curate open-ended prompts paired with many high-quality answers over which users would be expected to display heterogeneous latent preferences. Instead of persona-prompting LLMs based on high-level attributes (e.g., user's race or response length), which yields homogeneous preferences relative to humans, we develop a method that can simulate a large user base with diverse preferences from a set of pre-trained reward models. Our dataset and generated personalities offer an innovative testbed for developing personalization algorithms that grapple with continual data sparsity--few relevant feedback from the particular user--by leveraging historical data from other (similar) users. We explore basic in-context learning and meta-learning baselines to illustrate the utility of PersonalLLM and highlight the need for future methodological development. Our dataset is available at https://huggingface.co/datasets/namkoong-lab/PersonalLLM

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

    cs.LG 2026-07 reject novelty 6.0 of 10

    IRIS learns and iteratively refines natural-language user personas from implicit interaction streams; on 100 Reddit AITA commenters it predicts held-out verdicts at 61% accuracy, within a statistically untested 56-61% band.

  2. When Does Personality Composition Matter for Multi-Agent LLM Teams?

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Low agreeableness massively shifts multi-agent LLM communication yet barely hurts coding milestones, while the same prompt sharply degrades research milestones and collapses bargaining agreements.

  3. PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization

    cs.CL 2025-06 conditional novelty 6.0 of 10

    PersonaFeedback provides a human-labeled benchmark showing current LLMs, including strong reasoners, score only about 65-70 percent on hard personalization choices, and explicit persona information helps more than retrieval.

Pith tools