REVIEW 19 cited by
Personalization of Large Language Models: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Personalization of Large Language Models (LLMs) has recently become increasingly important with a wide range of applications. Despite the importance and recent progress, most existing works on personalized LLMs have focused either entirely on (a) personalized text generation or (b) leveraging LLMs for personalization-related downstream applications, such as recommendation systems. In this work, we bridge the gap between these two separate main directions for the first time by introducing a taxonomy for personalized LLM usage and summarizing the key differences and challenges. We provide a formalization of the foundations of personalized LLMs that consolidates and expands notions of personalization of LLMs, defining and discussing novel facets of personalization, usage, and desiderata of personalized LLMs. We then unify the literature across these diverse fields and usage scenarios by proposing systematic taxonomies for the granularity of personalization, personalization techniques, datasets, evaluation methods, and applications of personalized LLMs. Finally, we highlight challenges and important open problems that remain to be addressed. By unifying and surveying recent research using the proposed taxonomies, we aim to provide a clear guide to the existing literature and different facets of personalization in LLMs, empowering both researchers and practitioners.
Forward citations
Cited by 19 Pith papers
-
Tailored untruths: How personalisation challenges LLM safeguards
A 1.6-million-text study of eight LLMs in four languages finds that adding demographic personae to disinformation prompts raises jailbreak rates from 78% to 82%.
-
PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing
A training-free four-stage LLM agent that routes papers into personal folksonomy folders by inspecting member papers and metadata, lifting Recall@1 from 0.39 to 0.61 on real libraries.
-
Auditing Alignment Controllability in LLMs via Political Axes
On a 63,700-response Political Compass stress test of seven frontier LLMs, system-prompt framing dominates model identity, and steerability needs dispersion, symmetry, saturation, and refusal-floor metrics.
-
Personalized Image Aesthetic Assessment via Preference-rich Sample Mining and Cohort Merging
PRAC mines preference-rich images and merges LoRA adapters from aesthetically similar users to achieve state-of-the-art personalized aesthetic rating prediction.
-
FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
A tile-based memory layout for mobile GPUs that unifies forward and backward data access, eliminating most transpose/reshape overhead and speeding LLM fine-tuning 2.2–5.7× in the paper's measurements.
-
Synthetic Interaction Data for Scalable Personalization in Large Language Models
PersonaGym simulates noisy multi-turn user–assistant interactions to build PersonaAtlas, and PPOpt learns to rewrite user prompts from interaction history, improving judged personalization on synthetic benchmarks.
-
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
A broad measurement study shows that false refusals in LLMs depend more on model and task choice than on sociodemographic personas, with newer models refusing far less.
-
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
PREF is a reference-free, two-stage LLM judge that personalizes a quality rubric with a user profile and scores candidates against it, beating reminder-only baselines on the PrefEval implicit preference subset.
-
Impact of Rankings and Personalized Recommendations in Marketplaces
In a stylized marketplace model, public rankings provide zero welfare gains under capacity constraints, while personalized recommendations yield welfare gains that grow with preference heterogeneity and heavy-tailed i...
-
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.
-
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
PERSONACONVBENCH is a new Reddit-based benchmark showing that LLMs predict sentiment, community scores, and next replies better when given a user's multi-turn conversation history, and it releases public data and code.
-
High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making
Individual investing exposes four structural limits of LLM personalization—evolving contradictory memory, long-horizon thesis drift, style-vs-signal conflict, and no ground-truth labels—requiring new architectures bey...
-
Effects of Personality- and Opinion-Alignment in Human-AI Interaction
People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.
-
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
The authors introduce an entropy-based framework that uses user behavior prediction as a measure of LLM generalization, and find GPT-4o outperforms GPT-4o-mini and Llama-3.1 on movie and music recommendation tasks.
-
Reasoning Meets Personalization: Unleashing the Potential of Large Reasoning Model for Personalized Generation
Large reasoning models underperform general chat models on personalization tasks, but a structured template plus self-checking and self-referencing restores and improves performance.
-
PrefReward: Learning User Preference Matrix for Personalized Text Generation
PrefReward selects the most style-aligned LLM output via a KL-divergence reward against an explicit user preference matrix, beating retrieval baselines on LongLaMP.
-
Matching Game Preferences Through Dialogical Large Language Models: A Perspective
This perspective paper proposes the D-LLM framework, which couples the authors' GRAPHYP knowledge graphs with LLMs to personalize AI responses and make reasoning traceable, but no empirical validation is presented.
-
Personalized Image Generation from an Author Writing Style
LLM-generated text-to-image prompts derived from author style sheets produce images that ten raters judged as moderately faithful (4.08/5), but the evaluation has no control condition and the dataset link is a placeholder.
-
PersonaBOT: Bringing Customer Personas to Life with LLMs and RAG
A RAG chatbot augmented with synthetic personas generated from customer success stories raised its average accuracy rating from 5.88 to 6.42 at Volvo CE, with few-shot prompting producing more complete personas than c...
Discussion (0). Sign in to comment.