Pith. sign in

REVIEW 2 cited by

Optimization Methods for Personalizing Large Language Models through Retrieval Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05970 v1 pith:WLJJGIHV submitted 2024-04-09 cs.CL cs.IR

classification cs.CLcs.IR
keywords languagemodelsretrievalgenerationlargemodeloptimizationpersonalized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies retrieval-augmented approaches for personalizing large language models (LLMs), which potentially have a substantial impact on various applications and domains. We propose the first attempt to optimize the retrieval models that deliver a limited number of personal documents to large language models for the purpose of personalized generation. We develop two optimization algorithms that solicit feedback from the downstream personalized generation tasks for retrieval optimization -- one based on reinforcement learning whose reward function is defined using any arbitrary metric for personalized generation and another based on knowledge distillation from the downstream LLM to the retrieval model. This paper also introduces a pre- and post-generation retriever selection model that decides what retriever to choose for each LLM input. Extensive experiments on diverse tasks from the language model personalization (LaMP) benchmark reveal statistically significant improvements in six out of seven datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

    cs.CL 2026-08 conditional novelty 6.0 of 10

    AlignXada uses verbal reinforcement learning to learn reusable text-rewriting policies that compress universal user preference profiles into task-specific ones, improving downstream personalization accuracy on most te...

  2. Personalized Graph-Based Retrieval for Large Language Models

    cs.CL 2025-01 reject novelty 4.0 of 10

    PGraphRAG adds neighbor-user reviews to LLM prompts and claims improved personalized generation, but its own ablations show the user's history contributes little beyond item context.

Pith tools