Pith. sign in

REVIEW 10 cited by

LongLaMP: A Benchmark for Personalized Long-form Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11016 v3 pith:5KK55DAI submitted 2024-06-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords generationlong-textpersonalizedbenchmarklonglamptasksapplicationsimportance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Long-text generation is seemingly ubiquitous in real-world applications of large language models such as generating an email or writing a review. Despite the fundamental importance and prevalence of long-text generation in many practical applications, existing work on personalized generation has focused on the generation of very short text. To overcome these limitations, we study the problem of personalized long-text generation, that is, generating long-text that is personalized for a specific user while being practically useful for the vast majority of real-world applications that naturally require the generation of longer text. In this work, we demonstrate the importance of user-specific personalization for long-text generation tasks and develop the Long-text Language Model Personalization (LongLaMP) Benchmark. LongLaMP provides a comprehensive and diverse evaluation framework for personalized long-text generation. Extensive experiments on LongLaMP for zero-shot and fine-tuned language tasks demonstrate the effectiveness of the proposed benchmark and its utility for developing and evaluating techniques for personalized long-text generation across a wide variety of long-text generation tasks. The results highlight the importance of personalization across a wide variety of long-text generation tasks. Finally, we release the benchmark for others to use for this important problem.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ClawRec: A Claw-Native Recommender System

    cs.IR 2026-07 conditional novelty 6.5 of 10

    ClawRec turns cross-platform behavior into a temporally managed user state and role-aware complementary slates, beating agentic baselines on a new synthetic life-event benchmark.

  2. Know It, Act on It: Investigating Memory Utilization in LLM Personalization

    cs.CL 2026-07 conditional novelty 6.0 of 10

    LLM agents often pass a direct recall question about a user's preference yet fail to act on the same preference in a realistic request — a Know–Act gap that persists even in the best systems and is widest, on average,...

  3. Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Persona2Web is a new open-web benchmark where agents must infer a user's preferences from synthetic browsing history to solve intentionally ambiguous queries; current best agents score 13% success.

  4. Evaluating Style-Personalized Text Generation: Challenges and Directions

    cs.CL 2025-08 reject novelty 6.0 of 10

    A new style-discrimination benchmark for personalized text generation shows ensemble metrics give only a marginal, possibly test-fitted, edge over the best single judge.

  5. PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    PREF is a reference-free, two-stage LLM judge that personalizes a quality rubric with a user profile and scores candidates against it, beating reminder-only baselines on the PrefEval implicit preference subset.

  6. From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    ProxyReward trains long-form generation models by rewarding how well an AI judge can answer generated yes/no questions about the response, improving open-source models on ProxyQA.

  7. CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs

    cs.CL 2026-01 conditional novelty 5.0 of 10

    CURP represents users as sparse combinations of discrete prototype codebook embeddings and uses them as frozen-LLM prefixes, outperforming personalization baselines on four text-generation tasks with about 20M trainab...

  8. PrefReward: Learning User Preference Matrix for Personalized Text Generation

    cs.CL 2026-07 conditional novelty 4.0 of 10

    PrefReward selects the most style-aligned LLM output via a KL-divergence reward against an explicit user preference matrix, beating retrieval baselines on LongLaMP.

  9. Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs

    cs.AI 2026-01 conditional novelty 4.0 of 10

    PsPLUG, a soft-prompt plug-in trained with style-conditioned preference pairs, preserves user identity under explicit style instructions and lets users tune personalization strength via an α scalar.

  10. Position: It's Time to Act on the Risk of Efficient Personalized Text Generation

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Fine-tuned open LLMs can imitate individual writing styles from small samples, evade detection tools, and are not yet addressed by current safeguards or law.

Pith tools