Pith. sign in

REVIEW 3 major objections 5 minor 66 references

Personalized LLM for Generating Customized Responses to the Same Query from Different Users

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A single shared LLM can learn to answer the same query differently for different users.

desk verdict Good dataset and a sensible contrastive training recipe, but the architecture has no querier-identity input, so the same-query headline in Eq. 1 does not actually hold. read the letter →

arxiv 2412.11736 v2 pith:LZHJLSUF submitted 2024-12-16 cs.CL

classification cs.CL
keywords PersonalizationChatLLMsQuerier-AwareContrastivelearningDual-towerarchitectureLow-rankadaptationMQDialogdatasetResponsegeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish a new form of LLM personalization: instead of only varying the responder's role, the model should adapt to who is asking, so that identical queries receive different replies tailored to each querier's personality and relationship to the responder. To test this, the authors build a multi-querier dialogue dataset (MQDialog) from TV scripts and real WeChat records, and design a dual-tower architecture that separates a shared general encoder from a low-rank querier-specific encoder. A querier-contrastive loss, restricted to clusters of similar queries, pulls same-querier dialogues together and pushes different queriers apart. Reported results show relative gains of 8.4% to 48.7% in ROUGE-L and a 65.8% average GPT-4-judged winning rate over baselines, with case studies showing the same query answered differently for different queriers.

What carries the argument

The load-bearing mechanism is the dual-tower architecture: a general encoder initialized from a pretrained LLM captures the responder's cross-querier personality, while a separate specific encoder, whose feedforward layers are decomposed into two low-rank matrices, captures querier-specific personality; the two towers are fused by element-wise addition before the language-model head. Training is driven by a querier-contrastive loss, which maximizes a lower bound on mutual information between a dialogue representation and the querier's global representation, with multi-view augmentation (two projection views) and a query-similarity clustering step that confines contrastive pairs to dialogues with similar queries.

What would settle it

Collect a held-out set in which many different queriers ask literally the same question (e.g., 'What is gene sequencing?') to the same responder, then measure whether the model's responses vary systematically with querier identity, for example by computing the cosine distance between generated responses from different queriers versus responses from the same querier asked twice. If the between-querier distance is not significantly larger than the within-querier distance, the central claim is falsified.

Watch

Extended reading notes

Core claim

The central claim is that a unified, one-for-all model can internalize querier identity: even when two users ask exactly the same question, the model should produce responses whose distribution differs per querier, reflecting the querier's personality and relationship with the responder (formalized as P(y|x; Q_i; Θ) ≠ P(y|x; Q_j; Θ)). The paper argues this is achievable by decomposing dialogue personality into a cross-querier general component (full transformer) and a sparse, low-rank querier-specific component, trained jointly with language modeling and a querier-contrastive loss. Because identical queries across queriers are rare in real data, the method clusters dialogues by query-embedding similarity and performs contrastive learning within clusters. The paper further contributes MQDialog, a 173-querier, 12-responder benchmark built from English and Chinese scripts plus real WeChat records, and reports consistent improvements in BLEU/ROUGE and LLM-judged winning rates against zero-shot, fine-tuning, profile-based, and few-shot baselines.

Load-bearing premise

The central same-query claim is only directly tested in a few case studies; the quantitative evaluation relaxes 'same query' to 'similar queries' clustered by embedding similarity, so if similar-query similarity is not a good stand-in for identical queries, the reported numbers do not directly support the paper's headline claim.

Editorial extensions

If this is right

  • A single shared model can serve many users without per-user fine-tuning, since both encoders are shared and the added parameters are about 1% of the pretrained LLM.
  • Including the querier side, not just the responder role, improves response quality over role-profile, few-shot, and fine-tuning baselines across English and Chinese.
  • Query-similarity clustering is necessary: contrastive learning without it hurts performance, and removing the contrastive loss also degrades BLEU/ROUGE, showing both components matter.
  • The model's representations of dialogues become separated by querier in t-SNE plots, whereas fine-tuning mixes them.
  • On same-query case studies, the model responds differently to different queriers while fine-tuning produces near-identical replies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the central claim holds, then building paired same-query benchmarks (identical queries posed by many users) would likely show larger measured gains than the similar-query evaluation in the paper, because the relaxed clustering setup understates the contrastive signal.
  • The approach suggests a practical cold-start compromise: for a new querier with no history, the shared towers still produce generic responses, and personalization improves as dialogues accumulate; profile-clustering could scale the method to million-scale user bases, as the paper notes.
  • The querier-contrastive objective can be seen as a form of user-embedding learning; it might transfer to other personalized generation tasks such as recommendation explanations or customer support, where the responder is fixed and the querier varies.
  • Because the dataset is built from scripted dialogues and one author's WeChat records, results may depend on how cleanly querier identity is expressed in scripted versus real conversations; a test on naturally occurring multi-querier customer-service logs would be a strong external check.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies querier-aware LLM personalization, where the model should produce different responses to the same query depending on who asks it. The authors propose a dual-tower architecture consisting of a cross-querier general encoder and a shared low-rank specific encoder, trained with a querier-contrastive loss, multi-view augmentation, and query-similarity clustering. They construct MQDialog, a multi-querier dialogue dataset from English and Chinese scripts and WeChat records, and report BLEU/ROUGE gains over zero-shot, fine-tuned, profile-based, few-shot, and querier-characteristic baselines, plus GPT-4 and human win rates. The paper also provides ablations, t-SNE visualizations, and same-query case studies.

Significance. The task is well-motivated, and the empirical package is extensive: four baseline families, ablations, a new dataset, GPT-4 win rates, a small human evaluation, and public code. If the central claim were established, the proposed parameter-efficient design would be a useful contribution to personalized dialogue. However, the current manuscript does not establish the central claim: the architecture has no querier-identity input at inference, and the quantitative evaluation does not test the identical-query condition of Eq. (1). The case studies are suggestive but not a substitute for a controlled paired-query experiment. These issues are substantive enough that the revision should be major.

major comments (3)
  1. [Section 4.2-4.3, Eq. (1)] The architecture cannot satisfy Eq. (1) as stated. Both the general encoder G_g and the specific encoder G_s take only the dialogue text x as input and are shared across all queriers; the per-querier global representation e_i is used only in the contrastive losses (Eqs. (5) and (8)) and is never injected into the fused representation G_g(x)+G_s(x) that feeds the language-modeling head. Consequently, for any token-identical inputs x_i^k = x_j^l, the encoders produce identical representations and identical output distributions, so P(y_i^k | x_i^k; Q_i; Θ) = P(y_j^l | x_j^l; Q_j; Θ), contradicting Eq. (1). Stochastic sampling from the same distribution would generate different strings but not different distributions. To make the claim meaningful, the model must take an explicit querier-identity signal (e.g., e_i or a learned querier embedding) as input at inference, and the paper must specify how that signal is provided.
  2. [Section 4.1, Section 6.2, Figure 8] The quantitative evaluation does not test the same-query condition. Section 4.1 explicitly relaxes x_i^k = x_j^l to similar queries x_i^k ≈ x_j^l, and Section 6.2 measures BLEU and ROUGE on ordinary held-out dialogues, where each querier's context is generally different. The only same-query demonstrations are the case studies in Figure 8 and Appendix C, which are anecdotal and do not report whether the inputs were strictly identical, whether speaker names are part of x, or the decoding parameters used. A paired evaluation is needed: for the same surface-form query and the same context history, with only the querier identity varied, the responses or their distributions should be compared, and the input format should be stated precisely.
  3. [Section 4.3, Lemma 1 and Theorem 1] The proof of Theorem 1 does not satisfy the conditions of Lemma 1. The lemma requires X to be a set of n features of random samples, with exactly one sample drawn from the conditional distribution p(x_i|y) and the rest from the marginal p(x); here E consists of one running-average global representation per querier, which is not a fresh random sample from any conditional distribution given z_i^k, and the score function is a deterministic function of the model's own parameters. The mutual-information bound is imported from prior work (Ref. [48]) and cannot be invoked without verifying these sampling assumptions. If the bound is intended only as intuition, that should be stated explicitly; if it is a formal claim, the proof needs to be repaired.
minor comments (5)
  1. [Section 4.3, Eq. (4)] Equation (4) and the surrounding text: the quantity f_QC is an increasing function of cosine similarity, so the phrase 'proportional to the cosine distance' is misleading; please say 'cosine similarity' or adjust the expression accordingly.
  2. [Appendix D.1] The text says 'dual-town structure'; this appears to be a typo for 'dual-tower structure'.
  3. [Table 3] The '-' entries for the Real Person row in the RPG and QCG columns are not explained; please state explicitly why those baselines were not run for that responder.
  4. [Section 6.3] The human evaluation reports win rates but does not report inter-annotator agreement; please include a statistic such as Cohen's kappa or a comparable measure.
  5. [Sections 6.2 and 6.3] The decoding method (greedy, top-p, temperature, etc.) used to generate responses is never specified; this matters for interpreting the same-query case studies and win-rate comparisons.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; empirical gains are self-contained, though same-query personalization is not directly testable from the architecture.

full rationale

Score 0 (no circularity). The paper's quantitative claims (BLEU, ROUGE, GPT-4 winning rates) are empirical comparisons on held-out test dialogues against external baselines. No fitted parameter is re-labeled as a prediction, and Theorem 1 is a standard InfoNCE bound imported from Oord et al. (reference [48]), an external work. The querier-specific global representation e_i is a running average used only inside the contrastive loss (Eqs. 5 and 8) and is not fed to the language-model head, so no quantity is predicted from its own fitted value. There is no self-citation chain: the present authors do not appear in the reference list. The main weakness is a non-circular validity gap: Eq. (1)'s identical-query premise is relaxed to similar queries in Section 4.1, and the encoders G_g and G_s take only dialogue text x, so for exactly identical x the model yields identical distributions; the same-query case studies (Figure 8, Appendix C) may therefore reflect sampling rather than querier conditioning. This is a correctness/scope concern, not a circular derivation, so it does not raise the circularity score.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

No new physical or conceptual entities are postulated. The querier-specific encoder and the global representation e_i are learned model components, not independent entities with external falsifiable handles.

free parameters (6)
  • temperature tau = not reported
    Temperature in the querier-contrastive loss (Eq. 4) controls the sharpness of the softmax over queriers; chosen by hand, and the value is not given in the paper.
  • number of dialogue clusters K = 10
    Number of k-means clusters over query embeddings; set to 10 in the main experiments and studied via ablation in Figure 5. The cluster-restricted contrastive design depends on this choice.
  • LoRA rank = 16
    Rank of LoRA adaptation in the general block and of the low-rank matrices in the specific block; controls the capacity of the querier-specific encoder (Section 6.1).
  • WeChat dialogue splitting interval = 3 hours
    A new dialogue begins when a message arrives more than 3 hours after the previous one; this empirically chosen threshold shapes the WeChat portion of MQDialog (Section 5.3).
  • minimum dialogues per querier = 20
    Queriers with fewer than 20 dialogues with a responder are filtered out; this threshold determines the querier set of MQDialog (Section 5.2).
  • training hyperparameters = lr 2e-4 to 1e-4, batch 4, 20 epochs, max tokens 592, FP16
    Standard fine-tuning settings chosen for the experiments; they affect all models equally, so they are not central to the claim but are free choices.
assumptions (5)
  • standard math The InfoNCE lower bound (Lemma 1 from van den Oord et al. [48]) applies to the querier-contrastive loss.
    Theorem 1 relies on the categorical cross-entropy form of InfoNCE to conclude that L_QC maximizes a lower bound on mutual information between the dialogue representation and the querier global representation.
  • ad hoc to paper The global representation e_i can be treated as the positive sample drawn from the conditional distribution p(x|y) in Lemma 1.
    In practice e_i is a running average of the model's own representations, not a random draw from a conditional distribution, so the proof sketch in Section 4.3 assumes a property the training procedure does not literally provide.
  • domain assumption Querier-specific personality is sparse enough to be captured by low-rank matrices.
    Section 4.2 justifies the low-rank feedforward design by an intuition about sparsity; no formal or empirical bound is given beyond the overall results.
  • domain assumption Similar queries can stand in for identical queries in contrastive learning and evaluation.
    Section 4.1 relaxes the exact equality x_i^k = x_j^l to semantic similarity measured by OpenAI text-embedding-ada-002 with k-means; the evaluation then also uses similar rather than identical queries.
  • domain assumption Ground-truth scripted and chat responses are the correct personalization target and capture querier-specific response behavior.
    The dataset construction assumes that the responder's lines in TV scripts and the author's WeChat replies are the desired personalized outputs, which is the basis of all supervised metrics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Personalized LLM for Generating Customized Responses to the Same Query from Different Users." pith.science (2026). https://pith.science/paper/LZHJLSUF

@misc{pith2026241211736,
  author       = {Pith},
  title        = {Pith review of: Personalized LLM for Generating Customized Responses to the Same Query from Different Users},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LZHJLSUF}},
  note         = {Machine review of arXiv:2412.11736}
}
read the original abstract

Existing work on large language model (LLM) personalization assigned different responding roles to LLMs, but overlooked the diversity of queriers. In this work, we propose a new form of querier-aware LLM personalization, generating different responses even for the same query from different queriers. We design a dual-tower model architecture with a cross-querier general encoder and a querier-specific encoder. We further apply contrastive learning with multi-view augmentation, pulling close the dialogue representations of the same querier, while pulling apart those of different queriers. To mitigate the impact of query diversity on querier-contrastive learning, we cluster the dialogues based on query similarity and restrict the scope of contrastive learning within each cluster. To address the lack of datasets designed for querier-aware personalization, we also build a multi-querier dataset from English and Chinese scripts, as well as WeChat records, called MQDialog, containing 173 queriers and 12 responders. Extensive evaluations demonstrate that our design significantly improves the quality of personalized response generation, achieving relative improvement of 8.4% to 48.7% in ROUGE-L scores and winning rates ranging from 54% to 82% compared with various baseline methods.

Figures

Figures reproduced from arXiv: 2412.11736 by the authors.

Figure 1
Figure 1. LLM personalization from responder-side diversity [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Dual-tower architecture. The feedfroward layer of [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Querier-contrastive loss in each cluster of dialogues [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: The winning rate judged by GPT-4. model and the baselines are then manually compared for quality and for which is more like what the responder would naturally say to the specific querier by annotators. The annotators are three volunteers in the lab. To maintain objecti…
Figure 5
Figure 5. Figure 5: Results of experiments with different numbers of [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Chinese dialogue representations of different [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Case study with Sheldon from The Big Bang Theory and Tong Xiangyu from My Own Swordsman. [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Prompt template for evaluating the winning rate for the English datasets. [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Prompt template for evaluating the winning rate for the Chinese datasets. [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Case study with Sheldon from The Big Bang Theory. [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Case study with Tong Xiangyu from My Own Swordsman. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

66 extracted references · 18 canonical work pages

  1. [48]

    Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. CoRR abs/1807.03748 (2018). arXiv:1807.03748 http://arxiv.org/abs/1807.03748

  2. [1]

    AI@Meta. 2024. Llama 3 Model Card. (2024). https://github.com/meta-llama/ llama3/blob/main/MODEL_CARD.md

  3. [2]

    Bilmes, and Karen Livescu

    Galen Andrew, Raman Arora, Jeff A. Bilmes, and Karen Livescu. 2013. Deep Canonical Correlation Analysis. In Proceedings of the 30th International Con- ference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013 (JMLR Workshop and Conference Proceedings, Vol. 28) . JMLR.org, 1247–1255. http://proceedings.mlr.press/v28/andrew13.html

  4. [3]

    Devon Hjelm, and William Buchwalter

    Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019. Learning Repre- sentations by Maximizing Mutual Information Across Views. InAdvances in Neu- ral Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina...

  5. [4]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...

  6. [5]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

  7. [6]

    Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, Aili Chen, Nianqi Li, Lida Chen, Caiyu Hu, Siye Wu, Scott Ren, Ziquan Fu, and Yanghua Xiao. 2024. From Persona to Per- sonalization: A Survey on Role-Playing Language Agents. CoRR abs/2404.18231 (2024). doi:10.48550/ARXIV.2404.18231 arXiv:2404.18231

  8. [7]

    Nuo Chen, Yan Wang, Haiyun Jiang, Deng Cai, Yuhan Li, Ziyang Chen, Longyue Wang, and Jia Li. 2023. Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with Characters. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Lingui...

Show all 66 references
  1. [8]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of...

  2. [9]

    Yi-Pei Chen, Noriki Nishida, Hideki Nakayama, and Yuji Matsumoto. 2024. Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodolo- gies, and Evaluations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language R...

  3. [10]

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023. A Survey for In-context Learning. CoRR abs/2301.00234 (2023). doi:10.48550/ARXIV.2301.00234 arXiv:2301.00234

  4. [11]

    Marco Federici, Anjan Dutta, Patrick Forré, Nate Kushman, and Zeynep Akata

  5. [12]

    Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomás Mikolov

    Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomás Mikolov. 2013. DeViSE: A Deep Visual- Semantic Embedding Model. In Advances in Neural Information Processing Sys- tems 26: 27th Annual Conference on Neural Information...

  6. [13]

    Tingchen Fu, Xueliang Zhao, Chongyang Tao, Ji-Rong Wen, and Rui Yan

  7. [14]

    Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, and Baoyuan Wang. 2023. LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Con- structed from Live Streaming. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1:...

  8. [15]

    Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks.CoRR abs/2303.15056 (2023). doi:10. 48550/ARXIV.2303.15056 arXiv:2303.15056 Querier-Aware LLM: Generating Personalized Responses to the Same Query from Differen...

  9. [16]

    Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Philip Bachman, Adam Trischler, and Yoshua Bengio

    R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Philip Bachman, Adam Trischler, and Yoshua Bengio. 2019. Learning deep representa- tions by mutual information estimation and maximization. In 7th International Conference on Learning Representations, ICLR 2...

  10. [17]

    Harold Hotelling. 1992. Relations between two sets of variates. In Breakthroughs in statistics: methodology and distribution . Springer, 162–190

  11. [18]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 20...

  12. [19]

    Qiushi Huang, Shuai Fu, Xubo Liu, Wenwu Wang, Tom Ko, Yu Zhang, and Lil- ian Tang. 2023. Learning Retrieval Augmentation for Personalized Dialogue Generation. In Proceedings of the 2023 Conference on Empirical Methods in Nat- ural Language Processing, EMNLP 2023, Singapore, De...

  13. [20]

    Lilian Tang

    Qiushi Huang, Yu Zhang, Tom Ko, Xubo Liu, Bo Wu, Wenwu Wang, and H. Lilian Tang. 2023. Personalized Dialogue Generation with Persona-Adaptive Attention. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty- Fifth Conference on Innovative Applications...

  14. [21]

    Deuk Sin Kwon, Sunwoo Lee, Ki Hyun Kim, Seojin Lee, Taeyoon Kim, and Eric Davis. 2023. What, When, and How to Ground: Designing User Persona-Aware Conversational Agents for Engaging Dialogue. In Proceedings of the The 61st Annual Meeting of the Association for Computational Li...

  15. [22]

    Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi Mi, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, Linkang Zhan, Yaokai Jia, Pingyu Wu, and Haozhen Sun. 2023. ChatHaruhi: Reviving Anime Character in Reality via Large Language Model. CoRR abs/2308.09597 (202...

  16. [23]

    Hao Li, Chenghao Yang, An Zhang, Yang Deng, Xiang Wang, and Tat-Seng Chua

  17. [24]

    Spithourakis, Jianfeng Gao, and William B

    Jiwei Li, Michel Galley, Chris Brockett, Georgios P. Spithourakis, Jianfeng Gao, and William B. Dolan. 2016. A Persona-Based Neural Conversation Model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, August 7-12, 2016, Berl...

  18. [25]

    Selvaraju, Akhilesh Gotmare, Shafiq R

    Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation. In Advances in Neural Information Processing Systems 34: Annual Confe...

  19. [26]

    Yingming Li, Ming Yang, and Zhongfei Zhang. 2019. A Survey of Multi-View Representation Learning. IEEE Trans. Knowl. Data Eng. 31, 10 (2019), 1863–1883. doi:10.1109/TKDE.2018.2872063

  20. [27]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013

  21. [28]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, ...

  22. [29]

    Zhengyi Ma, Zhicheng Dou, Yutao Zhu, Hanxun Zhong, and Ji-Rong Wen. 2021. One Chatbot Per Person: Creating Personalized Chatbots based on Implicit User Profiles. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Vir...

  23. [30]

    Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few- Shot Bot: Prompt-Based Learning for Dialogue Systems. CoRR abs/2110.08118 (2021). arXiv:2110.08118 https://arxiv.org/abs/2110.08118

  24. [31]

    Pierre-Emmanuel Mazaré, Samuel Humeau, Martin Raison, and Antoine Bordes

  25. [32]

    Seyed Mahed Mousavi, Simone Caldarella, and Giuseppe Riccardi. 2023. Response Generation in Longitudinal Dialogues: Which Knowledge Representation Helps?. In Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023), Yun-Nung Chen and Abhinav Rastogi (Eds....

  26. [33]

    Atsushi Otsuka, Kazuya Matsuo, Ryo Ishii, Narichika Nomoto, and Hiroaki Sugiyama. 2024. User-Specific Dialogue Generation with User Profile-Aware Pre-Training Model and Parameter-Efficient Fine-Tuning. CoRR abs/2409.00887 (2024). doi:10.48550/ARXIV.2409.00887 arXiv:2409.00887

  27. [34]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA. ACL, 311–318. ...

  28. [35]

    Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

    Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep Contextualized Word Rep- resentations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguisti...

  29. [36]

    Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021. AdapterFusion: Non-Destructive Task Composition for Trans- fer Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: M...

  30. [37]

    Xubin Ren, Wei Wei, Lianghao Xia, Lixin Su, Suqi Cheng, Junfeng Wang, Dawei Yin, and Chao Huang. 2024. Representation Learning with Large Language Models for Recommendation. In Proceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, May 13-17, 2024 , Tat-Seng Chua,...

  31. [38]

    Nafis Sadeq, Zhouhang Xie, Byungkyu Kang, Prarit Lamba, Xiang Gao, and Julian J. McAuley. 2024. Mitigating Hallucination in Fictional Character Role-Play. CoRR abs/2406.17260 (2024). doi:10.48550/ARXIV.2406.17260 arXiv:2406.17260

  32. [39]

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024. LaMP: When Large Language Models Meet Personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, A...

  33. [40]

    Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nat. 623, 7987 (2023), 493–498. doi:10.1038/S41586-023-06647-8

  34. [41]

    Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-LLM: A Trainable Agent for Role-Playing. In Proceedings of the 2023 Conference on Empir- ical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and K...

  35. [42]

    ShareGPT. 2023. ShareGPT Dataset. https://huggingface.co/datasets/shareAI/ ShareGPT-Chinese-English-90k

  36. [43]

    Tianhao Shen, Sun Li, and Deyi Xiong. 2023. RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models. CoRR abs/2312.16132 (2023). doi:10. 48550/ARXIV.2312.16132 arXiv:2312.16132

  37. [44]

    Haoyu Song, Weinan Zhang, Yiming Cui, Dong Wang, and Ting Liu. 2019. Exploit- ing Persona Information for Diverse Generation of Conversational Responses. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intel- ligence, IJCAI 2019, Macao, China, ...

  38. [45]

    Zhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu, Bing Yin, and Meng Jiang. 2024. Democratizing Large Language Models via Personalized Parameter- Efficient Fine-tuning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Mi...

  39. [46]

    Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive Multiview Coding. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XI (Lecture Notes in Computer Science, Vol. 12356), Andrea Vedaldi, Horst Bischof...

  40. [47]

    Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan. 2024. CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: L...

  41. [49]

    Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, 11 (2008)

  42. [50]

    Hongru Wang, Minda Hu, Yang Deng, Rui Wang, Fei Mi, Weichao Wang, Yasheng Wang, Wai-Chung Kwan, Irwin King, and Kam-Fai Wong. 2023. Large Language Models as Source Planner for Personalized Knowledge-grounded Dialogues. In Findings of the Association for Computational Linguisti...

  43. [51]

    Noah Wang, Z. y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. 2024. RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playi...

  44. [52]

    Weiran Wang, Raman Arora, Karen Livescu, and Jeff A. Bilmes. 2015. On Deep Multi-View Representation Learning. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 (JMLR Workshop and Conference Proceedings, Vol. 37) ...

  45. [53]

    Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao. 2024. InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews. In Proceedin...

  46. [54]

    Zijian Wang, Yadan Luo, Ruihong Qiu, Zi Huang, and Mahsa Baktashmotlagh

  47. [55]

    Weinberger, John Blitzer, and Lawrence K

    Kilian Q. Weinberger, John Blitzer, and Lawrence K. Saul. 2005. Dis- tance Metric Learning for Large Margin Nearest Neighbor Classification. In Advances in Neural Information Processing Systems 18 [Neural Informa- tion Processing Systems, NIPS 2005, December 5-8, 2005, Vancouv...

  48. [56]

    Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, Hui Xiong, and Enhong Chen

  49. [57]

    Yuwei Wu, Xuezhe Ma, and Diyi Yang. 2021. Personalized Response Generation via Generative Split Memory Network. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, On...

  50. [58]

    Changqing Zhang, Yajie Cui, Zongbo Han, Joey Tianyi Zhou, Huazhu Fu, and Qinghua Hu. 2020. Deep partial multi-view learning. IEEE transactions on pattern analysis and machine intelligence 44, 5 (2020), 2402–2415

  51. [59]

    Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I have a dog, do you have pets too?. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Aus...

  52. [60]

    A" or "B

    Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Libiao Peng, Jiaming Yang, Xiyao Xiao, Sahand Sabour, Xiaohan Zhang, Wenjing Hou, Yijia Zhang, Yuxiao Dong, Jie Tang, and Minlie Huang. 2023. CharacterGLM: Customizing Chinese Conversational AI...

  53. [62]

    World Wide Web (WWW) 27, 5 (2024), 60

    A survey on large language models for recommendation. World Wide Web (WWW) 27, 5 (2024), 60. doi:10.1007/S11280-024-01291-2

  54. [2018]

    Training Millions of Personalized Dialogue Agents. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018 , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Associatio...

  55. [2020]

    In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020

    Learning Robust Representations via Multi-View Information Bottleneck. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net. https://openreview.net/ forum?id=B1xwcyHFDr

  56. [2021]

    In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021

    Learning to Diversify for Single Domain Generalization. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021. IEEE, 814–823. doi:10.1109/ICCV48922.2021.00087

  57. [2022]

    There Are a Thousand Hamlets in a Thousand People’s Eyes: Enhancing Knowledge-grounded Dialogue with Personal Memory. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 ...

  58. [2024]

    CoRR abs/2406.05925 (2024)

    Hello Again! LLM-powered Personalized Agent for Long-term Dialogue. CoRR abs/2406.05925 (2024). doi:10.48550/ARXIV.2406.05925 arXiv:2406.05925

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.