Pith. sign in

REVIEW 5 major objections 6 minor 86 references

Adaptive Personalized Conversational Information Retrieval

T0 review · 5 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Conversational search improves when an LLM first labels each query turn as non-personalized, partially personalized, or personalized, and the system fuses personalized and neutral retrieval lists with weights tuned per level.

desk verdict A useful adaptive-personalization pipeline with consistent iKAT gains, but the paper never shows the per-level fusion weights beat one globally tuned triple, so the core mechanism is under-supported. read the letter →

arxiv 2508.08634 v1 pith:QBC2T27F submitted 2025-08-12 cs.IR cs.CL

classification cs.IRcs.CL
keywords conversationalinformationretrievaladaptivepersonalizationlevelidentificationqueryreformulationrankingfusionTRECiKATlargelanguagemodelspseudo-responseexpansion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that conversational information retrieval should not personalize every query turn the same way. Its APCIR system first asks an LLM to label each turn as non-personalized, partially personalized, or fully personalized, using in-context examples and step-by-step reasoning; it then rewrites the turn into several query variants, including a personalized query with a pseudo answer and a neutral query with a pseudo answer. A personalization-aware ranking fusion step combines the retrieval lists from these variants with weights that are optimized separately for each personalization level on validation turns and then applied to test turns. On the TREC iKAT 2023 and 2024 test collections, the paper reports consistent gains over prior systems, including cases where the method beats human-written query rewrites. The practical claim is that explicit, per-turn control of how much personalization to apply is better than relying on an LLM to implicitly decide inside a single prompt.

What carries the argument

The load-bearing mechanism is a personalization-aware ranking fusion controlled by an explicit personalization-level signal. An LLM labels each turn as one of three levels (non-personalized, partially personalized, or fully personalized) using chain-of-thought prompting and in-context examples; the same LLM then generates a personalized reformulation plus pseudo response and a non-personalized reformulation plus pseudo response. After min-max normalization of each retrieval list, the fusion weights for each level are chosen by grid search as the triple that maximizes an evaluation metric over validation turns of that level, subject to the weights summing to one, and the chosen triple is applied to test turns of the same predicted level (Eq. 5, Algorithm 1). The per-level weights are the mechanism that lets the system down-weight personalized results when a turn is predicted not to need them, which is what avoids over-personalization.

What would settle it

Use the human personalization annotations available in the TREC iKAT data as the level signal in place of the LLM's labels and compare fused results; if human-labeled levels substantially outperform LLM-labeled levels, the level classifier itself is a bottleneck, and if they do not, the paper's claim that the gap reflects annotation noise is supported.

Watch

Extended reading notes

Core claim

APCIR's central discovery is that personalized retrieval improves when the amount of personalization is decided per query turn and the decision is used to control fusion, not just prompt content. The paper shows in a preliminary experiment that a fully personalized LLM rewrite alone does not beat a non-personalized rewrite, but fusing the two lists does, and that even ground-truth binary selection of profile sentences does not beat fusion. The method therefore separates three levels of personalization, generates both personalized and non-personalized query variants with pseudo responses, and determines, for each level, the weight triple that maximizes a retrieval metric on a validation set (Eq. 5). Applying those level-specific weights at test time yields the reported gains: MRR 57.3 vs 45.7 on iKAT-23 and 82.7 vs 43.2 on iKAT-24 in the open comparison, with the framework also outperforming the compared baselines when retriever and re-ranker are held fixed across methods.

Load-bearing premise

The method assumes that queries sharing an LLM-predicted personalization level are similar enough across datasets that one weight triple per level, tuned on a validation collection, transfers to the test collection.

Editorial extensions

If this is right

  • Systems can avoid over-personalization by assigning a low fusion weight to personalized rewrites when the predicted level is non-personalized; the optimized weights differ across levels.
  • The adaptive mechanism transfers across backbone LLMs and retrievers: gains hold with GPT-4o, ChatGPT-3.5, LLaMA-3.1-8B and Mistral-2-7B, and with BM25, ANCE and SPLADE.
  • First-stage fusion can matter more than re-ranking: on iKAT-23, applying re-rankers to SPLADE runs degraded several systems, including APCIR, suggesting the fused list is already strong.
  • APCIR can exceed human-written query rewrites on iKAT-24, indicating that explicit level detection plus fusion can substitute for costly manual reformulation in personalized CIR.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If per-level weights generalize across collections, the same three-level taxonomy could be applied to other conversational tasks that mix personal and factual turns; one test is tuning weights on one corpus and transferring to a new domain.
  • Because the level signal is noisy (77.8%/47.8% and 79.6%/63.0% overlap with human judgments), a calibrated confidence score from the LLM could allow soft fusion weights rather than a hard per-level triple, which the paper leaves implicit.
  • Equation (5) optimizes weights against a validation metric, which is close to hyperparameter tuning; the sensitivity of the reported gains to the size and composition of the validation query set is not tested in the paper and would be a direct extension.
  • The finding that annotated personal information does not beat fusion suggests that retrieval-time blending may be more robust than query-rewrite-time filtering; a testable extension would be applying the same fusion idea to other rewrite-based retrieval pipelines.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This paper addresses the question of when and how much to personalize conversational information retrieval. The authors propose APCIR, which first uses GPT-4o with chain-of-thought and in-context examples to label each query turn as non-personalized, partially personalized, or personalized; then generates up to three reformulated query variants per turn (non-personalized query, query plus pseudo-response, and personalized query plus pseudo-response, with a second non-personalized variant for non-personalized turns); and finally fuses the ranking lists with a linear combination whose weights are selected per personalization level by grid search on a validation collection and transferred to the other iKAT collection. Experiments on TREC iKAT 2023 and 2024 report large gains over existing conversational query reformulation baselines in both open and aligned comparisons, with additional analyses of ablation components, weight-identification baselines, fusion strategies, retrievers/re-rankers, and LLM-versus-human personalization judgments.

Significance. The paper makes a clear and practically relevant proposal: instead of always injecting a full user profile into query reformulation, it conditions reformulation and fusion on an explicit personalization level. The aligned comparisons with fixed query rewriter, retriever, and reranker are a strength, as is the cross-dataset protocol for selecting fusion weights without test-label leakage and the release of code and data. The reported improvements are large (e.g., MRR 57.3 vs. 45.7 on iKAT-23 and 82.7 vs. 43.2 on iKAT-24 in the open comparison), and I found no circularity in the described evaluation protocol: the LLM personalization labels are not derived from the relevance labels, and the weight-fitting collection is different from the test collection. The main reason the significance is not yet fully established is that the per-level adaptive mechanism is not compared with a single globally optimized fusion weight, and the optimization metric used for weight selection is unspecified; both are needed to attribute the gains to adaptation rather than to weight tuning.

major comments (5)
  1. [Algorithm 1 and Sec. 5.3] The pseudocode and the experimental protocol are inconsistent. Algorithm 1 takes a single conversation session as input, builds level-specific lists L_l from that session, selects optimal weights by maximizing M on phi_fused(q_n) for the same turns in L_l (lines 14-28), and then applies these weights to the same turns (lines 29-32). If run as written, this fits and evaluates on the same query turns. The text in Sec. 5.3 instead says that weights are selected on one iKAT collection and applied to the other. Please rewrite Algorithm 1 to take a validation session or collection and a separate test session as inputs, and state clearly in the pseudocode where weight selection stops and test evaluation begins. This is load-bearing for every reported number.
  2. [Sec. 4.3, Eq. (5), Table 2] The claimed adaptive mechanism is not isolated. The per-level weight triples are compared with Random, Equal, User-based Entropy, and DEPS, but none of these baselines is allowed to optimize a single global weight triple on the same validation queries using the same grid search. The three fitted triples are numerically close (e.g., iKAT-24: (0.20,0.38,0.42), (0.28,0.36,0.36), and (0.23,0.40,0.37)), so a single triple may account for most of the gain over Equal Personalization. Please add a global-weight baseline and an ablation that removes the level conditioning; without this, the paper's central claim that adaptively conditioning fusion weights on the personalization level drives the Table 1 gains is not established.
  3. [Sec. 4.3, Eq. (5), Algorithm 1] The optimization metric M is never specified. It appears only as 'a retrieval evaluation metric' in Eq. (5) and as the accumulator in Algorithm 1. If M is, say, NDCG@3, then the MRR and Recall numbers in Tables 1-4 are not the objective being optimized, and the reported gains on those metrics could be incidental. Specify M, and ideally repeat the weight search for each reported metric to show that the conclusions are not metric-dependent.
  4. [Tables 1 and 2] No variance or repeated-run information is reported, and significance testing is incomplete. GPT-4o generation is stochastic, yet the paper reports point estimates and marks the dagger symbol only for selected comparisons in Table 1, with no significance tests for Table 2. Report means and standard deviations (or bootstrap confidence intervals) over multiple runs, state the sampling temperature used for the LLM calls, and apply significance tests to the weight-identification comparisons.
  5. [Sec. 5.3, Table 5] The cross-dataset transfer of per-level weights rests on the assumption that LLM-predicted levels are consistent enough across collections. The agreement rates in Table 5 are 77.8%/47.8% (iKAT-23) and 79.6%/63.0% (iKAT-24) for personalized/non-personalized classes; the paper attributes the gap to annotation discrepancy, but this is an interpretation, not a measurement. Please add a sensitivity analysis: fit weights on one split of each dataset and evaluate on the held-out split, and compare the results when the level assignment uses LLM labels versus human labels. This would separate level-prediction noise from the fusion benefit.
minor comments (6)
  1. [Sec. 3.2, Figure 2] The preliminary experiment lacks error bars, numeric values, and significance tests; please add them or state explicitly that the figure is illustrative.
  2. [Sec. 5.3] The reported fusion weight triples are not labeled with their personalization levels; please clarify the order (w1,w2,w3) and which level each triple corresponds to.
  3. [Algorithm 1, lines 5-6] The notation 'if l_n != a' and '{or q'_n,r'_n, q''_n,r''_n} if l_n == a' is ambiguous; rewrite these as explicit conditional statements.
  4. [Eq. (5)] The summation 'sum_{n=1}^N ..., l_n = l' is ambiguous; clarify that the sum is over turns in level l, as in Algorithm 1.
  5. [Table 1] The open comparison for iKAT-24 has '-' for PCIR and GtR; state explicitly that these numbers were not available and were therefore omitted, and consider reporting them with the same backbone if possible.
  6. [Abstract] The phrase 'The results confirm the effectiveness of adaptive personalization' overstates the evidence given the missing global-weight ablation; consider softening to 'suggest' or 'indicate' pending the new analysis.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: per-level fusion weights are fitted on a disjoint validation collection and then applied to test queries; the load-bearing personalization signal comes from an LLM, not from the relevance labels used for scoring.

full rationale

The paper's derivation chain is self-contained in the sense relevant to circularity. The personalization level in Eq. (2) is produced by an LLM with in-context examples and CoT, not fitted to the retrieval labels. The reformulated queries in Eqs. (3)-(4) are generated from the user profile, context, and identified level. The fusion weights in Eq. (5) and Algorithm 1 are optimized per personalization level on a validation set using the retrieval metric M, and the paper explicitly states that the weights are selected on one iKAT collection and applied to the other (Sec. 5.3: 'we use one dataset for it and apply the weights to another for inference, e.g., weights selection on iKAT-23 and test on iKAT-24, vice versa'). Thus the reported test numbers are not forced by the fitted parameters by construction; the validation/test split breaks the self-definitional loop. The personalization-level labels are LLM predictions, not derived from the relevance judgments used to compute MRR/NDCG, so the level identification is not a renamed version of the target metric. Self-citations appear as a baseline (PCIR) and as a supporting remark about annotation discrepancy (Sec. 6.6), but neither is load-bearing for the central claim; the central mechanisms are evaluated against external TREC iKAT collections and against non-self-cited baselines. The skeptic's concern about a missing comparison to a single globally optimized weight triple is a legitimate experimental-control question, but it is not a circularity reduction: failing to ablate the per-level conditioning does not mean the per-level conditioning is equivalent to its inputs by definition. Therefore no circular step can be exhibited from the paper's own equations, and the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 1 invented entities

The central result depends on three fitted weight triples per dataset (six effective degrees of freedom, since each triple sums to one), a hand-designed three-level taxonomy, and several domain assumptions about LLM level identification, cross-dataset transfer, and the surrogate metric used for weight selection. No physical entities or new forces are postulated.

free parameters (3)
  • Fusion weight triples (w1,w2,w3) per personalization level for iKAT-23 = level a: (0.36,0.17,0.47); level b: (0.35,0.2,0.45); level c: (0.25,0.2,0.55)
    Grid search over 0.01-step weight candidates maximizing an evaluation metric on iKAT-23 validation queries (Sec. 5.3); the resulting weights are applied to iKAT-24 test queries.
  • Fusion weight triples (w1,w2,w3) per personalization level for iKAT-24 = level a: (0.2,0.38,0.42); level b: (0.28,0.36,0.36); level c: (0.23,0.4,0.37)
    Grid search on iKAT-24 validation queries (Sec. 5.3); the resulting weights are applied to iKAT-23 test queries.
  • Hand-crafted CoT and in-context demonstration examples = not numeric; randomly sampled between the two datasets
    The prompt includes manually designed examples (L_Example, QR_Example, CoT) that affect both level identification and query reformulation; Sec. 5.3 states the example is randomly selected between datasets, so the choice is an uncontrolled hand-tuned component.
assumptions (5)
  • domain assumption LLM personalization level identification is accurate enough to group query turns into useful fusion-weight buckets
    Stage 1 uses GPT-4o with CoT and in-context examples to label each turn as non, partial, or personalized. Table 5 shows only 77.8%/47.8% (iKAT-23) and 79.6%/63.0% (iKAT-24) overlap with human judgments, yet the method relies on these labels for weight assignment.
  • domain assumption Optimal fusion weights are shared across all turns with the same predicted personalization level and transfer between iKAT-23 and iKAT-24
    Sec. 4.3 states 'we assume that they remain consistent across all turns that share the same level of personalization,' and Sec. 5.3 selects weights on one dataset and applies them to the other. If the level groups are not homogeneous across datasets, the transferred weights may be suboptimal.
  • domain assumption The evaluation metric M in Eq. (5) is a valid surrogate for all reported metrics
    Algorithm 1 maximizes an unspecified metric M during grid search, while the paper reports MRR, NDCG@3, and Recall. Improvements on metrics that are not the optimization target are not directly controlled.
  • domain assumption LLM-generated pseudo responses improve retrieval when concatenated to reformulated queries
    The preliminary experiments in Sec. 3.2 support this for the tested settings, and the fusion pipeline always includes pseudo-response-expanded rewrites, but the assumption is validated only on the two iKAT datasets.
  • domain assumption Min-max normalization makes ranking scores from different query rewrites comparable for linear fusion
    Sec. 4.3 applies f_norm to each ranking list before linear combination; this assumes the normalized scores preserve relative usefulness across lists, which is a modeling choice rather than a proven property.
invented entities (1)
  • Three-level personalization taxonomy (non, partial, personalized)
    purpose: Categorical signal that conditions query reformulation and selects fusion weights for each turn
    The taxonomy is introduced by the authors as part of the LLM instruction. No external benchmark validates the taxonomy itself; its usefulness is measured only through downstream retrieval scores on the two iKAT datasets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adaptive Personalized Conversational Information Retrieval." pith.science (2026). https://pith.science/paper/QBC2T27F

@misc{pith2026250808634,
  author       = {Pith},
  title        = {Pith review of: Adaptive Personalized Conversational Information Retrieval},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QBC2T27F}},
  note         = {Machine review of arXiv:2508.08634}
}
read the original abstract

Personalized conversational information retrieval (CIR) systems aim to satisfy users' complex information needs through multi-turn interactions by considering user profiles. However, not all search queries require personalization. The challenge lies in appropriately incorporating personalization elements into search when needed. Most existing studies implicitly incorporate users' personal information and conversational context using large language models without distinguishing the specific requirements for each query turn. Such a ``one-size-fits-all'' personalization strategy might lead to sub-optimal results. In this paper, we propose an adaptive personalization method, in which we first identify the required personalization level for a query and integrate personalized queries with other query reformulations to produce various enhanced queries. Then, we design a personalization-aware ranking fusion approach to assign fusion weights dynamically to different reformulated queries, depending on the required personalization level. The proposed adaptive personalized conversational information retrieval framework APCIR is evaluated on two TREC iKAT datasets. The results confirm the effectiveness of adaptive personalization of APCIR by outperforming state-of-the-art methods.

Figures

Figures reproduced from arXiv: 2508.08634 by the authors.

Figure 1
Figure 1. Example of different personalized information [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Preliminary experimental results with four differ [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Overview of our APCIR framework and workflow. The personalization granularity is first identified for each given query (Stage 1) and then used to reformulate multiple queries (Stage 2). The personalization level and reformulated queries are provided for personalization-aware ranking fusion (Stage 3) to obtain optimal fusion weights and produce a final ranking list. The experimental results are presented in [PITH_FU… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Ablation studies on the effectiveness of each com [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Impact of using different retrievers and re-rankers. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

86 extracted references · 57 canonical work pages

  1. [1]

    Zahra Abbasiantaeb and Mohammad Aliannejadi. 2024. Generate then Retrieve: Conversational Response Retrieval Using LLMs as Answer and Query Generators. arXiv preprint arXiv:2403.19302 (2024)

  2. [2]

    Zahra Abbasiantaeb, Simon Lupart, Leif Azzopardi, Jeffrey Dalton, and Moham- mad Aliannejadi. 2025. Conversational gold: Evaluating personalized conversa- tional search system using gold nuggets. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3455–3465

  3. [3]

    Vaibhav Adlakha, Shehzaad Dhuliawala, Kaheer Suleman, Harm de Vries, and Siva Reddy. 2022. TopiOCQA: Open-domain Conversational Question Answering with Topic Switching.Transactions of the Association for Computational Linguistics 10 (2022), 468–483

  4. [4]

    Mohammad Aliannejadi, Zahra Abbasiantaeb, Shubham Chatterjee, Jeffery Dal- ton, and Leif Azzopardi. 2024. TREC iKAT 2023: The Interactive Knowledge Assistance Track Overview. arXiv preprint arXiv:2401.01330 (2024)

  5. [5]

    Raviteja Anantha, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021. Open-Domain Question Answering Goes Conversational via Question Rewriting. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 520–534

  6. [6]

    Haonan Chen, Zhicheng Dou, Kelong Mao, Jiongnan Liu, and Ziliang Zhao

  7. [7]

    Zhiyu Chen, Jie Zhao, Anjie Fang, Besnik Fetahu, Oleg Rokhlenko, and Shervin Malmasi. 2022. Reinforced Question Rewriting for Conversational Question Answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track . 357–370

  8. [8]

    Yiruo Cheng, Kelong Mao, and Zhicheng Dou. 2024. Interpreting Conversational Dense Retrieval by Rewriting-Enhanced Inversion of Session Embedding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2879–2893

Show all 86 references
  1. [9]

    Zhuyun Dai, Arun Tejasvi Chaganty, Vincent Y Zhao, Aida Amini, Qazi Mamunur Rashid, Mike Green, and Kelvin Guu. 2022. Dialog Inpainting: Turning Documents into Dialogs. In International Conference on Machine Learning. PMLR, 4558–4586

  2. [10]

    Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020. TREC CAsT 2019: The conversational assistance track overview. arXiv preprint arXiv:2003.13624 (2020)

  3. [11]

    Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2021. CAsT 2020: The Conver- sational Assistance Track Overview . Technical Report

  4. [12]

    Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2022. TREC CAsT 2021: The conversational assistance track overview. In In Proceedings of TREC

  5. [13]

    Zhicheng Dou, Ruihua Song, and Ji-Rong Wen. 2007. A large-scale evaluation and analysis of personalized search strategies. In Proceedings of the 16th international conference on World Wide Web. 581–590

  6. [14]

    Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber. 2019. Can You Un- pack That? Learning to Rewrite Questions-in-Context. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language...

  7. [15]

    Hung-Chieh Fang, Kuo-Han Hung, Chen-Wei Huang, and Yun-Nung Chen. 2022. Open-Domain Conversational Question Answering with Historical Answers. In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022 . 319–326

  8. [16]

    Jianfeng Gao, Chenyan Xiong, Paul Bennett, and Nick Craswell. 2022. Neural ap- proaches to conversational information retrieval. arXiv preprint arXiv:2201.05176 (2022)

  9. [17]

    Junjie Huang, Jiarui Qin, Jianghao Lin, Ziming Feng, Weinan Zhang, and Yong Yu. 2025. Unleashing the Potential of Multi-Channel Fusion in Retrieval for Personalized Recommendations. In Proceedings of the ACM on Web Conference

  10. [18]

    Yunah Jang, Kang-il Lee, Hyunkyung Bae, Seungpil Won, Hwanhee Lee, and Kyomin Jung. 2023. IterCQR: Iterative Conversational Query Reformulation without Human Supervision. arXiv preprint arXiv:2311.09820 (2023)

  11. [19]

    Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2023. InstructoR: Instructing Unsupervised Conversational Dense Retrieval with Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 6649–6675

  12. [20]

    Sungdong Kim and Gangwoo Kim. 2022. Saving dense retriever from shortcut dependency in conversational search. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 10278–10287

  13. [21]

    Vaibhav Kumar and Jamie Callan. 2020. Making Information Seeking Easier: An Improved Pipeline for Conversational Search. In Empirical Methods in Natural Language Processing

  14. [22]

    Oren Kurland and J Shane Culpepper. 2018. Fusion in information retrieval: Sigir 2018 half-day tutorial. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval . 1383–1386

  15. [23]

    Yilong Lai, Jialong Wu, Congzhi Zhang, Haowen Sun, and Deyu Zhou. 2025. AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment. InProceedings of the 31st International Conference on Computational Linguistics. 7698–7720

  16. [24]

    Carlos Lassance and Stéphane Clinchant. 2023. Naver Labs Europe (SPLADE)@ TREC Deep Learning 2022. arXiv preprint arXiv:2302.12574 (2023)

  17. [25]

    Carlos Lassance, Hervé Déjean, Thibault Formal, and Stéphane Clinchant. 2024. SPLADE-v3: New baselines for SPLADE. arXiv preprint arXiv:2403.06789 (2024)

  18. [26]

    Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. Contextualized Query Embeddings for Conversational Search. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 1004–1015

  19. [27]

    Sheng-Chieh Lin, Jheng-Hong Yang, Rodrigo Nogueira, Ming-Feng Tsai, Chuan- Ju Wang, and Jimmy Lin. 2020. Conversational question reformulation via sequence-to-sequence architectures and pretrained language models. arXiv preprint arXiv:2004.01909 (2020)

  20. [28]

    Wenhan Liu, Yujia Zhou, Yutao Zhu, and Zhicheng Dou. 2024. How to personal- ize and whether to personalize? Candidate documents decide. Knowledge and Information Systems 66, 9 (2024), 5581–5604

  21. [29]

    Simon Lupart, Zahra Abbasiantaeb, and Mohammad Aliannejadi. 2024. IRLab@ iKAT24: Learned Sparse Retrieval with Multi-aspect LLM Query Generation for Conversational Search. arXiv preprint arXiv:2411.14739 (2024)

  22. [30]

    Simon Lupart, Mohammad Aliannejadi, and Evangelos Kanoulas. 2025. DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval . 9–19

  23. [31]

    Kelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo, Zheng Liu, Tetsuya Sakai, and Zhicheng Dou. 2024. ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval. arXiv preprint arXiv:2404.13556 (2024)

  24. [32]

    Kelong Mao, Zhicheng Dou, Haonan Chen, Fengran Mo, and Hongjin Qian. 2023. Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search. In Findings of the Association for Compu- tational Linguistics: EMNLP 2023

  25. [33]

    Kelong Mao, Zhicheng Dou, Bang Liu, Hongjin Qian, Fengran Mo, Xiangli Wu, Xiaohua Cheng, and Zhao Cao. 2023. Search-Oriented Conversational Query Editing. In Findings of the Association for Computational Linguistics: ACL 2023 . 4160–4172

  26. [34]

    Kelong Mao, Zhicheng Dou, and Hongjin Qian. 2022. Curriculum Contrastive Context Denoising for Few-shot Conversational Dense Retrieval. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 176–186

  27. [35]

    Kelong Mao, Zhicheng Dou, Hongjin Qian, Fengran Mo, Xiaohua Cheng, and Zhao Cao. 2022. ConvTrans: Transforming Web Search Sessions for Conversa- tional Dense Retrieval. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 2935–2946

  28. [36]

    Kelong Mao, Hongjin Qian, Fengran Mo, Zhicheng Dou, Bang Liu, Xiaohua Cheng, and Zhao Cao. 2023. Learning Denoised and Interpretable Session Representation for Conversational Search. In Proceedings of the ACM Web Conference 2023 . 3193– 3202

  29. [37]

    Qiaozhu Mei and Kenneth Church. 2008. Entropy of search logs: how hard is search? with personalization? with backoff?. In Proceedings of the 2008 Interna- tional Conference on Web Search and Data Mining . 45–54

  30. [38]

    Chuan Meng, Negar Arabzadeh, Mohammad Aliannejadi, and Maarten de Rijke

  31. [39]

    Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2025. Query Performance Prediction using Relevance Judg- ments Generated by Large Language Models. ACM Transactions on Information Systems (TOIS) (2025)

  32. [40]

    Chuan Meng, Francesco Tonolini, Fengran Mo, Nikolaos Aletras, Emine Yilmaz, and Gabriella Kazai. 2025. Bridging the gap: From ad-hoc to proactive search in conversations. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information R...

  33. [41]

    Alessandro Micarelli, Fabio Gasparetti, Filippo Sciarrone, and Susan Gauch. 2007. Personalized search on the world wide web. The adaptive web: Methods and strategies of web personalization (2007), 195–230

  34. [42]

    Fengran Mo, Yifan Gao, Chuan Meng, Xin Liu, Zhuofeng Wu, Kelong Mao, Zhengyang Wang, Pei Chen, Zheng Li, Xian Li, et al. 2025. UniConv: Unifying Retrieval and Response Generation for Large Language Models in Conversations. In Proceedings of the 63rd Annual Meeting of the Assoc...

  35. [43]

    Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh, Boxing Chen, Qun Liu, and Jian-Yun Nie. 2024. CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search. arXiv preprint arXiv:2406.05013 (2024). Adaptive Personalized Conversational ...

  36. [44]

    Fengran Mo, Kelong Mao, Ziliang Zhao, Hongjin Qian, Haonan Chen, Yiruo Cheng, Xiaoxi Li, Yutao Zhu, Zhicheng Dou, and Jian-Yun Nie. 2025. A survey of conversational search. ACM Transactions on Information Systems (TOIS) (2025)

  37. [45]

    Fengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu, Kaiyu Huang, and Jian-Yun Nie

  38. [46]

    Fengran Mo, Chuan Meng, Mohammad Aliannejadi, and Jian-Yun Nie. 2025. Con- versational search: From fundamentals to frontiers in the LLM era. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 4094–4097

  39. [47]

    Fengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao, Yutao Zhu, Peng Li, and Yang Liu. 2023. Learning to Relate to Previous Turns in Conversational Search. In 29th ACM SIGKDD Conference On Knowledge Discover and Data Mining (SIGKDD)

  40. [48]

    In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics

    ConvGQR: Generative Query Reformulation for Conversational Search. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. 4998–5012

  41. [49]

    Fengran Mo, Bole Yi, Kelong Mao, Chen Qu, Kaiyu Huang, and Jian-Yun Nie. 2024. ConvSDG: Session Data Generation for Conversational Search. In Companion Proceedings of the ACM on Web Conference 2024 . 1634–1642

  42. [50]

    Fengran Mo, Jinghan Zhang, Yuchen Hui, Jia Ao Sun, Zhichao Xu, Zhan Su, and Jian-Yun Nie. 2025. ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense Retrieval. arXiv preprint arXiv:2508.04001 (2025)

  43. [51]

    Fengran Mo, Chen Qu, Kelong Mao, Tianyu Zhu, Zhan Su, Kaiyu Huang, and Jian-Yun Nie. 2024. History-Aware Conversational Dense Retrieval.arXiv preprint arXiv:2401.16659 (2024)

  44. [52]

    Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Docu- ment Ranking with a Pretrained Sequence-to-Sequence Model. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 708–718

  45. [53]

    OpenAI. 2024. https://openai.com/index/hello-gpt-4o/. blog (2024)

  46. [54]

    Fengran Mo, Longxiang Zhao, Kaiyu Huang, Yue Dong, Degen Huang, and Jian- Yun Nie. 2024. How to leverage personal textual knowledge for personalized conversational information retrieval. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Manag...

  47. [55]

    Paul Owoicho, Jeffrey Dalton, Mohammad Aliannejadi, Leif Azzopardi, Johanne R Trippas, and Svitlana Vakulenko. 2022. TREC CAsT 2022: Going beyond user ask and system retrieve with initiative and response generation. NIST Special Publication (2022), 500–338

  48. [56]

    Quinn Patwardhan and Grace Hui Yang. 2023. Sequencing Matters: A Generate- Retrieve-Generate Model for Building Conversational Agents. arXiv preprint arXiv:2311.09513 (2023)

  49. [57]

    Arnold Overwijk, Chenyan Xiong, and Jamie Callan. 2022. ClueWeb22: 10 billion web documents with rich information. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3360–3362

  50. [58]

    Chen Qu, Liu Yang, Cen Chen, Minghui Qiu, W Bruce Croft, and Mohit Iyyer

  51. [59]

    Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at TREC-3. Nist Special Publication Sp 109 (1995), 109

  52. [60]

    Hongjin Qian and Zhicheng Dou. 2022. Explicit Query Rewriting for Conversa- tional Dense Retrieval. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 4725–4737

  53. [61]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investi- gating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...

  54. [62]

    Bin Tan, Xuehua Shen, and ChengXiang Zhai. 2006. Mining long-term search history to improve search accuracy. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining . 718–723

  55. [63]

    Jaime Teevan, Susan T Dumais, and Daniel J Liebling. 2008. To personalize or not to personalize: modeling queries with variation in user intent. In Proceedings of the 31st annual international ACM SIGIR conference on Research and development in information retrieval. 163–170

  56. [64]

    Mirco Speretta and Susan Gauch. 2005. Personalized search based on user search histories. In The 2005 IEEE/WIC/ACM International Conference on Web Intelligence (WI’05). IEEE, 622–628

  57. [65]

    Christophe Van Gysel and Maarten de Rijke. 2018. Pytrec_eval: An Extremely Fast Python Interface to trec_eval. In SIGIR. ACM

  58. [66]

    Nikos Voskarides, Dan Li, Pengjie Ren, Evangelos Kanoulas, and Maarten de Rijke. 2020. Query resolution for conversational search with limited supervision. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 921–930

  59. [67]

    Hongning Wang, Xiaodong He, Ming-Wei Chang, Yang Song, Ryen W White, and Wei Chu. 2013. Personalized ranking model adaptation for web search. In Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval. 323–332

  60. [68]

    Svitlana Vakulenko, Shayne Longpre, Zhucheng Tu, and Raviteja Anantha. 2021. Question rewriting for conversational question answering. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining . 355–363

  61. [69]

    Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, and Gaurav Singh Tomar

  62. [70]

    Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate Nearest Neighbor Neg- ative Contrastive Learning for Dense Text Retrieval. In International Conference on Learning Representations

  63. [71]

    Yiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu, Wenjie Wang, Fuli Feng, Hamed Zamani, Xiangnan He, and Tat-Seng Chua. 2025. Personalized Generation In Large Model Era: A Survey. arXiv preprint arXiv:2503.02614 (2025)

  64. [72]

    Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query Expansion with Large Language Models. arXiv preprint arXiv:2303.07678 (2023)

  65. [73]

    Chanwoong Yoon, Gangwoo Kim, Byeongguk Jeon, Sungdong Kim, Yohan Jo, and Jaewoo Kang. 2024. Ask Optimal Questions: Aligning Large Language Models with Retriever’s Preference in Conversational Search. arXiv preprint arXiv:2402.11827 (2024)

  66. [74]

    Shi Yu, Jiahua Liu, Jingqin Yang, Chenyan Xiong, Paul Bennett, Jianfeng Gao, and Zhiyuan Liu. 2020. Few-shot generative conversational query rewriting. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval . 1933–1936

  67. [75]

    Shi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng, and Zhiyuan Liu. 2021. Few- shot conversational dense retrieval. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 829–838

  68. [76]

    Trippas, Jeffrey Dalton, and Filip Radlinski

    Hamed Zamani, Johanne R. Trippas, Jeffrey Dalton, and Filip Radlinski. 2022. Conversational Information Seeking. Found. Trends Inf. Retr. 17 (2022), 244–456. https://api.semanticscholar.org/CorpusID:246210119

  69. [77]

    Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz. 2023. Enhancing Conver- sational Search: Large Language Model-Aided Informative Query Rewriting. In Findings of the Association for Computational Linguistics: EMNLP 2023. 5985–6006

  70. [78]

    Zhehao Zhang, Ryan A Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, et al. 2024. Per- sonalization of large language models: A survey. arXiv preprint arXiv:2411.00027 (2024)

  71. [79]

    Yujia Zhou, Zhicheng Dou, and Ji-Rong Wen. 2020. Encoding history with context-aware representation learning for personalized search. In Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval. 1111–1120

  72. [80]

    Yujia Zhou, Qiannan Zhu, Jiajie Jin, and Zhicheng Dou. 2024. Cognitive per- sonalized search integrating large language models with an efficient memory mechanism. In Proceedings of the ACM on Web Conference 2024 . 1464–1473

  73. [82]

    Tianhua Zhang, Kun Li, Hongyin Luo, Xixin Wu, James Glass, and Helen Meng

  74. [83]

    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

    Adaptive Query Rewriting: Aligning Rewriters through Marginal Probabil- ity of Conversational Answers. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 13444–13461

  75. [2020]

    InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval

    Open-retrieval conversational question answering. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 539–548

  76. [2022]

    CONQRR: Conversational Query Rewriting for Retrieval with Reinforce- ment Learning. (2022)

  77. [2023]

    In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval

    Query Performance Prediction: From Ad-hoc to Conversational Search. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2583–2593

  78. [2024]

    In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Generalizing Conversational Dense Retrieval via LLM-Cognition Data Augmentation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2700–2718

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.