{"id":"50d3b502-9cb6-4f5b-a190-a0e8ccdb9505","arxiv_id":"2411.09166","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured survey of dialogue systems that use unstructured text as external knowledge, organizing datasets, retrieval and generative model components, evaluation metrics, and future directions.","lead":"This paper surveys computer systems that make chatbots more informative by feeding them unstructured text such as Wikipedia articles or personal profiles, and sorts the field into retrieval-based and generation-based approaches. It is a useful map of pre-2021 research for anyone entering knowledge-grounded dialogue, though it stops before the large language model era.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bearing concern: the survey asserts 'first systematic review' and future trends without any documented search protocol, and its bibliography ends in 2021 while the arXiv posting is 2024; coverage completeness is therefore unverified.","rationale":"The reader's weakest_assumption isolates exactly the load-bearing premise: the surveyed literature is complete and current enough to support both the 'first systematic review' and the future-trends projections. I agree with that identification because the paper makes these claims without a search protocol and without engaging any work after 2021, despite the 2024 arXiv posting. The concern is not about internal consistency of the taxonomy; the component-based decomposition of retrieval and generative models is coherent and the historical tables are a useful contribution for the pre-2021 period. The problem is that the central contribution is framed as a current, systematic review, and that framing cannot be validated from the manuscript as submitted. A reproducible literature search is the natural check: it will either confirm that no earlier or more recent systematic survey exists, or it will show that the 'first' and 'future trends' claims need to be explicitly scoped to an earlier period. I do not think this rises to rejection, because a conditional acceptance with a request to reframe as a historical survey, add a methodology section, and update or remove the future-trends claims is proportionate. I also considered the Section 6.3 claim that 'adding a copy mechanism is always helpful,' which appears overstated given heterogeneous experimental settings, but it is a secondary conclusion rather than the paper's central claim; the coverage and protocol issue is more load-bearing for the survey's overall validity.","tokens_in":54516,"tokens_out":6242,"duration_ms":77382,"concrete_test":"Reproduce the coverage claim with a documented literature search: query ACL Anthology, DBLP, and Google Scholar for 2016-01-01 to 2024-11-14 using combinations such as ('knowledge-grounded dialogue' OR 'document-grounded conversation' OR 'background-based conversation' OR 'unstructured text dialogue') AND ('survey' OR 'review' OR 'overview'); then apply explicit inclusion criteria (peer-reviewed survey/review whose scope covers open-domain dialogue grounded in unstructured text) and screen the hits against the paper's Table 3, Table 5, and Table 6. If any prior work, such as Yu et al. 2020 or a 2021-2024 survey of RAG-based dialogue, satisfies those criteria, the 'first systematic review' and the future-trend claims must be narrowed or revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central value proposition is that it is the first systematic review of UTEDS (Section 1: 'As far as we know, we are the first to make a systematic review of the UTEDS') and that its Section 7 future-trends analysis reflects where the field is heading. Both claims presuppose that the set of works reviewed is the complete relevant literature. The manuscript provides no search methodology: no databases, query strings, years, inclusion/exclusion criteria, screening process, or PRISMA-style flow. The reference list's newest entries are from 2021 (e.g., Wu et al. 2021, Fan et al. 2021, Deriu et al. 2021, and the DSTC9 track), matching the ACM TOIS copyright line, while the arXiv submission is dated 14 Nov 2024. Thus no reader can verify that earlier surveys (the paper cites Guo et al. 2021, Santhanam and Shaikh 2019, Yu et al. 2020 and dismisses them as 'incomplete') were not also systematic within their stated scope, nor that 2021-2024 developments (retrieval-augmented generation, LLM-centric knowledge grounding, long-context in-context learning, agentic retrieval) are reflected. If the UTEDS field evolved substantially after 2021, the Section 7 trends (e.g., model-based evaluation, lifelong learning) are a 2020 snapshot and may mislead a 2024 reader. The historical taxonomy in Sections 3-6 may remain a useful map of the pre-2021 literature; the claim to be a systematic, current review is what is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a survey of open-domain dialogue systems that augment response generation or retrieval with unstructured text as external knowledge (UTEDS). It defines UTEDS-related concepts, organizes datasets by source and type, and introduces a component-based taxonomy: retrieval models are decomposed into Fusion, Matching, and Ranking; generative models into Dialogue and Knowledge Encoding, Knowledge Selection, and Response Generation. Sections 5 and 6 review automatic and human evaluation metrics and compile reported results from the primary literature, and Section 7 proposes six future research directions. The authors state in Section 1 that this is the first systematic review of UTEDS.","tokens_in":54781,"tokens_out":6843,"duration_ms":75731,"significance":"If the claims are properly scoped, the paper offers a coherent map of the pre-2021 UTEDS literature. The dataset table (Table 3) and the model tables (Tables 5-11) are broad, and the component-wise decomposition is a useful organizing device for newcomers. The paper is also careful to state in Section 6 that cross-paper numbers 'can only be used as a reference' because of differing data-processing schemes, and it interprets the compiled results rather than merely transcribing them. These strengths make the historical taxonomy genuinely useful. However, the survey's status as a 'systematic' and current review is not established: no search protocol is reported, and the bibliography ends around 2021 while the arXiv version is dated November 2024. Those issues bear directly on the paper's central claims, not on the historical taxonomy itself.","major_comments":[{"comment":"The central claim 'As far as we know, we are the first to make a systematic review of the UTEDS' is not verifiable as stated. The manuscript reports no search methodology: no databases, query strings, years, inclusion/exclusion criteria, screening steps, or PRISMA-style flow are given. The references cited as prior work (Guo et al. [61], Santhanam and Shaikh [155], Yu et al. [224]) are dismissed as 'incomplete' without a documented overlap analysis. Because the paper is posted to arXiv in November 2024 while the reference list's newest entries are from 2021 (e.g., [31], [41], [68], [80]), the reader cannot determine whether the surveyed set is complete even for the claimed pre-2021 scope, let alone whether the survey is the first systematic treatment. I recommend adding a methodology section and either updating coverage to 2024 or explicitly reframing the paper as a historical survey of work up to 2021.","section":"Section 1"},{"comment":"The future-trends discussion is presented as guidance for current research but does not engage with post-2021 developments that are directly relevant to UTEDS, such as retrieval-augmented generation, LLM-based knowledge grounding, long-context in-context learning, and agentic retrieval. As a consequence, Section 7's six directions (limited resources, semantic representation, knowledge mining, interpretable reasoning, evaluation, lifelong learning) read as a 2020-2021 snapshot. If the authors intend the arXiv version to be current, Section 7 must be revised to address these developments, or the entire paper must be explicitly dated as a survey of the literature up to 2021. This is load-bearing because the paper's value proposition includes projecting 'future development trends' for the field.","section":"Section 7"}],"minor_comments":[{"comment":"Some statistics are marked as computed by the authors ('*'), but the estimation method is not described; please state how averages and totals were computed and how missing entries (e.g., T-Chat Words/Text) are handled.","section":"Section 2, Table 4"},{"comment":"The description of FIRE is hard to follow because the superscript/subscript notation for the intermediate tensors is not introduced; a small diagram or a notation table would improve readability.","section":"Section 3.2"},{"comment":"The conclusion 'Adding a copy mechanism is always helpful' is too strong; the compiled tables do not isolate the copy mechanism in controlled comparisons, and several strong models in Tables 10-11 do not use it. Please soften this to 'copy mechanisms have been reported to help on specific datasets.'","section":"Section 6.3"},{"comment":"The display of the R@1/2/5 columns is malformed, with repeated headers and dashes that make some rows ambiguous; please re-typeset the table so each entry is readable.","section":"Table 8"},{"comment":"There are numerous typos and inconsistent abbreviations (e.g., 'Tabel 1', 'CUMDoG', 'metircs', 'A verage length', and the alternation between UTEDS and UTED); a careful copy-edit would improve the presentation.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main obstacle is the gap between the 'systematic, first, current' framing and the absence of any search protocol plus a 2021-era bibliography on a November 2024 arXiv posting. This is fixable in revision, but it must be addressed head-on. Note also that several surveyed models are the authors' own (CAT [110], GDR [165], Persona-CVAE [166], RCDG [167]); the reporting appears balanced, but a declaration of interest would be useful. The ACM TOIS 2021 copyright line should be reconciled with the arXiv version's date."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, the taxonomy is genuinely good: splitting retrieval models into Fusion/Matching/Ranking and generative models into Encoding/Knowledge Selection/Response Generation gives a clean map of the pre-2021 UTEDS literature, and the dataset tables plus the careful performance caveats (numbers not directly comparable across papers) are worth having. Second, it is not a 2024 survey. The bibliography and the ACM TOIS copyright line stop in 2021, and the arXiv posting is November 2024. There is no stated search protocol, no inclusion criteria, and no discussion of retrieval-augmented generation or LLM-based grounding, which is where the field actually went after 2021.\n\nWhat the paper does well: the module decomposition is a real organizing contribution, not just a list of models. The dataset comparison (knowledge source, who sees the text, who leads, labeling) is thorough and useful for newcomers. The conclusions in Section 6 — copy mechanisms help, pretrained models help, knowledge selection is central, automatic metrics are unreliable — are actually supported by the compiled tables, and the authors are appropriately careful about not over-reading cross-paper numbers.\n\nThe soft spots are real but proportionate. The claim to be the first systematic review of UTEDS is unverifiable without a search methodology; the earlier surveys cited (Guo et al., Santhanam and Shaikh, Yu et al.) are dismissed as incomplete, but the grounds for that dismissal are not argued. The future-trends discussion reads as a 2020 snapshot — model-based evaluation and lifelong learning are plausible, but the paper says nothing about in-context learning, Agentic recall, or the RAG paradigm that defines current practice. The self-citations (CAT, GDR, Persona-CVAE, RCDG) are fine; those are legitimate works in the surveyed area, not padding.\n\nNone of this invalidates the historical taxonomy. If the paper is reframed as a systematic survey of the 2018--2021 period, it holds up well. As a 2024 submission, it needs either a substantial update or an explicit historical framing with a dated scope statement.\n\nI'd send this to a serious referee, but with a clear mandate: the authors must either add a methodology section and update through at least 2023, or honestly scope the paper as a historical survey. The taxonomy deserves to be available; the 'first' and 'future trends' claims do not.","headline":"A useful module-based taxonomy of knowledge-grounded dialogue up to 2021, but the 'first systematic survey' and future-trends claims are undercut by a missing search protocol and a reference list that stops before the LLM era.","tokens_in":55358,"tokens_out":2007,"would_cite":true,"duration_ms":25669,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first systematic survey of open-domain dialogue systems that ground responses in unstructured text, and it organizes the field into a six-module blueprint.","keywords":["unstructured text enhanced dialogue system","open-domain dialogue","knowledge-grounded conversation","document-grounded dialogue","knowledge selection","retrieval-based dialogue","generative dialogue models","dialogue evaluation"],"falsifier":"Run a documented search of a major NLP anthology for knowledge-grounded or document-grounded dialogue restricted to publications before 2021 and count peer-reviewed papers that the survey neither cites nor fits into its six components; if the count is substantial, the completeness behind the 'first systematic review' and the trend projections is not supported.","tokens_in":54263,"feed_emoji":"🗺️","tokens_out":7948,"duration_ms":82406,"temperature":0.7,"pith_summary":"This paper claims to be the first systematic review of open-domain dialogue systems that use unstructured text, such as Wikipedia articles, movie plots, personas, and reviews, as external knowledge; the paper calls these UTEDS. It sorts the field into a component-level map: retrieval models are built from Fusion, Matching, and Ranking modules, while generative models are built from Dialogue and Knowledge Encoding, Knowledge Selection, and Response Generation modules. The survey compiles the datasets, architectures, training objectives, evaluation metrics, and reported scores behind those components, and draws comparative conclusions from the compiled tables. The authors argue that mining unstructured text during conversation is the future direction for open-domain dialogue research, because most human knowledge is stored in raw text.","feed_headline":"First survey maps text-grounded chatbot research into six modules","feed_subtitle":"Retrieval systems split into fusion, matching, ranking; generative systems into encoding, selection, generation.","key_machinery":"The central organizing device is a module decomposition: retrieval systems as Fusion, Matching, and Ranking, and generative systems as Dialogue and Knowledge Encoding, Knowledge Selection, and Response Generation. Knowledge Selection is identified as the core component, because the model must decide which piece of raw text is semantically and logically relevant before the response can be grounded. The decomposition does the argument's work by giving every surveyed model a coordinate, letting the authors compare datasets, losses, metrics, and reported results on a common grid.","core_discovery":"On the paper's own terms, the discovery is that the scattered UTEDS literature can be read as instances of a single blueprint, so that otherwise unrelated models become comparable module by module. Retrieval systems all fuse dialogue context with external knowledge, match the fused representation against candidate responses, and rank the candidates; generative systems all encode context and knowledge, select the relevant knowledge, and generate a response conditioned on that selection. Under that blueprint the survey catalogs roughly eighteen datasets, compares dozens of models on shared benchmarks such as WoW, Persona-Chat, CMUDoG, Holl-E, and Topical-Chat, and tabulates their reported automatic scores. Its comparative reading supports conclusions that copy mechanisms reliably help, pre-trained models help in most settings, knowledge selection is the central difficulty, and word-overlap automatic metrics do not track dialogue quality. The paper presents this unified view as the first systematic review of UTEDS and uses it to project six research directions: limited-resource learning, semantic representation learning, knowledge mining, interpretable reasoning, better evaluation, and lifelong learning.","pith_inferences":["The Fusion-Matching-Ranking and Encoding-Selection-Generation blueprint reads naturally as an early map of retrieval-augmented generation, where query-document fusion, reranking, and grounded decoding occupy the same coordinates; extending the map to modern large-language-model chat systems is a step the paper itself does not take.","The paper's negative result on automatic metrics points toward reference-free model-based evaluation; a testable extension is to score the compiled benchmark outputs with a modern reference-free judge and check correlation with the human judgments the survey collected.","The explicit-versus-implicit selection distinction corresponds to hard-versus-soft retrieval in current systems, and a testable extension is whether making selection explicit improves knowledge attribution when outputs are fact-checked against the source documents.","Because the compiled results run through roughly 2021, the future-trend list predates the large-language-model wave; updating the survey's six directions for decoder-only instruction-tuned models is a natural continuation that the authors leave implicit."],"forward_implications":["Failures in a grounded dialogue system can be localized to one component, so improvement efforts can target Fusion versus Matching versus Ranking, or Encoding versus Selection versus Generation, rather than treating the system as a black box.","The compiled evidence says copy mechanisms and pre-trained encoders are the interventions that pay off, so future grounded generation systems should treat both as default components.","Because attention blurs over long documents and explicit selection lets models measure selection accuracy, document-grounded dialogue needs explicit selection or hierarchical memory rather than plain attention.","Word-overlap automatic metrics are unreliable, so credible progress claims for UTEDS need model-based or human evaluation alongside automatic scores.","The paper's six future directions define the open problems of the field as the authors see them, from knowledge-sparse and low-resource settings to interpretable reasoning and lifelong learning."],"supporting_citations":[{"why":"Supplies the Wizard of Wikipedia dataset and the Transformer Memory Network that anchors the explicit knowledge-selection comparisons.","marker":"[33]"},{"why":"Supplies Persona-Chat and the key-value Profile Memory baseline used across retrieval and generative tables.","marker":"[227]"},{"why":"Supplies CMUDoG, the document-grounded conversation setting that motivates early interaction and copy-based generation.","marker":"[246]"},{"why":"Supplies Holl-E and the GTTP baseline, the reference for background-based conversation with labeled knowledge spans.","marker":"[127]"},{"why":"Supplies Topical-Chat, the multi-source factual knowledge dataset used as a common generative testbed.","marker":"[57]"},{"why":"Supplies the Conversing by Reading task that motivates on-demand machine reading and MRC-style knowledge selection.","marker":"[141]"},{"why":"Supplies the DSTC-7 sentence-generation track that makes the UTEDS task a community benchmark.","marker":"[223]"},{"why":"The earlier knowledge-enhanced text generation survey whose incomplete coverage of text-grounded dialogue motivates the paper's systematic claim.","marker":"[224]"},{"why":"An earlier dialogue-systems survey that the paper says did not cover UTEDS, supporting the claim of first systematic review.","marker":"[71]"}],"fun_headline_variants":["First systematic survey breaks open-domain chatbots into six modules","Chatbot research unified: six-module blueprint from new survey","Survey maps text-grounded chatbots: retrieval vs generative modules","Six-module framework for text-enhanced dialogue systems","First survey of text-grounded chatbots: fusion to ranking, encoding to selection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes the works it catalogs, gathered without a stated search protocol or inclusion criteria, are the complete relevant literature on text-grounded dialogue systems, so its 'first systematic review' and future-trend claims depend on nothing important being left out.","fun_headline_variants_meta":{"raw":{"variants":["First systematic survey breaks open-domain chatbots into six modules","Chatbot research unified: six-module blueprint from new survey","Survey maps text-grounded chatbots: retrieval vs generative modules","Six-module framework for text-enhanced dialogue systems","First survey of text-grounded chatbots: fusion to ranking, encoding to selection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00024,"raw_usage":{"total_tokens":1519,"prompt_tokens":948,"completion_tokens":571,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":564,"tokens_out":571,"duration_ms":6134,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:57:09.014718+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a documented search of a major NLP anthology for knowledge-grounded or document-grounded dialogue restricted to publications before 2021 and count peer-reviewed papers that the survey neither cites nor fits into its six components; if the count is substantial, the completeness behind the 'first systematic review' and the trend projections is not supported.","supporting_citations":[{"cited_title":"Dialog System Technology Challenge 7","cited_arxiv_id":"1901.03461","evidence_quote":"Supplies the DSTC-7 sentence-generation track that makes the UTEDS task a community benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"An earlier dialogue-systems survey that the paper says did not cover UTEDS, supporting the claim of first systematic review."}],"review_version":1}