{"id":"73606784-0472-47c4-956a-e309df2a5be3","arxiv_id":"2412.19823","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic survey of 108 papers classifies how large language models are used for communication network and service management across four network domains.","lead":"Large language models are being applied across mobile, vehicular, cloud, and edge networks to monitor, plan, deploy, and support communication services. This survey organizes that body of work into one taxonomy and lists open challenges for researchers and operators.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section I-D 'first extensive survey' claim rests on an unverified absence of prior multi-domain surveys and on a corpus that deliberately excludes most preprints; both assumptions need checking.","rationale":"The reader's weakest_assumption identifies the completeness and representativeness of the literature selection as the load-bearing assumption. I agree, and I would sharpen it to the two conditions that the 'first extensive survey' claim requires: the absence of any prior same-scope survey, and a representative corpus of primary studies. The paper documents a PRISMA flow and 108 included papers, which is genuine evidence of systematic effort, but it does not document a search for prior multi-domain surveys, and its own Fig. 2 discloses the exclusion of most preprints. Both gaps bear directly on the central novelty claim. The concrete test I propose targets the first condition directly: a search for prior surveys with the same multi-domain scope, run across the same databases plus arXiv without preprint exclusion. If the search finds no prior survey, the firstness claim is validated despite possible missed primary papers; if it finds one, the central claim fails. The second condition is harder to settle definitively, but a preprint-inclusive search would at least reveal whether the excluded literature changes the taxonomy. Since the reader already assigned CONDITIONAL, and my concern reinforces rather than overturns that verdict, I leave the verdict unchanged.","tokens_in":43582,"tokens_out":4571,"duration_ms":45112,"concrete_test":"Perform a systematic search for prior surveys or reviews with the same four-domain NSM scope (mobile/IoT, vehicular, cloud, fog/edge) using the paper's own sources plus arXiv, with queries such as: 'large language models' AND ('network management' OR 'service management') AND ('mobile' OR 'vehicular' OR 'cloud' OR 'edge') AND ('survey' OR 'review'), without excluding preprints. If this search surfaces a prior multi-domain LLM-for-NSM survey, the Section I-D firstness claim is undercut; if it does not, the central claim survives. In the same search, record how many relevant preprint-only works are found; substantial numbers would additionally call the representativeness of the 108-paper corpus into question.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is firstness: 'the first extensive survey conducted on LLM-enabled NSM in communication networks involving mobile networks and IoT technologies, vehicular networks, cloud-based networks, and fog/edge-based networks' (Section I-D). For that claim to hold, two conditions must be true: (1) no prior survey with the same multi-domain NSM scope exists, and (2) the survey's 108-paper corpus is representative of the field. Condition (1) is not established by the paper's methodology. The related-work comparison in Table I covers nine surveys, but the PRISMA methodology in Section I-E reports no systematic search for prior surveys or reviews; the keyword searches target primary studies, and the initial pool of 14 survey papers is not described as being screened for scope overlap with this work. Condition (2) is explicitly weakened by the authors' own Fig. 2, which states 'Excluded most preprints to ensure quality control' (Section I-E4). In a fast-moving 2023-2024 area like LLM-for-NSM, much influential work appears first as arXiv preprints, so excluding most of them biases the corpus toward already-published work and away from the state of the art at the stated access dates of September 6-8, 2024. A taxonomy built on this non-representative subset may miss NSM task categories that exist only in the preprint literature. The firstness claim is therefore under-supported by the paper's own documented search and selection process.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys applications of Large Language Models (LLMs) to communication network and service management (NSM). It uses a PRISMA-based methodology to select 108 papers and organizes them under a four-part taxonomy: network monitoring and reporting, AI-powered network planning, network deployment and distribution, and continuous network support, across four network domains (mobile/IoT, vehicular, cloud, and fog/edge). The paper also provides a tutorial on LLM fundamentals, summarizes related work in Table I, and concludes with challenges and future directions. The central claim, stated in Section I-D, is that this is the first extensive survey covering LLM-enabled NSM across those four multi-domain communication network types.","tokens_in":43836,"tokens_out":5651,"duration_ms":50185,"significance":"If the coverage is as complete as claimed, this survey would be a useful entry point for researchers and practitioners working on LLM-based network management. The explicit PRISMA flow diagram, the disclosed keyword search list and access dates, and the cross-domain taxonomy are methodological strengths that are often missing from other surveys. The paper also offers structured tables of primary works, which are convenient for locating relevant papers by domain and task. However, the significance of the survey is contingent on two assumptions that are not fully established: the validity of the 'first extensive survey' claim, and the representativeness of a corpus that deliberately excludes most preprints.","major_comments":[{"comment":"The 'first extensive survey' claim in Section I-D is not established by the methodology described in Section I-E. The PRISMA keyword search targets primary studies on LLMs for NSM, but the paper does not report a systematic search for prior surveys or reviews with comparable multi-domain scope, and Table I compares only nine selected surveys. To support the firstness claim, the authors should either perform and document a systematic survey-of-surveys search (including preprint servers and with explicit inclusion criteria) or revise the claim to a weaker statement that reflects the search actually conducted.","section":"Section I-D; Section I-E1"},{"comment":"The eligibility-stage removal of 'most preprints' (Section I-E4, Fig. 2) introduces a potential selection bias that is load-bearing for the survey's coverage claim. For a field like LLM-for-NSM, where a large share of 2023-2024 advances appears first as arXiv preprints, excluding most preprints at an access date of September 6-8, 2024 can systematically omit recent task categories and results. The authors should quantify how many preprints were retained versus excluded, justify the retention criteria, and add a limitation paragraph explaining the effect on taxonomic completeness.","section":"Section I-E4; Fig. 2"},{"comment":"Table VI contains citation and attribution errors that undermine the survey's reliability as a reference. The rows 'Gao et al. (2024) [142]' and 'Liu et al. (2024) [111]' are used for works that the text attributes to references [180] and [181], and the reference number [142] is already used for Fontana et al. in Section IV.B.1. In addition, the text after Section IV.D points to 'Table VI' when the mobile-network summary is Table V. The reference list and table cross-references should be audited and corrected.","section":"Table VI; Section V"},{"comment":"Table IV mixes reported and author-estimated values without per-entry provenance. The asterisk note '[∗] Indicates estimated values' is insufficient because the reader cannot tell whether the 1.8T-token GPT-4 figure, the 15T-token LLaMA 3 figure, or the hardware entries come from the cited papers or from the authors' assumptions. Each estimated cell should be explicitly marked, and the table should state the source basis for the estimates.","section":"Section II.C; Table IV"}],"minor_comments":[{"comment":"The heading reads 'Network nonfiguration' and should be 'Network configuration.'","section":"Section V.B.2"},{"comment":"The phrase 'Excluded most preprints to ensure quality control' would benefit from a precise count of retained versus excluded preprints; the flow diagram currently reports only the combined removal of 49 articles.","section":"Section I-E4; Fig. 2"},{"comment":"Several rows (e.g., Dandoush et al., Tong et al., Shao et al., Rong et al.) list 'LLM Solution' and 'Dataset Used' as '-'; consider replacing hyphens with 'Not specified' for clarity.","section":"Table V"},{"comment":"The discussion of MonitorAssistant and the anomaly-detection systems notes lack of public data for some works, but the work in [198] is described as having 'no comprehensive evaluation' without a concrete statement of what validation was performed; consider adding a short critical assessment of that work's evidence.","section":"Section VI.A.2"},{"comment":"The Emergence pipeline (reference [122]) is described in detail, but no figure number is cited for Fig. 19; ensure all figures are referenced in the text.","section":"Section VI.C.1"}],"recommendation":"major_revision","confidential_remarks":"The firstness claim is the main risk; a quick search on arXiv (e.g., 'LLM network management survey' 2024) may surface competing surveys that the authors did not compare. The citation collisions in Table VI are indicative of a broader proofreading issue; I recommend the editorial office ask for a full reference audit. The paper is otherwise within the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This survey earns its place as a roadmap. It covers four network domains (mobile/IoT, vehicular, cloud, fog/edge) under one four-part taxonomy—monitoring/reporting, AI-powered planning, deployment/distribution, continuous support—and it documents 108 papers with tables that list contributions, datasets, and metrics. That is a real organizational contribution. The PRISMA flow and the explicit search strings make the selection process auditable, which is more than most surveys in this area bother to do. The LLM fundamentals section is competent and accessible.\n\nThe soft spots are real but not disqualifying. The “first extensive survey” claim in Section I-D is the weakest part. The paper compares itself to nine prior surveys, but its methodology does not include a systematic search for prior surveys or reviews, so the absence of a competing multi-domain survey is asserted rather than demonstrated. The authors can keep the claim if they soften it to “to the best of our knowledge after a systematic search for related surveys,” but as written it outruns the evidence. The preprint exclusion (“excluded most preprints to ensure quality control” in Section I-E4) is disclosed, but in a 2023–2024 area where influential work circulates as arXiv preprints, this biases the corpus toward published work and may miss task categories that only exist in the preprint literature. That is a limitation, not a fatal flaw, and it should be stated as one.\n\nThere are also citation and cross-reference errors—for example, Table VI attributes [142] to Gao et al. while the text cites [142] as Fontana et al., and [111] is likewise reused for two different works. These are fixable but need a careful pass. The survey also reports performance numbers from cited papers without critical validation, which is normal for a survey but means readers should treat those numbers as claims, not verified fact. The many self-citations from the same research groups are noticeable but not disqualifying; the authors are active in exactly this area.\n\nBottom line: this is a solid, useful survey for anyone entering LLM-for-NSM or looking for a structured map of the field. It deserves peer review and would benefit from a revision that tightens the firstness claim, addresses the preprint bias, and fixes the reference errors. I would cite it as a roadmap.","headline":"A genuinely useful, broad-scope survey of LLMs for network/service management with a workable taxonomy and a disclosed PRISMA method, but the firstness claim outruns its own search procedure and the preprint exclusion could skew the corpus.","tokens_in":742,"tokens_out":704,"would_cite":true,"duration_ms":23970,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 108 studies maps LLM use across four network management domains under one four-part taxonomy.","keywords":["large language models","network and service management","communication networks","mobile networks","vehicular networks","cloud networks","fog and edge computing","taxonomy"],"falsifier":"Rerun the survey's search with the same databases and date cutoff but add synonyms the authors did not use, such as 'foundation model,' 'generative AI,' 'network operations,' and 'self-healing network,' and count how many additional peer-reviewed studies satisfy their inclusion criteria; a count large relative to the 108 included studies would undercut the claim of an extensive and representative map. Alternatively, locate a peer-reviewed survey published before December 2024 that already covers LLM-enabled NSM across all four network domains; one such survey would falsify the 'first' claim directly.","tokens_in":43396,"feed_emoji":"📡","tokens_out":7035,"duration_ms":61674,"temperature":0.7,"pith_summary":"This paper is a survey, and its claim is about how the literature fits together. The authors set out to show that the scattered recent work on using large language models (LLMs) to manage communication networks forms a single landscape once it is seen through the right lens. They organize 108 selected studies under a four-part taxonomy — network monitoring and reporting, AI-powered network planning, network deployment and distribution, and continuous network support — and they map each category onto four network domains: mobile networks and IoT, vehicular networks, cloud-based networks, and fog/edge-based networks. The paper also argues that this is the first survey to cover LLM-enabled network and service management (NSM) across all four domains at once, rather than focusing on a single domain such as telecom, transportation, or edge intelligence. If that claim is right, the contribution is a shared map: researchers can see which LLM-assisted NSM tasks are already studied, which domains are thin, and where to look for open problems.","feed_headline":"First survey maps LLM use across four network domains","feed_subtitle":"A four-part taxonomy organizes 108 studies on mobile, vehicular, cloud, and edge network management.","key_machinery":"The organizing device is the taxonomy itself: a two-dimensional matrix whose rows are the four network domains (mobile/IoT, vehicular, cloud, fog/edge) and whose columns are the four NSM task families (monitoring and reporting, AI-powered planning, deployment and distribution, continuous support). The survey also relies on a systematic literature-selection process that starts with keyword searches, removes duplicates, screens titles, abstracts, and conclusions, excludes most preprints, and ends with 108 papers; this pipeline is what turns the taxonomy from an arbitrary scheme into a claimed mapping of the actual literature. The taxonomy does the argument's work: by placing each study in exactly one domain-task cell, it makes gaps, clusters, and the breadth of the 'first survey' claim visible.","core_discovery":"The central claim is that the body of research on LLMs for NSM can be comprehensively and usefully classified by crossing four network domains with four NSM task families. The paper states that it is, to its knowledge, the first extensive survey of LLM-enabled NSM spanning mobile networks and IoT technologies, vehicular networks, cloud-based networks, and fog/edge-based networks. Under the proposed taxonomy, each application is categorized as monitoring and reporting, AI-powered network planning, network deployment and distribution, or continuous network support; the survey then reviews the selected papers within each cell, notes their methods and experimental results, and derives cross-cutting challenges and future directions. The authors' intended contribution is not a new algorithm or empirical result but a structured synthesis that reveals the state of the field and provides a roadmap for LLM-driven NSM.","pith_inferences":["The tabulated evidence suggests an uneven distribution: most existing studies concentrate on monitoring and detection, while continuous support and fog/edge deployment are comparatively thin; readers looking for open problems should start there.","Because the search cutoff is September 2024 and most preprints were excluded, the map is a snapshot; a regularly updated, openly maintained version of the domain-task matrix would extend its shelf life.","The same four-task taxonomy could be tested against foundation models that are not purely text-based, such as multimodal or vision-language models, to see whether the NSM task boundaries still hold."],"forward_implications":["Researchers working in one network domain gain a ready index of relevant LLM techniques and evaluation choices used in the other three domains.","The taxonomy exposes underserved combinations, such as continuous network support in fog/edge settings, as concrete targets for new work.","Practitioners can use the four task families to scope LLM deployments by matching a management need, such as intent-based configuration, to studies that already attempted it.","The survey's fundamentals section gives a common baseline for comparing general-purpose and domain-specific LLMs in NSM contexts."],"supporting_citations":[{"why":"Prior survey of LLM applications in telecom that this work extends by adding a network and service management focus across multiple domains.","marker":"[38]"},{"why":"Prior survey of mobile edge intelligence for LLMs, representing the edge-focused single-domain perspective this survey distinguishes itself from.","marker":"[39]"},{"why":"Prior survey of LLMs in intelligent transportation systems, one of the single-domain surveys this work contrasts with its cross-domain scope.","marker":"[47]"},{"why":"General LLM survey covering fundamentals and applications outside communication NSM, used as source material for foundational LLM content.","marker":"[48]"},{"why":"Supplies the systematic literature-review methodology that shapes the survey's identification, screening, eligibility, and inclusion steps.","marker":"[50]"},{"why":"Example primary study of LLM-based intrusion detection in wireless networks that anchors the mobile-network monitoring category.","marker":"[134]"}],"fun_headline_variants":["First survey maps LLM use across four network domains","108 studies, four domains: LLM network management mapped","Survey: LLMs tackle mobile, vehicular, cloud, edge networks","LLM network management: first multi-domain survey unveiled","Four network domains, one taxonomy: LLMs surveyed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the keyword searches, abstract-level screening, and decision to exclude most preprints captured essentially all significant LLM-for-NSM research in the four domains; if important work was missed, the claimed first-mover status and the shape of the taxonomy could be misleading.","fun_headline_variants_meta":{"raw":{"variants":["First survey maps LLM use across four network domains","108 studies, four domains: LLM network management mapped","Survey: LLMs tackle mobile, vehicular, cloud, edge networks","LLM network management: first multi-domain survey unveiled","Four network domains, one taxonomy: LLMs surveyed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1341,"prompt_tokens":983,"completion_tokens":358,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":277}},"tokens_in":599,"tokens_out":358,"duration_ms":4214,"temperature":1.0,"reasoning_tokens":277,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:11:41.639350+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the survey's search with the same databases and date cutoff but add synonyms the authors did not use, such as 'foundation model,' 'generative AI,' 'network operations,' and 'self-healing network,' and count how many additional peer-reviewed studies satisfy their inclusion criteria; a count large relative to the 108 included studies would undercut the claim of an extensive and representative map. Alternatively, locate a peer-reviewed survey published before December 2024 that already covers LLM-enabled NSM across all four network domains; one such survey would falsify the 'first' claim directly.","supporting_citations":[],"review_version":1}