{"id":"d87c7405-42dc-4cc7-a7db-89f2c6e89143","arxiv_id":"2501.06932","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review and taxonomy of large language model applications for natural disaster management, with a public dataset catalog and a list of research challenges.","lead":"This paper surveys how large language models are used across the four phases of disaster management: mitigation, preparedness, response, and recovery. It introduces a taxonomy of applications, compiles public datasets, and lists open challenges for researchers and emergency managers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No search methodology plus demonstrable mis-categorizations in Table 1 undermine the central claim of a comprehensive, accurate survey.","rationale":"The central claim fails only if Table 1 contains a meaningful proportion of false or misattributed entries, which is exactly what the audit would determine.","tokens_in":27021,"tokens_out":3773,"duration_ms":37895,"concrete_test":"Independently audit every entry in Table 1: for each cited reference, retrieve the actual title and abstract and check (a) whether it addresses a disaster-management task, (b) whether it uses an LLM, and (c) whether the assigned Phase, Application, Task, and Architecture match the cited content. Quantify the false-positive and misclassification rates. If more than 5% of entries fail any of these checks, the taxonomy cannot be considered a reliable map, and the paper requires a documented re-selection protocol and corrected table before the central claim is accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is to provide the \"first systematical review\" and a taxonomy that gives a reliable map of LLM applications in disaster management. For this to hold, Table 1's entries must accurately reflect the cited works and the inclusion process must be sound. Neither condition is currently met. The Introduction and Limitations contain no search strategy, inclusion/exclusion criteria, database sources, or date range, so the comprehensiveness claim cannot be audited. More concretely, Table 1 lists Conneau (2019), the XLM-RoBERTa model paper, as a disaster-management Need Classification study with a novel method, and Section 3.3.3 cites the authors' own Lei et al. (2025) and Lei et al. (2022), which are spatial-temporal forecasting and bot-detection papers, not disaster tweet classification. These are not harmless typos: they indicate that the categorization pipeline can assign non-disaster, non-task-specific papers into the taxonomy, which directly threatens the accuracy of the phase/task distribution in Figure 1 and the utility of the proposed taxonomy. Additional mechanical defects, including a duplicated heading in Section 3.3.3 and a stray inserted passage near Figure 2, further reduce confidence that the entries were verified against their sources. The claim is not internally inconsistent, but it is unsupported until the survey corpus and Table 1 are audited.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys applications of large language models across the four disaster-management phases of mitigation, preparedness, response, and recovery. It proposes a taxonomy organized by application scenario, specific task, and model architecture (encoder-based, decoder-based, and multimodal), compiles a list of publicly available datasets, and discusses challenges and future directions. The central claims are that this is the first systematic review of LLM applications in disaster management and that the proposed taxonomy provides a reliable map of the field.","tokens_in":27243,"tokens_out":5238,"duration_ms":47405,"significance":"If its contents are accurate, the survey is a useful resource for the disaster-informatics and NLP communities: it organizes a scattered literature, offers a reasonable task/architecture taxonomy, and collects many relevant datasets in one place. The authors are also candid in the Limitations section about the focus on natural disasters and the subset of datasets included. However, the value of the survey rests on the correctness of its corpus and categorizations, and the concrete reliability problems detailed below currently prevent me from treating the map as trustworthy. The paper does not provide code or machine-checked artifacts, but the detailed appendix tables and the explicit limitations statement are positive features.","major_comments":[{"comment":"The Introduction claims to provide the 'first systematical review' of LLM applications in disaster management, but the manuscript gives no search strategy, databases, query terms, date range, or inclusion/exclusion criteria. Without an auditable methodology, the comprehensiveness claim cannot be verified and the selection of surveyed papers is unanchored. Please add a methodology subsection describing how papers were retrieved and screened, or explicitly temper the systematic/comprehensive claim.","section":"Section 1 and Limitations"},{"comment":"Table 1 lists Conneau (2019), the XLM-RoBERTa model paper, as a Response-phase Need Classification study with a novel method. In Section 3.3.3 the citation is used only as the embedding model in a cosine-similarity retrieval approach, not as a disaster-management paper. This is a demonstrable mis-categorization: it means Table 1 is not a reliable inventory of the surveyed literature, and it also affects the phase/task distribution reported in Figure 1.","section":"Table 1, Conneau (2019) row"},{"comment":"The sentences 'augmentation strategies such as manual hashtag annotation ... and self-training with soft labeling ... are employed to enhance classification performance (Lei et al., 2025)' and 'Multimodal LLMs can integrate rich data from social media ... (Lei et al., 2022)' cite Lei et al. (2025), which is the ST-FIT spatial-temporal forecasting paper, and Lei et al. (2022), which is the BIC Twitter bot-detection paper. Neither paper addresses disaster tweet classification or disaster information coordination. These citations should be removed or replaced with the intended disaster-related works; as written, they indicate that source verification is lacking.","section":"Section 3.3.3, Lei et al. citations"}],"minor_comments":[{"comment":"The heading 'Relevance Classification with Encoder-based LLMs' appears twice; the second occurrence should read '(Encoder-)Decoder LLMs' to match the text that follows.","section":"Section 3.3.1"},{"comment":"The sentence 'encoder-based LLMs have been fine-tuned to for binary classification' contains a stray 'to for' and should be 'fine-tuned for binary classification.'","section":"Section 3.3.2"},{"comment":"The phrase 'encoder-base LLMs' should be 'encoder-based LLMs.'","section":"Section 3.3.3"},{"comment":"The example tweets from CrisisLexT6 appear as body text immediately before Figure 2; if they are part of the figure, they should be placed inside it, and if they are intended as text, they need proper formatting and a citation.","section":"Figure 2"},{"comment":"CrisisFACTS is labeled with Application 'DIC' (Disaster Issue Consultation), but it is a summarization/timeline dataset and should be categorized under Disaster Information Coordination; please verify all Application labels in Table 2.","section":"Table 2"},{"comment":"The 'Did You Feel It' dataset entry is cited as '(Mousavi et al., 2024)', though the dataset originates from Atkinson and Wald (2007); cite the original source and use the 'Used in' column for Mousavi et al.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The two self-citations in Section 3.3.3 that point to unrelated papers are a particular concern; I do not impute intent, but the authors should be asked to verify every self-citation carefully. The absence of a methodology section is the most serious obstacle to the 'systematic review' claim, and the Table 1 mis-categorizations mean the paper needs a full audit of its survey corpus before it can be relied upon."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper's real contribution is organizational: a four-phase taxonomy (mitigation, preparedness, response, recovery) crossed with application scenarios, tasks, and model architectures, plus a dataset appendix. That is genuinely useful for a young, scattered subfield. The limitations section is refreshingly candid about scope—natural disasters only, open datasets only. I would use this as a starting map.\n\nThe soft spots are real and mostly in the machinery. The biggest is the absence of any search or inclusion protocol. \"First systematical review\" is a strong claim, but there is no database list, no date range, no inclusion/exclusion criteria, so coverage cannot be audited. The stress-test note is right about Table 1: Conneau (2019) is the XLM-RoBERTa model paper, not a disaster study, and it has no business in a need-classification row. The two self-citations in Section 3.3.3 (Lei et al. 2025, 2022) point to forecasting and bot-detection papers, not disaster tweet classification. Those look like citation errors, but they matter because they show the categorization pipeline can admit papers that do not fit the taxonomy. The duplicated heading in 3.3.1 and the odd formatting around Figure 2 are minor but consistent with the same story: the table was not verified entry by entry.\n\nThe taxonomy itself does not collapse. Most of the entries look right, and the phase/task distribution, while approximate, matches the literature's obvious response-phase bias. The challenges section is standard-issue but not wrong.\n\nWho is this for? Disaster informatics researchers and NLP groups looking for a map and dataset pointers. It deserves a serious referee, but not acceptance as-is. My recommendation: send it to review with a request for a documented search methodology and a full re-audit of Table 1 and the inline citations. Fix those and the survey becomes citable; as it stands, I would be reluctant to cite its counts.","headline":"Useful taxonomy and dataset catalog, but the 'systematic review' claim needs a documented search protocol and a table audit before I would trust the counts.","tokens_in":27777,"tokens_out":2897,"would_cite":false,"duration_ms":29549,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps large language models onto every phase of disaster management—mitigation, preparedness, response, and recovery—and distills the field's open problems.","keywords":["large language models","disaster management","taxonomy","disaster phases","social media analysis","multimodal LLMs","disaster datasets","survey"],"falsifier":"Run a reproducible literature search with specified databases, query terms for each disaster phase and LLM architecture, and explicit inclusion rules, then compare the retrieved corpus with the papers in Table 1; if the search finds a material cluster of LLM-disaster studies the survey omits, or if a nontrivial share of Table 1 entries turn out not to be about disaster management (for example, a multilingual language model paper listed as a novel disaster method), the comprehensiveness and accuracy claims fail.","tokens_in":26827,"feed_emoji":"🌪️","tokens_out":10114,"duration_ms":87741,"temperature":0.7,"pith_summary":"The paper sets out to give the first systematic overview of how large language models are used in natural disaster management, and to make that overview useful rather than merely enumerative. It organizes the surveyed studies into a taxonomy that crosses the four disaster phases with application scenarios, task types, and model architectures, and it compiles the public datasets those studies rely on. The survey then names the obstacles that currently block deployment: datasets skewed toward classification, efficiency limits for real-time use, hallucination in generated warnings and plans, and a lack of shared evaluation standards. If the map is accurate, it gives researchers and emergency managers a way to see which tasks are well served, which phases are neglected, and where new work would count most.","feed_headline":"Large language models now span all four phases of disaster management","feed_subtitle":"The first systematic review maps dozens of studies by disaster phase, task, and architecture, then lists open problems.","key_machinery":"The load-bearing structure is the taxonomy that maps each paper to a cell in a multi-axis grid: disaster phase, application scenario, task type, and model architecture. The four task types—classification, estimation, extraction, and generation—are the common downstream operations, and the three architecture families are encoder-based LLMs (for example BERT), decoder-based LLMs (for example GPT), and multimodal LLMs. This grid does the argument's work: it turns a long reference list into a coverage map from which the authors read both the field's concentration in response-phase classification and the gaps in mitigation, preparedness, and recovery.","core_discovery":"On its own terms, the paper's central claim is that a coherent field exists and that it has a recognizable shape: most LLM work in disasters concentrates on the response phase, especially classifying social media posts, while mitigation, preparedness, and recovery are comparatively sparse. The authors propose a taxonomy that assigns every surveyed study to a cell defined by disaster phase, application scenario (such as vulnerability assessment, disaster prediction, evacuation planning, or damage assessment), a task type among classification, estimation, extraction, and generation, and an architecture among encoder-based, decoder-based, and multimodal models. They further provide a table of publicly available datasets organized by the same task types, and they enumerate four open problems: biased and classification-heavy datasets, efficiency for real-time use, hallucination in generated content, and the absence of a unified evaluation protocol.","pith_inferences":["The same phase-task-architecture grid would likely transfer to adjacent crisis domains such as man-made disasters, public-health emergencies, or climate adaptation, making the taxonomy more general than the natural-disaster scope it is presented under.","Because relevance classification dominates the literature, benchmark saturation is plausible; a testable prediction is that fine-tuned encoder models will continue to match or beat prompt-based LLMs on those benchmarks, pushing the field's value toward multimodal and generation tasks.","The hallucination risk the paper identifies points to a missing shared resource: a safety-oriented evaluation set in which generated evacuation routes, warnings, and recovery plans are checked against authoritative ground truth."],"forward_implications":["A researcher can use the taxonomy to locate the sparse cells—estimation and generation outside the response phase—and target new work where the survey shows little competition.","Practitioners can treat the dataset tables as a starting point for building or evaluating disaster LLMs, with the caveat that most listed resources are text-only and classification-oriented.","The paper's four challenge areas (data quality, efficiency, hallucination, and unified evaluation) define a concrete agenda; if the survey's map is accurate, progress in disaster LLMs will come from fixing these rather than from inventing new architectures.","The survey's own statistics show that most studies apply existing models and cluster in the response phase, so the near-term bottleneck is task formulation and data rather than model design."],"supporting_citations":[{"why":"Supplies the four-phase disaster management framework (mitigation, preparedness, response, recovery) that organizes the entire taxonomy.","marker":"(Sun et al., 2020)"},{"why":"Defines BERT, the canonical encoder-based LLM whose fine-tuning underlies most classification and extraction studies in the survey.","marker":"(Devlin, 2018)"},{"why":"Defines GPT, the canonical decoder-based LLM used for generation, question answering, and prompting in the surveyed work.","marker":"(Brown, 2020)"},{"why":"Provides CrisisLexT6, the relevance-classification dataset used to illustrate the disaster-response task in Figure 2.","marker":"(Olteanu et al., 2014)"},{"why":"Provides CrisisBench, the consolidated benchmark that anchors the survey's classification-dataset list.","marker":"(Alam et al., 2021b)"},{"why":"Provides CrisisNLP, a large human-annotated corpus defining humanitarian categories used across the reviewed classification studies.","marker":"(Imran et al., 2016)"},{"why":"Provides CrisisMMD, the multimodal text-image dataset that grounds the survey's account of multimodal LLM classification.","marker":"(Alam et al., 2018)"},{"why":"Provides DisasterQA, the question-answering benchmark used to illustrate generation tasks and evaluation challenges.","marker":"(Rawat, 2024)"},{"why":"Provides CrisisFACTS, the multi-stream summarization dataset that supports the survey's generation-task and evaluation sections.","marker":"(McCreadie and Buntain, 2023)"}],"fun_headline_variants":["LLM disaster research skewed to response, survey shows","New taxonomy maps LLM use across disaster phases","Disaster LLMs: response phase dominates, survey finds","Four open problems limit LLM disaster management use","Systematic survey reveals LLM gaps in disaster phases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions stand on the assumption that its selection of papers is representative and its phase, task, and architecture labels are accurate, yet the paper describes no search strategy or inclusion criteria, leaving the taxonomy vulnerable to arbitrary or mistaken assignments.","fun_headline_variants_meta":{"raw":{"variants":["LLM disaster research skewed to response, survey shows","New taxonomy maps LLM use across disaster phases","Disaster LLMs: response phase dominates, survey finds","Four open problems limit LLM disaster management use","Systematic survey reveals LLM gaps in disaster phases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1394,"prompt_tokens":820,"completion_tokens":574,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":499}},"tokens_in":436,"tokens_out":574,"duration_ms":5814,"temperature":1.0,"reasoning_tokens":499,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:49:02.918279+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a reproducible literature search with specified databases, query terms for each disaster phase and LLM architecture, and explicit inclusion rules, then compare the retrieved corpus with the papers in Table 1; if the search finds a material cluster of LLM-disaster studies the survey omits, or if a nontrivial share of Table 1 entries turn out not to be about disaster management (for example, a multilingual language model paper listed as a novel disaster method), the comprehensiveness and accuracy claims fail.","supporting_citations":[],"review_version":1}