{"id":"f7677a29-ea48-487e-a11b-615e95070cde","arxiv_id":"2511.00422","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A PRISMA-based review of 83 papers classifies spatio-temporal model structures, finding hierarchical and additive models dominate while reproducibility is low.","lead":"This paper reviews 83 recent studies to map how statistical models for data that changes across space and time are built and used in five fields, and proposes a simple classification scheme. It finds hierarchical models and additive space-time components dominate, while economics relies on flat lag-based models, and it documents limited code sharing.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection filters likely skew field-level architecture claims; 'no hierarchical models in economics' rests on a small, keyword-restricted sample.","rationale":"The reader's weakest assumption marks sample representativeness as the load-bearing point, and I agree. The paper's central empirical contribution is a cross-domain synthesis of model structures; that synthesis stands or falls on whether the 83 papers fairly represent current practice in each field. The authors explicitly restrict by exact keywords, venue ranking, and a subjective QA process, and they offer no evidence that the keyword filter—which they say removes 'nearly twice as many records'—does not bias the sample toward certain traditions (e.g., disease mapping with BYM models) and away from others (e.g., Bayesian spatial econometrics). The most striking findings—zero hierarchical models in economics, only hierarchical models in criminology—are based on six papers each, so even a small selection bias could reverse the statement. The authors' claim in Section 2.2 that the restriction does not 'substantially limit coverage' is exactly the unsupported assertion that needs testing. The test I propose is a targeted re-run with relaxed keywords in those two fields; it is concrete, feasible, and would settle whether the concern lands. I did not find a more fundamental internal problem: the classification scheme is clearly defined, the PRISMA process is transparent, and the repository is promised. The word 'significantly' in Section 4.3.2 is used without a statistical test, but this is a presentational flaw rather than a central threat. The subjectivity of the QA criteria is real but secondary: it affects all papers and is less likely to create a systematic field-specific bias than the keyword filter. Therefore, the reader's CONDITIONAL verdict remains appropriate; my read does not change it.","tokens_in":32979,"tokens_out":6415,"duration_ms":64255,"concrete_test":"Perform a sensitivity analysis by re-running the Scopus query from Table 9 without the EXACTKEYWORD restriction (or with an expanded set of terms: 'spatio-temporal', 'spatial panel', 'Bayesian spatial', 'hierarchical spatio-temporal') for 2021–2025, limited to the same Q1/CORE-A venues, and apply the same title/abstract/full-text and QA screening. Then recompute the architecture counts in Tables 14 and 15. If any hierarchical models appear in economics, or any flat models in criminology, the claimed field preferences are artifacts of the selection criteria. A lighter alternative: if the authors archived the 88 records excluded by the EXACTKEYWORD filter, inspect whether any of those would meet the inclusion criteria and contain hierarchical models in these two fields.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's empirical core—field-wise frequencies of model architectures—depends on the representativeness of the 83 included papers. The sample is constrained by three sequential filters: (i) the exact Scopus EXACTKEYWORD restriction to 'Spatio-temporal Models'/'Spatiotemporal Analysis' (Section 2.2), (ii) the Q1/CORE-A venue requirement (Section 2.1), and (iii) the subjective QA criteria (Section 2.6). The authors assert that the keyword restriction does not 'substantially limit coverage' (Section 2.2) but provide no sensitivity analysis against the unrestricted search, which retrieved nearly twice as many records. With only six economics papers, the claim that 'no hierarchical models are used' (Section 4.3.2, Table 14) may be an artifact of the search string—which includes 'spatial panel data model*' but omits terms such as 'Bayesian spatial econometrics' or 'hierarchical Bayes'—rather than a true field difference. Similarly, the criminology result (all six models hierarchical, Table 15) rests on a tiny sample. Because the paper's headline findings (RQ1, RQ2) are exactly these cross-domain structural contrasts, this selection sensitivity is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a PRISMA-guided systematic review of 83 publications (2021–2025) that apply spatio-temporal statistical models in applied domains. Its main contributions are (i) a classification scheme distinguishing flat versus hierarchical model architectures and further decomposing model characteristics into additive components, lag structures, intensity functions, and other strategies, and (ii) a cross-domain synthesis of how these structures are used in epidemiology, ecology, public health, economics, criminology, and a small 'others' category. The central empirical claims are that hierarchical models are the most frequent architecture overall, that additive spatio-temporal components dominate, and that field-specific preferences exist—most strikingly, no hierarchical models are found in economics while all six criminology papers use hierarchical models. The authors also report limitations in data quality, model assumptions, and reproducibility, and note that code is absent for a majority of reviewed papers.","tokens_in":33217,"tokens_out":6750,"duration_ms":67898,"significance":"If the sample is representative, the review fills a real gap: existing surveys are either domain-specific or model-class-specific, whereas this paper provides an explicit, cross-disciplinary classification scheme and links it to application contexts. The methodology is transparent and reproducible in important respects: the PRISMA flow is documented, search strings are given, and the repository is public. The mathematical formalization of the classification scheme is a useful resource for practitioners seeking to identify or compare model structures. However, the significance of the frequency-based claims, including the headline field contrasts, rests entirely on the representativeness of the 83 included papers. Because the selection procedure involves strong filters and no sensitivity analysis, the descriptive conclusions are currently more fragile than the narrative suggests. If the authors address this by either adding sensitivity analyses or carefully re-scoping the claims, the paper would be a valuable contribution to interdisciplinary statistical practice.","major_comments":[{"comment":"The exact-keyword Scopus filter, combined with the Q1/CORE-A venue restriction, is load-bearing for the central descriptive claims of RQ1 and RQ2. The authors state that the unrestricted search retrieved nearly twice as many records but 'without substantially limiting the coverage of the research field' (Section 2.2); no sensitivity analysis supports this assertion. The claim that no hierarchical models are used in economics (Section 4.3.2, Table 14) rests on six papers, and the search string omits terms such as 'Bayesian spatial econometrics' or 'hierarchical Bayes', so the finding may be an artifact of the keyword filter rather than a true field difference. Similarly, the all-hierarchical result for criminology (Table 15) is based on six papers. Please provide a sensitivity analysis comparing the included sample with the unrestricted Scopus results, or explicitly rephrase the cross-dom","section":"§2.2, §4.3.2, Tables 9, 14, 15"},{"comment":"The quality assessment is an additional selection step that removes 128 of 211 full-text papers, yet the manuscript does not report how many reviewers performed the QA judgments, how disagreements were resolved, or any inter-rater reliability measure. Since QA1–QA7 are qualitative and applied sequentially, this introduces unquantified reviewer-dependent variation into the final sample. Given that the paper's frequency findings depend on the set of 83 papers, the QA step deserves the same transparency as the database search. Please report the number of screeners and the disagreement-resolution protocol, or provide a robustness check showing that the main field-level patterns are stable under plausible QA variations.","section":"§2.6, Figure 2"},{"comment":"The flat/hierarchical architecture distinction is central to the review, but the mapping from individual papers to architecture labels is not fully auditable from the manuscript. Some papers appear more than once in Table 10 with only tick marks and no on-the-row indication of which extracted model corresponds to which classification decision. The repository may contain the full extraction, but the manuscript itself should either present the underlying per-model assignments or provide a clearer key so a reader can verify, for example, why the two entries for [8] receive different classifications. This is a reproducibility concern for the main classification scheme.","section":"§4.2.1, Table 10"}],"minor_comments":[{"comment":"The notation for the latent process contains a typo: it should read Y := {Y_{t,s} | (t,s) ∈ D_s × D_t}, not D_t × D_t. The same error appears in the surrounding text.","section":"§4.2.2"},{"comment":"The PRISMA flow diagram appears mislabeled in the rendered version: the 'Records excluded (n = 88)' label is attached to the 'Abstract screening' box in the figure, whereas the text (Section 4.1) states that the 88 exclusions occur during journal/conference ranking. Please align the figure with the text.","section":"Figure 2"},{"comment":"The code-availability sentence is ambiguous: 'Code is available for only 34 publications. However, for five papers, it is only available upon request.' Clarify whether the 34 includes the 5 upon-request cases; as written, it may appear to sum to 88 rather than 83.","section":"§5.1"},{"comment":"The statement that existing classifications were 'used both of these to validate the groups found' is vague. Please describe concretely how the textbook taxonomies [26, 29, 109] were used to validate or revise the inductively derived categories.","section":"§2.4"},{"comment":"The table legend defines B and F but does not explain the meaning of the x marks or the columns. Add a note clarifying that x indicates presence of the given feature for that row's model.","section":"Table 10"},{"comment":"The text reports '60 hierarchical models and 26 flat models,' while the review includes 83 publications. Since some papers contribute multiple models, this is not necessarily an error, but please state explicitly that the counts are per model, not per publication, to avoid apparent inconsistency.","section":"§4.3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, transparent systematic review and the classification scheme is genuinely useful. The main barrier is that the empirical frequency claims—especially the cross-field contrasts—are presented with more confidence than the search and QA filters can support. A sensitivity analysis or a careful re-scoping of the conclusions would materially change my assessment. The topic fits Statistical Science well if the claims are made commensurate with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick read for you. This is a PRISMA systematic review of 83 applied spatio-temporal modeling papers from 2021–2025, with a proposed classification scheme (flat vs. hierarchical; additive, lag, intensity, other) and a field-wise tabulation. The genuinely new piece is the empirical cross-domain tabulation: hierarchical models dominate overall, economics papers in the sample use only flat lag models, criminology is all hierarchical. The classification itself is a reasonable synthesis of Cressie & Wikle, Elhorst, and Wikle et al., but the counts and the cross-field contrasts are the contribution.\n\nWhat it does well: the protocol is transparent, the PRISMA flow diagram is clear, the QA criteria are explicit, and the GitHub repository with coded data is a real plus. The citation graph and the reproducibility check (47 of 83 without code) are useful. It is honestly written, including a ChatGPT ethics note. That is all to the good.\n\nThe soft spot is representativeness, and it is load-bearing. The Scopus EXACTKEYWORD restriction plus the Q1/A venue filter plus the 2021–2025 window makes the 83 papers a keyword-defined slice, not a census of the literature. The authors assert the keyword restriction 'does not substantially limit coverage' but give no sensitivity analysis, and the unrestricted search returned nearly twice as many records. The claim that no economics paper uses hierarchical models rests on six papers, and the search string omits terms like 'Bayesian spatial econometrics' or 'hierarchical Bayes' that would plausibly capture hierarchical spatial econometric work. A single missed paper flips that cell. Criminology (all six hierarchical) is equally fragile. The paper also uses 'significantly' without a statistical test, and the QA criteria are subjective, though the documentation helps. These are addressable: add a sensitivity analysis or qualify the field-difference claims, and the review's descriptive value improves without changing its nature.\n\nSo: a solid, useful review of a specific slice of the applied spatio-temporal literature. For a reader wanting a compact classification scheme and a map of how a certain group of papers uses additive vs. lag structures, it has value. But the headline findings about field preferences should not be taken at face value until the selection sensitivity is addressed.\n\nMy recommendation: send it to peer review. A serious referee can tighten the claims and push for a sensitivity check. It is not a desk reject, but it should come back with the selection-bias discussion substantially strengthened. I would not cite the current version for the field-difference numbers; the classification scheme is workable enough to cite if you need a compact taxonomy.","headline":"Useful, transparent PRISMA review with a workable classification, but the field-difference claims rest on a keyword-restricted sample and need a sensitivity check before I'd trust them.","tokens_in":33740,"tokens_out":3266,"would_cite":false,"duration_ms":36498,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a systematic, cross-disciplinary review of spatio-temporal statistical models was missing, and proposes a classification scheme that reveals hierarchical models dominate current practice while modeling preferences diff","keywords":["spatio-temporal models","systematic review","classification scheme","hierarchical models","additive structures","lag structures","application domains","reproducibility"],"falsifier":"Re-run the review's search without the exact-keyword restriction or including journals outside Q1/CORE A, and count hierarchical vs flat models in economics. If hierarchical models appear in economics at a comparable rate to other fields, the paper's central distributional claim fails; if the pattern persists, it is robust. A second check: search for applied spatio-temporal papers in economics journals that mention hierarchical or Bayesian models to see whether the zero count is an artifact.","tokens_in":32819,"feed_emoji":"🗂️","tokens_out":3954,"duration_ms":41253,"temperature":0.7,"pith_summary":"The paper argues that no existing review combines a classification of spatio-temporal statistical model structures with a cross-domain survey of applications, and that this gap matters because researchers in one field rarely borrow models from another. To close it, the authors systematically reviewed 83 papers from 2021-2025 across five application domains and propose a two-level classification scheme: first flat vs hierarchical architecture, then additive, lag, intensity, and other structures. They find hierarchical models are the most common, additive structures are used in at least half of models in every field, and fields differ sharply—economics uses only flat lag-based models while criminology uses only hierarchical ones. They also report that Bayesian methods dominate outside economics and that code is publicly available for only a minority of papers. If true, the review gives applied researchers a shared vocabulary for comparing model structures and a map of where cross-field borrowing is most needed.","feed_headline":"Review of 83 spatio-temporal papers finds hierarchical models dominate","feed_subtitle":"Additive structures are the norm; economics favors flat lag models, criminology uses only hierarchical.","key_machinery":"The classification scheme itself is the central mechanism. It builds on the distinction between flat and hierarchical architectures—where hierarchical models factor the joint distribution into a data model and a process model—and then classifies the dependence structure into additive spatio-temporal components (u_s + v_t + gamma_{t,s}), lag structures (spatial weight matrices, temporal autoregressions, and their combinations), intensity functions for point processes, and a residual 'other' category. This scheme is what lets the authors compare structures across fields and produce the paper's distributional findings.","core_discovery":"The central discovery is a working taxonomy that organizes the model structures actually used in current spatio-temporal statistics into a small number of recurring types. At the top level, models split into flat architectures, which model the data process directly, and hierarchical architectures, which factor the joint distribution into a data model and a latent process model. Below that, the spatio-temporal dependence is captured either by additive components (spatial, temporal, and spatio-temporal effects such as CAR/BYM, random walks, Gaussian processes, and splines), by lag structures (spatial, temporal, or spatio-temporal lags of covariates, errors, or the target), by intensity functio","pith_inferences":["If search-filter sensitivity is real, the striking 'no hierarchical models in economics' result may reflect the sample of Q1 journals and exact keywords rather than actual economic practice; a replication that includes lower-ranked or methods-focused outlets could overturn this specific distribution.","The taxonomy could be applied to pre-2021 literature to test whether model preferences are stable or shifting, e.g., whether hierarchical models are growing outside health fields.","The paper's focus on interpretable statistical models suggests a natural extension: classifying hybrid statistical-deep-learning models under the same scheme to see whether they follow the same field-specific patterns.","The near-absence of cross-domain citations among the reviewed papers implies that a shared taxonomy might actually change citation behavior; a before/after bibliometric test is possible."],"forward_implications":["Applied statisticians gain a simple checklist—flat or hierarchical, additive or lag or intensity—for locating their model in the broader landscape.","The finding that economics uses flat lag models while criminology uses only hierarchical models suggests fields can learn from each other's toolkits, e.g., hierarchical structures for nested economic data.","The predominance of additive structures across fields implies that new methods for space-time interactions can be plugged into most existing model families.","Documentation that Bayesian methods dominate 56 of 83 papers outside economics points to continued demand for scalable Bayesian computation.","The low code-availability rate (34 of 83, some on request) indicates a concrete reproducibility gap that journals could close with data/code policies."],"fun_headline_variants":["Hierarchical models dominate 83 spatio-temporal studies","New taxonomy reveals common spatio-temporal model patterns","Review: additive and hierarchical structures dominate spatio-temporal stats","From epidemiology to criminology: spatio-temporal model trends","2021-25 review maps spatio-temporal model structures"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The review's quantitative findings rest entirely on the 83 papers that survived its search filters—two databases, exact keyword restrictions, Q1 journals and A-ranked conferences, 2021-2025—so if those filters systematically exclude relevant work (for instance, hierarchical models in economics or applied work in lower-ranked venues), the observed field differences could be artifacts of selection rather than real practices.","fun_headline_variants_meta":{"raw":{"variants":["Hierarchical models dominate 83 spatio-temporal studies","New taxonomy reveals common spatio-temporal model patterns","Review: additive and hierarchical structures dominate spatio-temporal stats","From epidemiology to criminology: spatio-temporal model trends","2021-25 review maps spatio-temporal model structures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2358,"prompt_tokens":728,"completion_tokens":1630,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":1562}},"tokens_in":472,"tokens_out":1630,"duration_ms":11280,"temperature":1.0,"reasoning_tokens":1562,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:31:39.281524+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the review's search without the exact-keyword restriction or including journals outside Q1/CORE A, and count hierarchical vs flat models in economics. If hierarchical models appear in economics at a comparable rate to other fields, the paper's central distributional claim fails; if the pattern persists, it is robust. A second check: search for applied spatio-temporal papers in economics journals that mention hierarchical or Bayesian models to see whether the zero count is an artifact.","supporting_citations":[],"review_version":1}