{"id":"a8cc34d0-7335-4c0f-af68-8913ceaacadf","arxiv_id":"2506.22895","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Sparse autoregression with exact mixed-integer optimization and scalable temporal/spatial extensions quantifies dominant periodic lags in large mobility and climate time series.","lead":"This paper builds sparse autoregression models with exact mixed-integer optimization to isolate dominant periodic lags in time series. It adds scalable extensions for temporally and spatially varying data and tests them on city mobility and climate records.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The two-stage STV-SAR global-support approximation is the load-bearing point: if the Eq. (13) support is wrong for many locations, every spatial seasonality map in §7.2–7.3 is structurally biased, and the paper's own Fig. 8 check does not close this gap.","rationale":"The reader's weakest-assumption analysis correctly identifies the two-stage STV-SAR approximation as the most load-bearing point. The central claim of scalable, interpretable periodicity quantification on spatiotemporal data depends on the global support set Ω being representative of local dynamics. If Ω omits a lag that is dominant in some region, Eq. (15) sets that coefficient to zero and the resulting seasonality map misrepresents that region's periodic structure. The paper's own conclusion admits this bias, and the only validation (Fig. 8) covers one variable and two decades without comparing coefficient magnitudes or checking other datasets. I do not see a more fundamental flaw: the MIO formulations in §4–5 are standard and the TV-SAR experiments are suggestive, though the El Niño language in the abstract overstates what §7.3 shows. The reader's CONDITIONAL verdict is appropriate; my analysis does not require changing it, but it sharpens the condition: the authors should provide per-grid support and coefficient comparisons or explicitly qualify that the STV-SAR maps are conditional on the approximate global support.","tokens_in":17604,"tokens_out":2797,"duration_ms":36481,"concrete_test":"Re-run the two-stage STV-SAR on one decade of the 10 km minimum temperature data, then independently solve per-grid sparse AR (using MIO or exact NNSP with the same τ) on a random subset of 5,000 grids. Compare (i) the fraction of grids whose independent support differs from the global Ω, and (ii) the correlation and mean absolute error between the per-grid lag-12 coefficients obtained from the independent fits versus the two-stage fits. If more than about 5% of grids select lags outside Ω, or if the lag-12 coefficient maps differ materially at those grids, the two-stage approximation is not reliable for the seasonality maps. A complementary check is to reproduce Fig. 8 for maximum temperature, precipitation, and SST for the 2010s; if those histograms do not peak sharply at {1, 11, 12}, the global support generalization is unsupported outside the one variable tested.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The scalability claim for STV-SAR rests on the two-stage scheme in §6.2–6.3. Eq. (13) solves for a single global coefficient vector w from the aggregated Gram matrix P and response q; its support Ω is then fixed as the support for every local model in Eq. (15). This is valid only if supp(argmin_w Σ_{m,n,γ} ||x̃_{m,n,γ} − A_{m,n,γ}w||²) matches the support that would be chosen by jointly solving Eq. (9), or at least coincides with locally optimal supports at the vast majority of grids. The paper's own conclusion concedes that the two-stage approximation 'may lead to estimation bias.' The only empirical support (§7.2.2, Fig. 8) examines support frequencies on minimum temperature for two decades, but it does not quantify whether the fixed global Ω distorts the per-grid seasonality coefficients from which all maps in §7.2.3 and §7.3.2 are built, and no analogous check is reported for maximum temperature, precipitation, or sea surface temperature. Because Eq. (15) forces every coefficient outside Ω to zero, a wrong global support would not merely add noise; it would structurally remove the true periodicity from entire regions, making the headline spatial patterns uninterpretable as periodicity maps.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a sparse autoregression (SAR) framework with ℓ0-norm sparsity and non-negativity constraints, solved exactly via mixed-integer optimization (MIO). It extends SAR to temporally-varying settings (TV-SAR) with shared support across segments and a decision-variable-pruning acceleration, and to spatiotemporal settings (STV-SAR) with a two-stage scheme that first learns a global support set and then fits per-grid coefficients. Experiments on NYC ridesharing data and on North American/global climate data report interpretable daily, weekly, and yearly periodicities, including shifts attributed to COVID-19 and decadal climate variability. The central claim is that exact sparse AR isolates dominant periodic lags and scales to large real-world spatiotemporal datasets.","tokens_in":17996,"tokens_out":4697,"duration_ms":55938,"significance":"If the claims are fully supported, the framework is a useful interpretable alternative to dense AR and heuristic sparse methods: the MIO formulations are mathematically sound, the selected lags {1,24,168} and {1,11,12} match domain expectations, and the reported scale of millions of time series is impressive. The paper also demonstrates concretely that greedy NNSP can return spurious lags that MIO avoids. The main limitations are the unvalidated two-stage global-support approximation and the absence of uncertainty quantification, which currently weaken the quantitative periodicity claims.","major_comments":[{"comment":"The two-stage global-support approximation is load-bearing for every spatiotemporal map in §7.2–7.3. Eq. (15) forces every per-grid coefficient outside Ω to zero, so an incorrect global support would systematically remove true periodic lags from entire regions rather than merely adding noise. The only validation, Fig. 8, reports aggregate support frequencies for minimum temperature in two decades, but it does not quantify whether the fixed Ω distorts per-grid coefficient estimates, and no analogous check is given for maximum temperature, precipitation, or SST. The conclusion itself concedes that the two-stage scheme 'may lead to estimation bias.' Please add a diagnostic on a sample of grids for each variable/decade comparing the two-stage estimates against local unrestricted or per-grid MIO supports—for example, the fraction of grids where the global support misses the local top lags and the distribution of coefficient differences. Without such a check, the spatial seasonality maps cannot be interpreted as reliable periodicity maps.","section":"§6.2–6.3, Eq. (13)–(15)"},{"comment":"The empirical periodicity claims are made without uncertainty quantification. The text asserts, for example, that the weekly component 'decreases remarkably' in 2020, that 2019 and 2021–2023 show 'no significant differences,' and that seasonality intensifies in the 2000s, but no standard errors, confidence intervals, or hypothesis tests are reported. The phrase 'no significant differences' in §7.1.3 is therefore unsupported. Please add block-bootstrap or equivalent confidence bands for the reported coefficients and use interval comparisons or tests for the stability/change statements. This is essential because the paper's contribution is periodicity quantification, not merely support-set recovery.","section":"§7.1.3, §7.2.3, §7.3.2"},{"comment":"The paper claims that MIO gives 'more accurate and reliable periodicity quantification than conventional greedy methods,' but the only baseline is NNSP. There is no comparison with standard periodicity and seasonality tools such as the autocorrelation function, periodogram/spectral analysis, or seasonal-trend decomposition, and no quantitative evaluation on synthetic series with known periods. As a result, the central claim that the selected lags are the 'dominant periodicities' is validated only qualitatively. Please add a benchmark (synthetic ground truth and, ideally, a standard spectral method) or temper the claims about superiority and reliability.","section":"§4.3 and §7"}],"minor_comments":[{"comment":"The constraint in Eq. (13) is written as '0 ≤ w ≤ M · z, ∀m, n, γ,' but w is a single global vector; the quantifier should be over the coordinate k ∈ [d].","section":"Eq. (13)"},{"comment":"The caption contains a typo ('corresponds to to one month') and the phrase 'accumulated from January to the end of the given month' is unclear; it should specify exactly how the monthly time segments are constructed from the original series.","section":"Figure 6 caption"},{"comment":"Algorithm 1 relies on the NNSP method from [7], but the paper does not provide a self-contained description or pseudocode for NNSP, which is needed to reproduce the DVP experiments.","section":"Algorithm 1"},{"comment":"Figure 8 reports support-set frequencies for independent SAR models but does not state which solver was used (NNSP, MIO-DVP, or MIO) or provide any uncertainty quantification; please clarify.","section":"Figure 8"},{"comment":"The claim that MIO-DVP 'can be viewed as a backbone-type algorithm' is reasonable, but the sure-screening justification is not established; no screening property (e.g., that the true support is contained in S̃ with high probability) is proven or empirically tested.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious referee. What is genuinely new is the two-stage scheme for STV-SAR and the DVP screening, plus the demonstration that these scale to millions of climate time series. The SAR formulation itself is best subset selection on an AR design matrix, and TV-SAR is explicitly a special case of the authors' own sparsely-varying regression; they say so in Remark 1, so there is no attempt to hide it. The empirical work is also honest: the selected lags (1, 24, 168 for NYC mobility; 1, 11, 12 for monthly climate) line up with what domain knowledge would predict, and the MIO solutions beat the greedy baseline cleanly.\n\nThe load-bearing soft spot is the one the authors themselves concede in the conclusion: the two-stage STV-SAR approximation 'may lead to estimation bias.' The global support from Eq. (13) is fixed for every local model in Eq. (15), which means a wrong global support can zero out a region's true periodicity. The Fig. 8 check shows that the global support matches the most frequently selected lags on minimum temperature, but that does not close the gap: it does not tell you whether the coefficient maps in Figs. 9 and 14 are distorted where local dynamics differ. The same check is not reported for precipitation, maximum temperature, or SST. This is not a fatal flaw—the approximation is disclosed, and the shared-support assumption is plausible for climate seasonality—but the paper would be stronger with a per-grid comparison, at least on a subsample, between the two-stage estimates and directly fitted local models.\n\nTwo smaller issues. There is no uncertainty quantification anywhere: no error bars, no significance tests, and no sensitivity of the selected lags to noise beyond the phase-segmentation experiment. For a paper whose claims are about 'quantifying' periodicity, that is a real omission, though it is a minor one given the descriptive rather than inferential framing. Also, no code or data are released, which matters here because the runtime claims (3.3 million series in under a minute) are central to the scalability argument and cannot be independently reproduced.\n\nWho is this for? Researchers working on interpretable time series models, especially in climate or mobility applications, who want a principled way to select sparse lags at scale. The methods section is clean and the experiments are suggestive. A serious referee should engage, and the main requests should be uncertainty quantification, a direct validation of the two-stage approximation, and code release. I would not cite this for the novelty of SAR as a concept, but I would cite it for the STV-SAR scaling result if I were working on large spatiotemporal panels.","headline":"A useful, honest methods paper: exact sparse AR via MIO is not new, but the scaling tricks and the climate-scale maps are, and the two-stage STV-SAR approximation is the soft spot.","tokens_in":18460,"tokens_out":702,"would_cite":true,"duration_ms":10296,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","90C11"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that exact sparse autoregression via mixed-integer optimization extracts the dominant periodic lags from time series and, in temporally and spatially varying forms, scales to millions of real-world mobility and climate…","keywords":["sparse autoregression","mixed-integer optimization","periodicity quantification","temporally-varying autoregression","spatiotemporally-varying autoregression","seasonality detection","interpretable machine learning","time series analysis"],"falsifier":"Synthesize a spatial panel in which half the grid cells have a yearly cycle and half are pure noise, run the two-stage STV-SAR pipeline, and compare its global support and per-cell maps against independent per-cell sparse autoregressions; if the aperiodic cells receive nonzero yearly-seasonality coefficients or the yearly lag is missing from the global support, the two-stage approximation is wrong.","tokens_in":17447,"feed_emoji":"📈","tokens_out":9835,"duration_ms":95917,"temperature":0.7,"pith_summary":"This paper tries to establish that periodicity in real time series can be read off directly from a sparse autoregressive model: place an $\\ell_0$ constraint that allows at most $\\tau$ nonzero lags, solve the resulting combinatorial problem exactly with mixed-integer optimization, and the selected lags and coefficients become the quantitative description of the cycles. The authors show that the plain version, SAR, recovers the expected daily and weekly lags on hourly mobility data where a greedy baseline inserts a spurious lag. They then extend it in time (TV-SAR), enforcing one shared support set across segments, and in space and time (STV-SAR), pooling millions of locations into a global support estimate. On NYC ridesharing data the temporally varying model tracks the pandemic weakening of the weekly rhythm and the strengthening of the daily rhythm; on North American and global climate data the spatiotemporal model maps where yearly seasonality is strong and how those maps move across decades. If the approach holds, it gives one interpretable tool for quantifying when and where cyclical structure changes, without assuming a fixed seasonal decomposition.","feed_headline":"Exact sparse AR picks out the lags that carry real cycles","feed_subtitle":"Mixed-integer optimization finds dominant lags, revealing COVID mobility shifts and climate seasonality change.","key_machinery":"The load-bearing object is the binary support vector $z \\in \\{0,1\\}^d$ coupled to the coefficient vector by $0 \\le w \\le Mz$, so that $w_k$ can be nonzero only when $z_k=1$, together with the global budget $\\sum_k z_k \\le \\tau$. This turns the $\\ell_0$ sparsity constraint into a mixed-integer quadratic program that branch-and-bound solvers can solve to proven optimality, which is what distinguishes the method from greedy subspace pursuit. For scalability, the decision variable pruning strategy first screens candidate lags with a fast greedy pass and then solves the reduced MIO, while the two-stage STV-SAR scheme pools all spatial locations into a single global MIO via the trace identity $\\sum \\|\\tilde{x} - Aw\\|_2^2 = \\operatorname{tr}(ww^\\top P) - 2w^\\top q + C$ and then solves per-location quadratic programs on the shared support. The selected support set is the model's output: the list of lags that carry the periodicity.","core_discovery":"The central claim is that exact sparse autoregression, rather than a dense least-squares fit or a greedy approximation, is the right formulation for periodicity quantification. Concretely, the paper solves $\\min_{w,z} \\|\\tilde{x} - Aw\\|_2^2$ subject to $0 \\le w \\le Mz$, $z \\in \\{0,1\\}^d$, and $\\sum_{k=1}^d z_k \\le \\tau$, where the binary vector $z$ encodes which lags are active and $A$ is the lagged design matrix. A large positive coefficient at lag $k$ is read as periodicity with period $k$, because the model must use that past value to explain the present. The non-stationary extension TV-SAR partitions the series into segments and requires all segments to share the same support set, so that a change in a coefficient means a change in cycle strength, not a change in which cycles exist. The spatiotemporal extension STV-SAR first estimates one global support set from a pooled fit and then estimates per-location coefficients on that support, which is what makes the model tractable for millions of grid cells. The empirical claim is that these models reproduce known cycles, expose structural breaks such as pandemic mobility shifts and Arctic sea surface temperature changes, and flag equatorial Pacific dynamics associated with El Niño.","pith_inferences":["Beyond the paper's descriptive use, the same selected-lag output could serve as a feature-selection front end for forecasting: if the support is stable across segments, a forecaster could restrict attention to those lags and estimate predictive coefficients without the dense noise.","The authors concede that the global-support shortcut can bias estimates when local dynamics differ; a natural test is to allow region-specific supports through the paper's sparsely-varying relaxation with a small support-difference budget and compare the resulting seasonality maps on heterogeneous climate zones.","The method should transfer to other strongly periodic systems, such as energy load, network traffic, or epidemiological counts, where the quantity of interest is precisely which lags carry the cycle and how their weights shift.","A higher-resolution mobility experiment could test whether the daily lag splits into sub-daily lags when sampling is finer, extending the same support-recovery logic to intra-day periodicities."],"forward_implications":["On hourly NYC ridesharing data with order $d=168$ and sparsity $\\tau=4$, TV-SAR selects lags $\\{1,24,167,168\\}$, identifying local autocorrelation plus daily and weekly cycles; the exact MIO solution achieves a lower objective than non-negative subspace pursuit, which inserts a spurious lag at $k=53$.","The monthly TV-SAR coefficients show the weekly lag weakening and the daily and local lags strengthening during 2020, then returning to pre-pandemic levels by 2021–2023.","On the North American Daymet data, STV-SAR maps yearly seasonality with support $\\{1,11,12\\}$, strong high-latitude signals, an intensification in the 2000s, and consistent patterns at 5, 10, and 20 km resolutions.","On global sea surface temperatures, the lag-12 coefficient maps show strong seasonality near mid-latitude continents, weaker seasonality in equatorial El Niño zones, and declining Arctic seasonality over recent decades.","The two-stage STV-SAR pipeline fits over 3.3 million time series at 5 km resolution in under a minute, demonstrating practical scalability."],"supporting_citations":[{"why":"It supplies the classical autoregressive model that SAR modifies by adding sparsity constraints.","marker":"[4]"},{"why":"It provides the autoregression background and notation used throughout the formulations.","marker":"[5]"},{"why":"It contributes the mixed-integer optimization machinery for sparse regression with support-set consistency that TV-SAR and STV-SAR build on.","marker":"[6]"},{"why":"It provides the non-negative subspace pursuit algorithm used as the greedy baseline and inside the DVP pruning step.","marker":"[7]"},{"why":"It is the origin of subspace pursuit, the screening algorithm that DVP uses to reduce the candidate lag set before MIO.","marker":"[25]"},{"why":"It established that best-subset selection via mixed-integer optimization is practically solvable, the foundation for exact SAR.","marker":"[27]"},{"why":"It is the Daymet climate dataset on which the North American seasonality maps are estimated.","marker":"[34]"}],"fun_headline_variants":["Exact sparse AR pins down dominant periodic lags","Sparse AR with MIO reveals cycle shifts at scale","Integer-optimized sparse AR quantifies periodicity","Sparse autoregression: exact lags for periodic patterns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that one global list of lags, learned by pooling all locations into a single autoregressive fit, is the correct support list for every location; if local cycle structures differ, the two-stage shortcut used for the climate maps can distort the spatial patterns, a bias the authors acknowledge.","fun_headline_variants_meta":{"raw":{"variants":["Exact sparse AR pins down dominant periodic lags","Sparse AR with MIO reveals cycle shifts at scale","Integer-optimized sparse AR quantifies periodicity","Sparse autoregression: exact lags for periodic patterns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00044,"raw_usage":{"total_tokens":2275,"prompt_tokens":1034,"completion_tokens":1241,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":1176}},"tokens_in":650,"tokens_out":1241,"duration_ms":64663,"temperature":1.0,"reasoning_tokens":1176,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:55:39.381267+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Synthesize a spatial panel in which half the grid cells have a yearly cycle and half are pure noise, run the two-stage STV-SAR pipeline, and compare its global support and per-cell maps against independent per-cell sparse autoregressions; if the aperiodic cells receive nonzero yearly-seasonality coefficients or the yearly lag is missing from the global support, the two-stage approximation is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the autoregression background and notation used throughout the formulations."},{"cited_title":"Slowly varying regression under sparsity,","cited_arxiv_id":null,"evidence_quote":"It contributes the mixed-integer optimization machinery for sparse regression with support-set consistency that TV-SAR and STV-SAR build on."},{"cited_title":"Correlating time series with interpretable convolutional kernels,","cited_arxiv_id":null,"evidence_quote":"It provides the non-negative subspace pursuit algorithm used as the greedy baseline and inside the DVP pruning step."},{"cited_title":"Subspace pursuit for compressive sensing signal reconstruction,","cited_arxiv_id":null,"evidence_quote":"It is the origin of subspace pursuit, the screening algorithm that DVP uses to reduce the candidate lag set before MIO."},{"cited_title":"Best subset selection via a modern optimization lens,","cited_arxiv_id":null,"evidence_quote":"It established that best-subset selection via mixed-integer optimization is practically solvable, the foundation for exact SAR."},{"cited_title":"Daymet: Monthly climate summaries on a 1-km grid for north america, version 4 r1,","cited_arxiv_id":null,"evidence_quote":"It is the Daymet climate dataset on which the North American seasonality maps are estimated."}],"review_version":1}