{"id":"d49919dd-b1da-44cc-b23f-9a32bab81e24","arxiv_id":"2505.07975","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Time-varying tensor VARs with one evolving CP loading and a conditional DIC plus knee point rule recover true configurations in simulations and reveal time-varying fMRI connectivity.","lead":"This paper introduces a time-varying version of tensor vector autoregression, where the coefficient matrix is a CP-decomposed tensor with one of its three loadings evolving over time. It shows that a conditional DIC with knee point detection selects the right model in simulations and finds time-varying brain connectivity in fMRI story-reading data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DIC_c,1 configuration selection is validated only when the true model lies in the one-time-varying-loading CP family; the fMRI 'time-varying dynamics' conclusion may be an artifact of that restricted class.","rationale":"The reader's weakest assumption is the one-time-varying-loading CP restriction, and that is also the concern I regard as most load-bearing. The paper's internal logic is coherent: the Gibbs sampler is standard for the stated state-space model, the DIC variants are correctly motivated by non-identifiability of the margins, and the simulation evidence supports the claim within the assumed model class. The gap is external validity: the abstract's 'accurately identifies true model configurations' is established only in the correctly specified case, while the fMRI conclusion requires the model class to be a reasonable approximation of real connectivity dynamics. The paper acknowledges the restriction in Section 2.2 and even notes that multiple time-varying loadings are possible but excluded, so the missing comparison is a stated limitation. A misspecification simulation of the kind described above would settle whether the selection rule and the empirical interpretation survive relaxation of the CP assumption. On that basis I do not think the reader's CONDITIONAL verdict should change; it already requires additional evidence, and the proposed test is a natural condition. I agree with the reader rather than adding a separate concern, because the structural restriction is the point where the central claim is least secure.","tokens_in":20193,"tokens_out":13198,"duration_ms":142111,"concrete_test":"Run a simulation with the same dimensions as Section 5.1 (N=3, P=3, T=200) but generate A_t from (i) a full TVP-VAR with independent random-walk coefficients and no CP constraint, and (ii) a CP model with two time-varying loadings (e.g., B_{t,1} and B_{t,2} both driven by random walks with Q=0.01). Apply the paper's DIC_c,1 + kneedle selection over the four configurations. Record (a) how often TVAR is selected when the true coefficients are time-varying, and (b) mean absolute error between the posterior mean A_t under the selected configuration and the true A_t. Frequent TVAR selection in case (i) would show the fMRI time-varying conclusion is not robust; poor A_t recovery in case (ii) would show the configuration label is unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DIC_c,1 with knee-point detection 'accurately identifies true model configurations' (Abstract; Section 5.2) and that the fMRI analysis supports time-varying brain connectivity (Sections 6.2-6.3). The supporting Monte Carlo study in Section 5 generates every dataset from exactly one of the four candidate classes: TVAR(3) or TVP-TVAR(3,j) with fixed rank 3 (Section 5.1). The selection problem is therefore one of choosing among correctly specified, non-nested CP models. In real data the coefficient tensor A_t is not guaranteed to have an exact CP decomposition with fixed rank and exactly one time-varying loading; Section 2.2 explicitly restricts to one time-varying loading for identifiability and computation and never fits or simulates models with two or three time-varying loadings, or a full TVP-VAR. Under misspecification, DIC_c,1 can select the best of several approximating models even when the true dynamics are qualitatively different, so the selected configuration label ('response loading varies') loses its interpretation. The empirical conclusion that dynamics are time-varying rests on comparing TVAR with TVP-TVAR(4,1) inside this restricted family; a full random-walk TVP-VAR or a CP model with multiple time-varying loadings could be better represented by TVP-TVAR(4,1) than by TVAR without the restricted model being true. Because the paper itself acknowledges that multiple time-varying loadings are possible (Section 2.2), this is a stated limitation rather than a hidden one, but it is still the load-bearing gap between the simulation evidence and the application claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a time-varying parameter tensor vector autoregression (TVP-TVAR) in which the VAR coefficient tensor has a CP decomposition and exactly one of the three loadings evolves as a random walk; the other two loadings remain time-invariant. Posterior inference is carried out with a Gibbs sampler using forward-filtering backward-sampling and closed-form full conditionals. The paper compares conditional and marginal DIC variants, recommends the conditional variant DIC_c,1, and uses the kneedle algorithm to select the CP rank. A Monte Carlo study with 3-variable, 200-observation data sets generated from TVAR(3) and TVP-TVAR(3,j), j=1,2,3, reports configuration recovery around 90% or better and improved rank selection with knee-point detection. The method is applied to 32 fMRI runs from the Wehbe et al. (2014) story-reading data, selecting TVP-TVAR(4,1) for most runs and reporting time-varying Granger causality networks. The paper claims over 90% reduction in parameters relative to standard VARs and argues that the empirical results support time-varying brain connectivity.","tokens_in":20474,"tokens_out":5277,"duration_ms":54649,"significance":"If the selection and rank procedures are reliable, the paper offers a practical way to fit high-dimensional TVP-VARs with strong parameter reduction and interpretable dynamic connectivity. The paper is clear about the state-space representation and provides explicit full conditionals, making the sampler implementable. The Monte Carlo study uses an external synthetic-data benchmark with known true configuration and rank, and the confusion matrix in Table 1 is encouraging. The knee-point idea addresses a real overfitting tendency of DIC for rank selection. The main significance is conditional on the restricted one-time-varying-loading CP family; the evidence does not yet establish that the selected configuration is meaningful under misspecification.","major_comments":[{"comment":"The configuration-selection evidence is confined to data generated from exactly the four model classes considered, with rank fixed at 3. Section 2.2 explicitly states that multiple time-varying loadings are possible but are not modeled. The fMRI conclusion that brain connectivity dynamics are time-varying rests on comparing TVAR with three one-time-varying-loading CP models; a data-generating process with two or three time-varying loadings, or an unrestricted TVP-VAR, could be best approximated by TVP-TVAR(4,1) without that model being true. To support the abstract's claim that DIC_c,1 'accurately identifies true model configurations' outside the correctly specified family, the authors should add Monte Carlo scenarios with multiple time-varying loadings and with a full TVP-VAR, and report configuration selection and DIC gaps for those scenarios.","section":"Section 5 (Table 1) and Section 2.2"},{"comment":"The claim that DIC_c,1 has lower Monte Carlo error than DIC_c,2 and DIC_m is demonstrated only for data generated from TVP-TVAR(3,1) with rank 3, and the histograms in panels (b) and (c) exclude several data sets by truncation. The text says the truncated panels 'account for 93 data sets,' but the behavior of the remaining 7 data sets is not described. Please provide Monte Carlo error comparisons for all four true configurations and a range of ranks, or explicitly qualify the reliability claim to the single scenario examined.","section":"Section 5.2 (Figure 2)"},{"comment":"Knee-point detection is used as the primary rank-selection device, but no distributional or consistency justification is given. The method assumes the DIC curve is monotone decreasing and selects the point of maximum curvature; the paper does not report how often this assumption fails or how sensitive the selected rank is to the normalization step. In Figure 5a, even with knee-point detection, fewer than 60 of 100 data sets recover the true rank 3, with most errors at ranks 2 and 4. The authors should report rank-selection accuracy numerically for all configurations and compare the knee-point rule with alternatives such as marginal-likelihood-based selection or a formal penalized criterion.","section":"Section 4 and Figure 5"},{"comment":"The fMRI narrative-alignment conclusion is based on one subject-run data set, with Granger causality thresholds delta=0.01 and p*=99.9% set without sensitivity analysis. The temporal pattern in Figure 7 and the network interpretation in Figure 8 could change materially under small threshold perturbations, so a robustness check across thresholds and across the 32 data sets is needed before the narrative-progression claim can be accepted as a general empirical finding.","section":"Section 6.3 (Figures 7-8)"}],"minor_comments":[{"comment":"The third margin in the CP decomposition is written as beta_3 twice; it should be beta_1 outer beta_2 outer beta_3.","section":"Equation (2.3)"},{"comment":"The last two column headers are both labeled TVP-TVAR(3,2); the final column should read TVP-TVAR(3,3).","section":"Table 1"},{"comment":"The caption says the middle and right panels restrict display to maxima of 350 and 500 and account for 93 data sets, but the number of excluded data sets and their Monte Carlo errors are not stated; please clarify whether the excluded values are extreme outliers or truncation artifacts.","section":"Figure 2 caption"},{"comment":"The entry for TVP-TVAR(4,3) is '/', but the configuration is fitted in the selection step; please explain why no parameter count is reported, and state how the averaged counts are computed across data sets of different length.","section":"Table 3"},{"comment":"The text refers to 'Nievell' and 'Precental gyrus'; these should be 'Neville' and 'Precentral gyrus'. Also, the threshold p*=99.9% should be motivated or cited.","section":"Section 6.3 and Appendix D"},{"comment":"No code or data availability statement is included; releasing the sampler and simulation scripts would strengthen reproducibility, which is particularly valuable given the high computational cost of the MCMC implementation.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is publishable in principle, but the empirical claim about time-varying fMRI dynamics currently depends on the one-time-varying-loading CP family. The expanded simulations requested in Major Comment 1 are essential before acceptance. The relationship to Zhang et al. (2021), who also fit time-varying tensor VARs to the same story-reading data, should be sharpened to clarify the incremental contribution of the DIC-based configuration and rank selection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a credible, workmanlike extension of tensor VARs to time-varying coefficients. The genuinely new pieces are the three interpretable configurations (exactly one CP loading time-varying) and the systematic comparison of DIC variants with knee-point rank selection. The state-space samplers are standard but correctly derived, and the simulation confirms that DIC_c,1 recovers the true configuration in the family considered. That is real value.\n\nWhat the paper does well: it carefully motivates why conditional DIC beats marginal DIC here — the CP sign/permutation indeterminacy inflates the variance of margin-based plug-ins — and the trace plots support that claim. The knee-point detection is a sensible fix to DIC's tendency to overfit rank. The fMRI application is a nice demonstration with interpretable Granger connectivity patterns.\n\nSoft spots, in proportion: the main one is that all validation is inside the one-time-varying-loading CP family. The Monte Carlo study only generates from the four candidate classes, so it never tests misspecification — e.g., two or three time-varying loadings, or a full TVP-VAR. Section 2.2 explicitly restricts to one time-varying loading for identifiability and computation, and the paper acknowledges that multiple time-varying loadings are possible, so this is a stated limitation, not a hidden one. But it means the fMRI conclusion that 'dynamics are time-varying' is really a statement about the best approximation within a restricted family. That should be softened.\n\nOther soft spots: the simulation is limited to N=3, T=200 with a single innovation setup; no code is provided (a reproducibility problem for a methods paper); the DIC and Granger quantities in the application have no uncertainty quantification; the abstract says 'over 90% reduction relative to standard VARs' when the comparison in Table 3 is actually against the full TVP-VAR, not the time-invariant VAR; and Table 1 has a typo (the fourth column header repeats TVP-TVAR(3,2) instead of (3,3)). None of these are fatal.\n\nRecommendation: send to peer review. The core method is sound and the DIC comparison is useful, but ask for robustness checks under misspecification, code, and a more careful statement of the empirical claim. I'd bring it to a reading group as a good example of a Bayesian tensor model selection study.","headline":"A credible, workmanlike extension of tensor VARs to time-varying coefficients; the DIC selection story holds within its stated model family and the paper deserves peer review with requests for robustness checks and code.","tokens_in":21093,"tokens_out":3757,"would_cite":true,"duration_ms":31104,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"A plug-in conditional DIC reliably selects the configuration and rank of a time-varying tensor VAR.","keywords":["time-varying parameter vector autoregression","tensor decomposition","CP decomposition","Deviance information criterion","knee point detection","Granger causality","fMRI","Bayesian inference"],"falsifier":"Simulate 100 data sets from a TVP-TVAR in which two loadings (say response and predictor) both follow random walks, then run the paper's four-configuration comparison: if $DIC_{c,1}$ selects one of the restricted configurations with high confidence rather than flagging misspecification, the criterion's configuration choice is not reliable for multi-loading dynamics. A cheaper check is to fit a model allowing all three loadings to vary on the fMRI data and see whether the time-varying connectivity pattern or the parameter reduction survives.","tokens_in":19927,"feed_emoji":"🧠","tokens_out":8311,"duration_ms":70161,"temperature":0.7,"pith_summary":"Time-varying parameter vector autoregressions become hard to fit in high dimensions because every coefficient may change at every time point. This paper compresses the time-varying coefficient matrix into a third-order tensor with a CP decomposition in which exactly one of three loadings (response, predictor, or temporal) evolves as a random walk, giving three model configurations. It then asks how to choose the configuration and the decomposition rank from MCMC output, and claims that one conditional DIC variant, $DIC_{c,1}$, which plugs in the posterior mean of the coefficient tensor, is substantially more reliable than conditional or marginal DICs based on latent margins: it has lower Monte Carlo error and recovers the true configuration in simulations. Applied to fMRI story-reading data, the chosen models cut parameter counts by over 90% and suggest that brain connectivity dynamics are time-varying.","feed_headline":"A plug-in DIC selects the true time-varying tensor VAR","feed_subtitle":"It scores the coefficient tensor instead of latent loadings, cutting Monte Carlo error and recovering the true rank.","key_machinery":"The central object is the third-order coefficient tensor $A_t \\in \\mathbb{R}^{N\\times N\\times P}$ in the VAR $\\boldsymbol{y}_t = A_{t,(1)}\\boldsymbol{x}_t + \\boldsymbol{\\epsilon}_t$, subject to a CP decomposition in which one of the three factor families is time-varying, for example $A_t = \\sum_{r=1}^R \\boldsymbol{\\beta}^{(r)}_{1,t} \\circ \\boldsymbol{\\beta}^{(r)}_2 \\circ \\boldsymbol{\\beta}^{(r)}_3$. The vectorized time-varying loading follows a random walk, giving a state-space model sampled by a Gibbs sampler with forward-filtering backward-sampling. Model selection is carried by $DIC_{c,1}$, which plugs in the posterior mean of $A_t$ itself, and by the 'kneedle' knee-point detector applied to the sequence of $DIC_{c,1}$ values across ranks. The machinery works because the composite tensor is identified even though its margins are not, so the plug-in deviance is computed from a well-mixing, identifiable quantity.","core_discovery":"The paper's central claim is that a specific conditional DIC, $DIC_{c,1}$, solves model selection for TVP-TVARs by scoring the identified composite tensor rather than the weakly identified loading factors. Because CP loadings are identifiable only up to sign switching and permutation, their MCMC chains mix poorly, and DIC variants that plug in loading means inherit large Monte Carlo error; $DIC_{c,1}$ avoids this by plugging in the posterior mean of the coefficient tensor $A_t$ itself. Simulations over 100 data sets per configuration show that $DIC_{c,1}$ selects the true configuration in nearly all cases and, combined with knee point detection on the DIC-versus-rank curve, recovers the true rank instead of defaulting to the maximum. The paper therefore recommends $DIC_{c,1}$ with the 'kneedle' algorithm for configuration and rank choice. On 32 fMRI data sets from story reading, the selected model is most often TVP-TVAR(4,1), meaning a time-varying response loading, with parameter reductions above 90%, and Granger causality counts rise and fall with the narrative.","pith_inferences":["The same logic suggests a general principle: in any Bayesian latent-variable model where a low-rank or compressed parameter is identified while its factors are not, model selection criteria should be evaluated on the identified composite parameter rather than on the factors themselves.","A natural stress test is to generate data from a TVP-TVAR with two or three loadings varying jointly; if $DIC_{c,1}$ then confidently prefers one of the three restricted configurations, the fMRI conclusion that connectivity dynamics are time-varying may be an artifact of the restricted model family.","The knee-point procedure's performance may depend on the chosen maximum rank $R^*$ and on the number of ranks evaluated; a systematic sensitivity check would clarify when the improvement is robust.","The Granger causality interpretation could be validated out-of-sample by checking whether the narrative-linked rise and fall in connectivity counts reproduces across subjects or across chapters."],"forward_implications":["Practitioners can select both the configuration and the rank of a TVP-TVAR from a modest number of MCMC runs, using $DIC_{c,1}$ plus knee point detection rather than fitting the full model space exhaustively.","The number of estimated parameters scales as $(2N+P)R$ instead of $N^2P$, so high-dimensional VARs become feasible; the fMRI application reports over 90% parameter reduction.","The finding that the response loading is time-varying for most subjects implies that Granger-causal connectivity patterns are time-dependent and that static VAR estimates would miss narrative-linked structure.","Knee point detection should be used rather than minimum DIC when rank is the target, because DIC underpenalizes overfitted tensor ranks.","Conditional DICs based on identified composite quantities can outperform marginal DICs in state-space models whose latent components are only weakly identified."],"supporting_citations":[{"why":"Defines the conditional and marginal DIC variants for latent-variable models that the paper compares.","marker":"Celeux et al. (2006)"},{"why":"Introduces the DIC and effective number of parameters that the paper adopts as its model selection metric.","marker":"Spiegelhalter et al. (2002)"},{"why":"Supplies the 'kneedle' algorithm used to detect the knee point in the DIC-versus-rank curve.","marker":"Satopaa et al. (2011)"},{"why":"Documents the uncertainty and complexity-bias concerns about conditional DICs that the paper argues do not apply to $DIC_{c,1}$.","marker":"Chan and Grant (2016b)"},{"why":"Provides evidence that DIC separates underfitted but not overfitted models, motivating knee point detection for rank choice.","marker":"Maity et al. (2021)"},{"why":"Establishes the random-walk evolution of coefficients in TVP-VARs that the paper uses for the time-varying loading.","marker":"Primiceri (2005)"},{"why":"Introduced the tensor VAR with CP decomposition whose coefficient representation the paper extends to the time-varying setting.","marker":"Wang et al. (2022)"},{"why":"Gives the interpretation of response, predictor, and temporal loadings that defines the three TVP-TVAR configurations.","marker":"Luo and Griffin (2025)"},{"why":"Provides a Bayesian time-varying tensor VAR for fMRI effective connectivity and the lag order and ROI choice used in the application.","marker":"Zhang et al. (2021)"},{"why":"Contributed the story-reading fMRI data set used in the empirical analysis.","marker":"Wehbe et al. (2014)"}],"fun_headline_variants":["A plug-in DIC scores the tensor, not loadings, for TVP-VARs","Conditional DIC beats marginal DIC for time-varying tensor VARs","Knee-point DIC selects true rank in TVP tensor VARs","Time-varying tensor VARs: 90% fewer parameters via CP decomposition","DIC on coefficient tensor recovers true TVP-TVAR configuration"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the true data-generating process has a fixed-rank CP tensor form with exactly one time-varying loading; if real dynamics involve several loadings changing together, all four compared configurations are misspecified and the selected one is only the best of a restricted family.","fun_headline_variants_meta":{"raw":{"variants":["A plug-in DIC scores the tensor, not loadings, for TVP-VARs","Conditional DIC beats marginal DIC for time-varying tensor VARs","Knee-point DIC selects true rank in TVP tensor VARs","Time-varying tensor VARs: 90% fewer parameters via CP decomposition","DIC on coefficient tensor recovers true TVP-TVAR configuration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00059,"raw_usage":{"total_tokens":2791,"prompt_tokens":993,"completion_tokens":1798,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":1698}},"tokens_in":609,"tokens_out":1798,"duration_ms":12931,"temperature":1.0,"reasoning_tokens":1698,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:08:06.184813+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate 100 data sets from a TVP-TVAR in which two loadings (say response and predictor) both follow random walks, then run the paper's four-configuration comparison: if $DIC_{c,1}$ selects one of the restricted configurations with high confidence rather than flagging misspecification, the criterion's configuration choice is not reliable for multi-loading dynamics. A cheaper check is to fit a model allowing all three loadings to vary on the fMRI data and see whether the time-varying connectivity pattern or the parameter reduction survives.","supporting_citations":[{"cited_title":"P., and Titterington, D","cited_arxiv_id":null,"evidence_quote":"Defines the conditional and marginal DIC variants for latent-variable models that the paper compares."},{"cited_title":"kneedle","cited_arxiv_id":null,"evidence_quote":"Supplies the 'kneedle' algorithm used to detect the knee point in the DIC-versus-rank curve."},{"cited_title":"K., Basu, S., and Ghosh, S","cited_arxiv_id":null,"evidence_quote":"Provides evidence that DIC separates underfitted but not overfitted models, motivating knee point detection for rank choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the random-walk evolution of coefficients in TVP-VARs that the paper uses for the time-varying loading."},{"cited_title":"and Griffin, J","cited_arxiv_id":null,"evidence_quote":"Gives the interpretation of response, predictor, and temporal loadings that defines the three TVP-TVAR configurations."},{"cited_title":"Bayesian Time-Varying Tensor Vector Autoregressive Models for Dynamic Effective Connectivity","cited_arxiv_id":"2106.14083","evidence_quote":"Provides a Bayesian time-varying tensor VAR for fMRI effective connectivity and the lag order and ROI choice used in the application."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributed the story-reading fMRI data set used in the empirical analysis."}],"review_version":1}