{"id":"224ddae0-7484-498d-a1b0-15d595a38ab0","arxiv_id":"2607.18209","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ATLAS is a multi-environment factor model estimator that separates invariant (shared-loading) factors from environment-specific factors and uses auxiliary labels to align and transfer prediction-relevant signals.","lead":"The paper introduces ATLAS, a method that separates latent factors shared across environments from environment-specific factors, and uses auxiliary labels to align the predictive factors for transfer. It provides non-asymptotic error bounds and shows that invariant-only prediction is worst-case optimal without auxiliary labels.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 1(a) is unfalsifiable and silently violated by common shared artifacts; ATLAS will classify such artifacts as invariant, so the transfer guarantees apply only when this untestable condition happens to hold.","rationale":"The reader's weakest-assumption analysis and my own stress-test converge on Assumption 1(a): the maximum invariant subspace condition is the load-bearing, unfalsifiable premise on which all recovery and transfer guarantees rest. The paper itself proves this assumption is instance-level necessary (Theorem B.1), which rules out any purely data-driven way to verify it. The most plausible real-world failure mode is a shared heterogeneous loading direction across environments — exactly the scenario the reader describes. Because the theoretical claims are explicitly conditional on Assumption 1, this concern does not invalidate the mathematics; it does, however, mean the abstract's unconditional-sounding claim that ATLAS 'unveils invariant and transferable latent factors' is only as strong as an assumption that may be false in the motivating EHR applications. The reader's CONDITIONAL verdict already captures this. I do not see a separate internal inconsistency in the proofs as presented, and the paper deserves credit for a genuine identification theorem and detailed finite-sample rates. The proposed simulation would make the failure mode concrete and quantify how far the transfer error can deviate from the stated rate when the assumption is violated. Since the reader already identified this as the weakest assumption and conditioned the verdict accordingly, no change in the verdict is needed.","tokens_in":47794,"tokens_out":5527,"duration_ms":58591,"concrete_test":"Simulate |E|=2 environments with r_I=2 and r_H=2, choosing A^(1) and A^(2) to share one identical column a ∉ col(B), so Assumption 1(a) fails while the rest of the model is regular. Give that shared heterogeneous factor zero predictive effect on Y in the training environments but a large environment-dependent effect in held-out test environments. Run ATLAS with the true ranks r^(e), r_I supplied, so no rank-selection ambiguity. Record (i) the number of components IHD declares invariant, (ii) the subspace distance between col(B) and the estimated invariant loading space, and (iii) the transfer L2 error of ATLAS(wo/Z) relative to the oracle invariant-only predictor. If IHD returns r_I+1 invariant directions and the transfer error exceeds the bound implied by Theorem 4.6 (computed as if Assumption 1 held), the concern is confirmed: the method cannot detect or correct a silent violation of i","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central identification claim — and hence every downstream rate in Theorem 4.3 and Theorem 4.6 — rests on Assumption 1(a): ∩_e col([B, A^(e)]) = col(B). This condition is not just strong; it is unfalsifiable from observational data, as the paper itself acknowledges (Section 2.1) and proves as instance-level necessary in Theorem B.1. If two or more environments share a heterogeneous loading direction a ∉ col(B) — e.g., the same coding artifact, laboratory convention, or documentation pattern across hospitals — then the true intersection is col([B, a]) and IHD will output a as an 'invariant' factor. The recovered F_I then contains a spurious component that is not invariant in the causal/invariance sense used to define the uncertainty set in Proposition 2.2. If that component has any environment-dependent association with Y in unseen environments, the transferred predictor carries non-transportable signal, and the oracle-style guarantee in Theorem 4.6 no longer applies. The paper provides no diagnostic or sensitivity analysis to detect this failure. This is a scope limitation rather than an internal inconsistency: conditional on Assumption 1 holding, the proofs appear coherent. But the practical claim that ATLAS 'uncovers invariant factors' is only as strong as this unverifiable premise, and the paper gives no way to know when the premise is false.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a multi-environment linear factor model in which each environment's covariates are driven by invariant factors with shared loadings and environment-specific heterogeneous factors. The authors show that, under a maximum-invariant-subspace condition plus block-uncorrelatedness, the invariant and heterogeneous factors are identifiable up to invertible transformations, and they prove instance-level necessity of the identification condition. They then propose ATLAS, a three-stage estimator: (i) an invariance-heterogeneity decomposition (IHD) that estimates shared and environment-specific loading subspaces and constructs diversified projections for the factors; (ii) a spectral step using auxiliary labels Z to select and align prediction-invariant heterogeneous factors; and (iii) a pooled GLM fit in labeled environments to estimate the invariant prediction rule and transfer it to new environments, with or without auxiliary labels. The theoretical core supplies non-asymptotic, dimension-free-in-direction sub-Gaussian error bounds for factor recovery, for aligning prediction-invariant factors, and for the transferred prediction error. Simulations and a temporal EHR application on rheumatoid arthritis are used to illustrate the method.","tokens_in":48131,"tokens_out":7968,"duration_ms":81658,"significance":"If the results hold as stated, the paper makes a useful contribution to multi-environment factor analysis and transfer learning. The identification theory is carefully developed: the maximum-invariant-subspace assumption is not merely imposed but shown to be instance-level necessary, and the use of auxiliary labels to go beyond invariant-factor-only prediction is a genuine extension of invariant causal prediction ideas to latent factor models. The non-asymptotic bounds are detailed and plausible, and the dimension-free sub-Gaussian control is a valuable technical contribution. The real-data demonstration, while limited, is appropriate for the motivating EHR setting. The main value is in providing a statistically rigorous method with explicit rates for a problem that is usually treated heuristically or under stronger distributional assumptions.","major_comments":[{"comment":"The entire identification and all downstream rate guarantees are conditional on the intersection condition ∩_e col([B,A^(e)]) = col(B) (or its quantitative version, Condition 4.2). The manuscript correctly acknowledges that this assumption is unfalsifiable from observational data and proves instance-level necessity in Theorem B.1. However, the practical claims in the title, abstract, and real-data section go beyond this conditional statement. If two or more environments share a heterogeneous loading direction a∉col(B) — a realistic scenario with shared coding conventions, laboratory normalizations, or documentation artifacts — then Condition 4.2 fails with ϵ_A=0, and IHD will include a in the estimated invariant subspace. The recovered 'invariant' factor may have environment-dependent association with Y in unseen environments, and the transfer guarantee in Theorem 4.6 no longer applies.","section":"Section 2.1 (Assumption 1(a)); Theorems 4.3 and 4.6"},{"comment":"The abstract and introduction describe the non-asymptotic bounds as 'sharp.' The only lower-bound argument in the paper is for the single-environment benchmark in Lemma 4.1, where the λ^{-1/2} noise floor is identified. For the multi-environment results — Theorem 4.3 for invariant/heterogeneous factor recovery, Theorem 4.5 for prediction-invariant factor alignment, and Theorem 4.6 for the transferred prediction error — no minimax lower bounds are provided. Terms such as sqrt(r^(e)/n_x) in δ_FI and the first-order terms in δ_Z are asserted as tight, but no matching lower-bound analysis is offered. The 'sharp' claim is therefore unsupported as stated. Please either provide lower bounds for these multi-environment problems or replace 'sharp' with 'near-oracle'/'non-asymptotic' and specify precisely which rates are known to be optimal.","section":"Abstract; Section 1.2; Theorems 4.3–4.6"}],"minor_comments":[{"comment":"In the caption and surrounding text, 'ATALS' appears to be a typo for 'ATLAS.'","section":"Section 5.1, Figure 2(d) caption"},{"comment":"The theoretical results require choosing λ_ihd in an interval [C δ_WI, ϵ_A − C δ_WI] that depends on unknown population quantities, but the algorithm takes λ_ihd as an input without a data-driven selection rule. The same issue applies to λ_sel in Section 3.2. In the real-data experiment, λ_IHD is fixed at 0.01 and r^(e)=64 in Appendix E.2; a brief sensitivity analysis for these choices would strengthen the practical claims.","section":"Section 3.1, Algorithm 1 and Section 4"},{"comment":"The display '∥Q^{-⊤} bβ − β*∥_2 / eC2' appears to have a typographical/subscript formatting issue; the intended statement is that the norm is bounded by eC2 times δ_y. Please correct the notation.","section":"Theorem 4.6, Eq. following (4.11)"},{"comment":"The paragraph correctly assumes F_I ⊥⊥ F_H for the nonparametric no-Z extension. It may be worth stating explicitly that the derivation uses independence, not merely the block-uncorrelatedness used elsewhere, to justify E[g_H(F_SH)|F_I] = E[g_H(F_SH)].","section":"Section 6, 'No auxiliary labels Z'"}],"recommendation":"major_revision","confidential_remarks":"The technical core of the paper is sound conditional on Assumption 1, and the authors are transparent about its unfalsifiability and instance-level necessity. The main risks are (i) overclaiming practical identification of 'invariant' factors without a sensitivity analysis or diagnostic for violations of Assumption 1(a), and (ii) the unsupported use of 'sharp' for rates that lack matching lower bounds. These are fixable within the manuscript's scope by adding a sensitivity analysis and either supplying lower bounds or moderating the language. The paper is not ready for acceptance in its current form, but rejection would be disproportionate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The takeaway: this is a solid, carefully worked-out paper that deserves referee time, but it is not the breakthrough its abstract hints at. The genuinely new contribution is using auxiliary labels to align prediction-invariant heterogeneous factors — going beyond shared-loading-only methods like AJIVE and Personalized PCA — combined with a partialling-out step that handles non-orthogonal loadings. The non-asymptotic bounds are detailed and plausible, and the isometric (direction-wise) sub-Gaussian error control is a real technical achievement. I checked the identification argument: it uses covariance geometry and GLM exogeneity, so the reliance on the authors' earlier invariant-subspace work is background, not circular.\\n\\nThe main soft spot is the load-bearing assumption, stated in Eq. (1.4) and Assumption 1(a): the intersection of the loading spaces across environments must equal the invariant loading space exactly. The paper correctly admits this is unfalsifiable, and the stress-test concern is real: if two sites share a coding artifact or documentation convention, ATLAS will label it an invariant factor, and the 'invariant' predictor will carry non-transportable signal. The paper offers no diagnostic or sensitivity analysis for this. That limits the practical claims but does not break the internal mathematics. A second, smaller issue: the abstract calls the bounds 'sharp,' but the paper supplies multi-environment lower bounds only for the single-environment benchmark, so 'sharp' overstates the case. The real-data section reports AUCs without error bars, does not specify how ranks were selected, and no code or data are provided — acceptable for a stats paper but it weakens the empirical claim.\\n\\nAll said, the identification theory is careful, the algorithm is new, and the rates make sense. I'd send this to a serious referee — one who will push on the unfalsifiability point and ask for diagnostics, and who will check the real-data uncertainty. For anyone working on multi-environment factor models or transfer learning, this is worth reading and citing. It won't reorganize the field, but it is a legitimate, helpful step.","headline":"Solid multi-environment factor-model paper with a genuinely new auxiliary-label alignment step; the main catch is the central invariant-subspace assumption is unfalsifiable and the real-data analysis is thin.","tokens_in":48600,"tokens_out":3195,"would_cite":true,"duration_ms":29521,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62J12","62H12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Invariant latent factors in multi-environment data are identifiable from unlabeled covariates alone, and auxiliary labels then make the full latent signal transportable at near-oracle error.","keywords":["factor model","invariance","heterogeneous environments","multi-environment data","transfer learning","latent factor regression","auxiliary labels","diversified projection"],"falsifier":"Simulate two or more environments with a true invariant loading B and environment-specific loadings A^(e), then add one extra loading vector shared by all environments (violating Assumption 1(a)) and run ATLAS: the recovered invariant space will include the extra direction, and the transferred predictor's worst-case out-of-sample risk will visibly exceed the oracle built on the true B. A cheaper check: compute the average of the per-environment top-subspace projectors and inspect the eigenvalue just after the r_I-th; if it is not separated from 1 (ϵ_A near zero in Condition 4.2), the identific","tokens_in":47663,"feed_emoji":"📐","tokens_out":7537,"duration_ms":72073,"temperature":0.7,"pith_summary":"The paper tries to establish a clean separation: when data are collected from several environments whose covariate distributions differ, the latent structure splits into invariant factors (shared loadings across environments) and heterogeneous factors (environment-specific loadings), and this split is recoverable from unlabeled covariates alone under one geometric condition — the intersection of the per-environment loading spaces must be exactly the shared loading space. Given that split, ATLAS uses auxiliary labels — cheap, noisy proxies observed in a subset of environments — to find which of the heterogeneous factors predict the response in the same way everywhere, and to align them across environments. The result is a prediction rule that transfers: with auxiliary labels in the new environment, the method transports the full invariant signal; without them, it falls back to the invariant factors only, which the paper proves is the unique worst-case-optimal choice. The paper backs this with non-asymptotic guarantees: recovered factors carry dimension-free sub-Gaussian errors, and the transported predictor's error decomposes into one term from estimating the response coefficient, one from aligning factors with auxiliary labels, and one from recovering the latent factors themselves. If correct, this gives a principled recipe for building site-transportable predictors from abundant unlabeled data plus a few labels.","feed_headline":"Intersecting loading spaces uncovers transferable latent factors","feed_subtitle":"ATLAS then aligns response-relevant factors with auxiliary labels, so predictors transfer at near-oracle error.","key_machinery":"The load-bearing object is the maximum invariant subspace: the intersection of the per-environment loading spaces, ∩_e col([B, A^(e)]) = col(B). ATLAS computes it empirically by averaging the top principal-subspace projectors of each environment's covariance matrix and taking the leading eigenspace of the average; this turns a set-theoretic intersection into a spectral step with a measurable eigen-gap (the parameter ϵ_A in Condition 4.2). The second piece is the diversified projection construction: each environment gets its own projection matrix that maps X to factor proxies while partialling out the heterogeneous factors from the invariant projection, so invariant and heterogeneous scores a","core_discovery":"On its own terms, the paper's central claim is that invariant and heterogeneous latent factors can be disentangled without any supervision, provided that the intersection of the column spaces of the per-environment loading matrices equals the shared loading space (Assumption 1(a)) and that invariant and heterogeneous factors are block-uncorrelated (Assumption 1(b)). Under those conditions, the invariant factors are identified up to a single common invertible transformation across all environments, and the heterogeneous factors up to environment-specific transformations; with a rank condition on the auxiliary-label coefficients, the response-relevant subset of each block is identified up to a","pith_inferences":["Editorial inference: the method's guarantees inherit the assumption that no spurious 'shared' direction exists across sites; in practice, a coding or measurement convention common to all observed environments will masquerade as an invariant factor. Institutions applying ATLAS should define environments to straddle known convention breaks (coding systems, note templates, lab vendors) so artifacts a","Editorial inference: the eigen-gap of the averaged projector suggests a practical diagnostic — plot the spectrum of the averaged projection matrix; if the (r_I+1)-th eigenvalue sits close to 1, the environments are not providing the exhaustive heterogeneity the identification needs, and transfer claims should be downgraded.","Editorial inference: the paper's sketch for adapting to a new environment implies a streaming variant — a new site's loading subspace could be intersected with the stored invariant space using only its own covariance and a handful of auxiliary labels, letting a federation of institutions update the transferable model without re-running the entire pipeline."],"forward_implications":["A fixed number of environments — as few as two, when heterogeneity is exhaustive — suffices to identify the invariant factors; the number of environments does not need to grow with the latent dimension.","Without auxiliary labels, restricting prediction to the invariant factors is the unique worst-case-optimal strategy under rotation uncertainty; heterogeneous factors cannot be transferred without additional supervision.","Auxiliary labels improve efficiency even when no heterogeneous factor is response-relevant, by reducing the variance of the estimated invariant signal.","The theory covers weak factors (loading strength need not scale with √d) and gives direction-wise, dimension-free sub-Gaussian error control, so downstream error does not accumulate over the factor dimension.","The framework extends to nonlinear mean functions: replacing the final regression step with nonparametric estimation substitutes a nonparametric rate for the parametric estimation error, leaving the factor-alignment machinery unchanged."],"fun_headline_variants":["Space intersection separates invariant and heterogeneous factors","ATLAS aligns latent factors with labels for near-oracle transfer","Loading-space intersection uncovers transferable latent factors","Unsupervised disentanglement of invariant factors for robust prediction"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire identification rests on Assumption 1(a): the intersection of loading spaces across environments is exactly the shared invariant space, meaning no environment-specific factor happens to load in the same direction at every observed site; if one does (a shared coding artifact, a common assay drift), it will be classified as invariant, the 'invariant' predictor will carry that spurious signal, and the failure is undetectable from the data.","fun_headline_variants_meta":{"raw":{"variants":["Space intersection separates invariant and heterogeneous factors","ATLAS aligns latent factors with labels for near-oracle transfer","Loading-space intersection uncovers transferable latent factors","Unsupervised disentanglement of invariant factors for robust prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000514,"raw_usage":{"total_tokens":2335,"prompt_tokens":751,"completion_tokens":1584,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1522}},"tokens_in":495,"tokens_out":1584,"duration_ms":11364,"temperature":1.0,"reasoning_tokens":1522,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:40:08.019014+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate two or more environments with a true invariant loading B and environment-specific loadings A^(e), then add one extra loading vector shared by all environments (violating Assumption 1(a)) and run ATLAS: the recovered invariant space will include the extra direction, and the transferred predictor's worst-case out-of-sample risk will visibly exceed the oracle built on the true B. A cheaper check: compute the average of the per-environment top-subspace projectors and inspect the eigenvalue just after the r_I-th; if it is not separated from 1 (ϵ_A near zero in Condition 4.2), the identific","supporting_citations":[],"review_version":1}