{"id":"50d58a30-6676-4ac8-a439-c89490c63d5f","arxiv_id":"2607.16035","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A p-parameter noisy dynamical model is generically identifiable from the expected values of 2p+1 random Fourier features, provided the full data distribution already identifies the parameters.","lead":"This paper claims that a dynamical model with p unknown parameters, observed through noise, can be uniquely pinned down by 2p+1 random summary features and a simulation-matching procedure. The idea could make model calibration far less reliant on hand-crafted statistics, but the proof currently omits a key genericity step.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Random Fourier features are not shown to be prevalent: Theorem 2's 'almost surely' does not transfer from C^1 to the finite-dimensional family Eq. (1).","rationale":"The reader's weakest_assumption identifies exactly the same gap: the paper proves prevalence of embeddings in C^1, not prevalence (or almost-sure probability) within the specific random Fourier feature family. This is the central logical bridge from the genericity theorems to the identification claim. Without it, Theorem 2 only gives a conditional statement: if the random features embed, then identification works. The title claim that 2p+1 random features identify p parameters is therefore not established by the manuscript. I agree with the REJECT verdict and do not see a need to adjust it. The concern is not that the conclusion is known to be false, but that the argument as written does not prove it; the missing step may be repairable with additional assumptions, which is why the appropriate disposition is rejection of the current version, not acceptance.","tokens_in":12064,"tokens_out":7943,"duration_ms":85514,"concrete_test":"Test the transfer step for a simple family satisfying Assumptions 1–2, e.g., p=1, Θ=[-1,1], P_θ = N(θ,1), k=3. For this family Φ_i(θ) = exp(-Ω_i^2/2) cos(Ω_i θ + α_i). Compute analytically or by high-resolution grid search the measure of the set of (Ω_1,Ω_2,Ω_3,α_1,α_2,α_3) for which θ↦Φ(θ) is one-to-one on [-1,1]. If this set does not have full Lebesgue measure, the 'almost surely' claim is false for a model satisfying the paper's assumptions. If it does have full measure, the claim survives this case but still requires a direct proof for general P_θ; the current proof, which appeals only to prevalence in C^1, remains invalid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is Section II.D's assertion that the random cosine features of Eq. (1) yield an embedding 'almost surely, by Theorem 2.' Theorem 2 (Appendix A, Step 3) only proves that embeddings are prevalent in C^1(R^p,R^k), i.e., that 'almost every' C^1 map is one-to-one on Θ. But the feature maps generated by Eq. (1) form a finite-dimensional family parameterized by (Ω_i, α_i), i=1..k, of dimension k(d(m+1)+1). A finite-dimensional subset of an infinite-dimensional Banach space is shy, not prevalent, so the prevalence statement about C^1 says nothing about almost every draw from this restricted family. To justify the title claim, one must prove that, for P_θ satisfying Assumption 2, the probability over the random draw (Ω_i, α_i) that θ ↦ (E_θ[cos(Ω_1·X+α_1)], ..., E_θ[cos(Ω_k·X+α_k)]) is injective on Θ equals one. No such proof is given; the sentence 'Indeed, by Theorem 2, they will, almost surely' is a non sequitur. The simulations in Section III are two examples only and cannot establish the generic claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims that for a p-parameter dynamic model observed with non-iid noise, the expectation values of 2p+1 random Fourier features of length-(m+1) blocks identify the parameter vector. Theorems 2 and 3 are stated for stationary and nonstationary cases, respectively, and are supposed to follow from the Sauer–Yorke–Casdagli fractal Whitney embedding theorem: since almost every C^1 map from R^p to R^k is an embedding for k≥2p+1, and the random-feature expectation map is C^1, the random features should almost surely embed the parameter space. Applications to a noisily observed Hénon map and Lorenz-63 system are presented as illustrations. The formal theorems only prove prevalence among all C^1 maps; the transfer to the specific finite-dimensional random cosine family is not proved.","tokens_in":12360,"tokens_out":10480,"duration_ms":100530,"significance":"The intended contribution is valuable if correct: it would replace user-chosen summary statistics or trained network summaries by simple random Fourier features and give a clean identification principle for nonlinear dynamic models with dependent, non-Gaussian noise. The paper also has genuine strengths: it is explicit about the genericity axioms, the smoothness computation in Step 2 of the proof of Theorem 2 is coherent under the stated Fréchet differentiability assumption, and the focus on non-iid noise is a useful extension of noiseless embedding results. However, the core claim is not established. The paper's central step—that a random draw from the finite-dimensional family Eq. (1) produces an embedding almost surely—is asserted rather than proved, and the formal theorems do not even state this claim. The simulations are two examples and cannot substitute for the missing proof. As it stands, the rigorously established part of the paper is a conditional finite-feature reduction under full-distribution identifiability, not the headline identification theorem.","major_comments":[{"comment":"The step from prevalence in C^1 to the random Fourier feature family is unproved and is load-bearing. Step 3 of the proof of Theorem 2 only establishes that embeddings are prevalent in C^1(Θ,R^k), a restatement of Sauer et al. The feature maps generated by Eq. (1) form a finite-dimensional family, parameterized by (Ω_i, α_i), i=1,...,k, of dimension k(d(m+1)+1). A finite-dimensional subset of the infinite-dimensional space C^1(R^p,R^k) is shy, not prevalent, so the prevalence statement about C^1 gives no probability statement for random draws from this family. The sentence in §II.D, 'Indeed, by Theorem 2, they will, almost surely,' is therefore not a consequence of Theorem 2. One needs a direct proof that, under Assumption 2, the distribution over (Ω_i, α_i) makes θ ↦ (E_θ[cos(Ω_1·X+α_1)], ..., E_θ[cos(Ω_k·X+α_k)]) injective on Θ with probability one. No such argument appears; Theorem 3","section":"§II.D, Theorem 2; Appendix A Step 3"},{"comment":"The formal statement of Theorem 2 does not contain the headline claim. It concludes only that C⊂C^1(R^p,R^k) and that almost every C^1 map is one-to-one on Θ, with no assertion about the specific random feature map of Eq. (2). The identification claim is then added in the prose ('Suppose the randomly chosen functions ... Indeed, by Theorem 2, they will, almost surely'). Thus the stated theorem cannot support the abstract's claim. This is not a minor presentational issue: a reader cannot verify the main theorem from the formal result. The same structural problem applies to Theorem 3.","section":"§II.D, Theorem 2 statement"},{"comment":"Assumption 2.2 assumes θ ↦ P_θ is a bijection, which is essentially full distributional identifiability from the (m+1)-block. This is a strong primitive: it already grants the kind of identifiability the paper aims to establish, albeit at the level of the full distribution rather than 2p+1 features. The finite-feature reduction is a meaningful goal, but the paper should present it as such and should not imply identification is obtained from scratch. As written, the rigorously proved part of Theorem 2 is conditional both on full-distribution identifiability and on the unproved random-feature embedding, so the title overclaims.","section":"Assumption 2.2, §II.D"}],"minor_comments":[{"comment":"The map p takes values in the space of probability measures, which is not a vector space. Fréchet differentiability should be defined with respect to the ambient vector space of signed measures with total-variation norm. Please state this embedding explicitly.","section":"§II.D, Assumption 2.3"},{"comment":"The proof uses the fact that Lipschitz maps do not increase box-counting dimension. Please specify which box-counting dimension (upper/lower) is used and give the inequality explicitly, since the manuscript does not define the notion.","section":"§II.E, Step 3 of Theorem 3 proof"},{"comment":"The simulation section lacks implementation details: the optimization method, the number of simulations s, the lag m, the number of Monte Carlo replications behind the densities, and the handling of the rolling-window choices are not described. This limits reproducibility and makes the empirical illustrations hard to assess.","section":"§III"},{"comment":"The abstract promises practical estimation procedures, but the consistency of the time-average and rolling-window estimators is entirely deferred to the companion paper [21]. If the present paper is meant to stand alone, the relevant consistency statements should be restated or precisely referenced.","section":"Abstract and §IV"},{"comment":"The caption refers to color for distinguishing parameter ranges; please ensure the figure is also readable in black-and-white or add markers.","section":"Figure 1 caption"}],"recommendation":"reject","confidential_remarks":"The central gap identified in the stress-test note is real and, in my reading, fatal to the paper's main claim. The prevalence statement for C^1 maps does not transfer to the finite-dimensional random cosine family, and no alternative proof is supplied. The formal theorems do not state the headline result. A revision would require a substantially new argument, not a local correction; I would be willing to reconsider if the authors prove the missing random-feature embedding property under explicit assumptions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere's my read of the paper. The core idea—that noisy dynamical systems can be identified by matching 2p+1 generic features, with the features drawn at random—is genuinely attractive. It extends Sontag's noiseless experiment-count result and the Sauer/Takens embedding program to a setting with process and observation noise. The smoothness argument in Appendix A is coherent under Assumptions 1–2, and the two examples (Hénon and Lorenz) are at least suggestive.\n\nThe problem is the load-bearing step in Section II.D. The authors prove that embeddings are prevalent in C^1(R^p, R^k), and then assert that the random cosine features of Eq. (1) will therefore embed 'almost surely.' That doesn't follow. The cosine family is a finite-dimensional set parameterized by the frequencies and phases; a finite-dimensional subset of the infinite-dimensional C^1 space is shy, not prevalent. Prevalence over all C^1 maps tells you nothing about draws from this restricted family. To justify the title claim you would need a separate argument that for P_θ satisfying Assumption 2, the map θ ↦ (E_θ[cos(Ω_1·X+α_1)], ...) is injective on Θ with probability one over the random draws. No such proof appears.\n\nThere's a second, quieter issue. Assumption 2.2—θ ↦ P_θ is a bijection—is essentially full-data identifiability. The theorem then says that if the parameters are identified by the full distribution, they are identified by 2p+1 features. That's a useful reduction, but it is weaker than the title's promise of identification from random features without prior identifiability assumptions. The consistency results and the nonstationary treatments are deferred to a companion manuscript, so the present paper is really a statement of a principle plus an unproven generic-feature claim.\n\nThe two simulations are just examples; they don't establish the generic claim. Still, the flaw is a gap, not an incoherence. A revision that either proves the prevalence of the random cosine family or switches to a genuinely prevalent feature class (e.g., full random smooth functions) could make this work.\n\nWho is this for? People working in simulation-based inference and nonlinear time series. It's not there yet, but it deserves a careful referee who will push the authors on the finite-dimensional-versus-prevalence issue.\n\nMy recommendation: send it to peer review, but with a clear signal that the centrality of the embedding step has to be addressed.\n\nBest.","headline":"The central genericity step doesn't go through—prevalence over C^1 maps doesn't transfer to the finite-dimensional random cosine family—so the title claim isn't proven, though the idea is worth a serious revision.","tokens_in":12850,"tokens_out":2689,"would_cite":false,"duration_ms":24872,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that the parameters of a noisy dynamical model can be uniquely recovered from 2p+1 randomly chosen features of the observed time series, where p is the number of unknown parameters.","keywords":["parameter identification","random features","embedding theorem","dynamical systems","noisy time series","simulation-based inference","prevalence","Fourier features"],"falsifier":"Take a parametric family satisfying the paper's assumptions and, for a fixed k=2p+1, numerically estimate the probability (over draws of Ω and α) that the feature-expectation map Φ is non-injective on a fine grid covering Θ; if for some family this probability does not approach zero, the almost-sure embedding claim is false. Equivalently, exhibit two parameters θ1≠θ2 such that E_{θ1}cos(Ω·X+α)=E_{θ2}cos(Ω·X+α) for a set of (Ω,α) of positive measure.","tokens_in":11925,"feed_emoji":"📈","tokens_out":8565,"duration_ms":80019,"temperature":0.7,"pith_summary":"This paper establishes that the parameters of a noisy dynamical model can be uniquely recovered from 2p+1 randomly chosen features of the observed time series, where p is the number of unknown parameters. The central object is the map sending a parameter vector to the expected values of random Fourier features; the authors show this map is, generically, a one-to-one immersion, so matching simulated and observed feature averages locates the true parameters. Unlike classical state-space reconstruction, which recovers the state only up to coordinate changes, this identifies the numerical parameters themselves, under weak assumptions about the noise (it may be non-iid, non-Gaussian, and state-dependent). The result applies both to stationary processes, via time-averaged features, and to nonstationary processes such as noisy differential equations, via rolling-window features. Practical estimators based on this principle are demonstrated on the Lorenz-63 and Hénon models, suggesting a route to simulation-based inference without hand-picked summary statistics.","feed_headline":"2p+1 random features identify p noisy-model parameters","feed_subtitle":"Matching expected random Fourier features uniquely recovers parameters of noisy differential and discrete-time models.","key_machinery":"The load-bearing object is the random Fourier feature map φ_i(x)=cos(Σ_j Ω_{ij} x_j + α_i) and its expectation Φ(θ)=E_θ[φ(X)]. The dimension count k≥2p+1 comes from the fractal embedding prevalence theorem, which says that almost every smooth map from R^p to R^k is one-to-one on a compact p-dimensional set and immersed on its smooth pieces. The paper applies this theorem to the expected-feature map, transferring genericity from the space of all smooth maps to the specific random-feature family; this transfer is what makes the identification claim an almost-sure property of the random draws rather than a statement about specially designed measurements.","core_discovery":"The paper's central claim is that identification of a p-dimensional parameter θ in a noisy dynamical model can be achieved through the expectation values of k=2p+1 random features φ_i(x)=cos(Ω_i·x+α_i), with Gaussian frequencies and uniform phases. Defining Φ(θ)=E_θ[φ(X)], the authors argue that for almost every draw of the random features, Φ is one-to-one on the compact parameter space Θ and is a C^1 diffeomorphism between Θ and its image, provided the map from parameters to stationary or local distributions is smooth and injective. Therefore Φ(θ1)=Φ(θ2) only when θ1=θ2, and the population discrepancy Q(θ)=∥Φ(θ0)-Φ(θ)∥ is nonzero away from the truth. This converts parameter identification i","pith_inferences":["The proof establishes prevalence of embeddings among all C^1 maps but does not explicitly prove that the random cosine family inherits this prevalence; verifying that the pushforward measure of the random (Ω,α) puts full mass on embeddings would close the gap between the conditional result ('if the features embed') and the almost-sure statement.","If the almost-sure embedding holds, the same machinery suggests a general recipe for automatic summary statistics in simulation-based inference: draw random features once, fix them, and match expectations; this could be tested on models outside dynamics, such as spatial point processes or random graphs.","The identification result is non-asymptotic in the sense that it concerns population expectations; the finite-sample question is whether the empirical discrepancies concentrate quickly enough, which the companion work addresses with asymptotics. A natural testable extension is to measure how the required sample size scales with p and with the feature count."],"forward_implications":["For any parametric time-series model satisfying the smoothness and injectivity assumptions, parameter estimation reduces to minimizing a discrepancy between observed and simulated random-feature averages, with no user-chosen summary statistics or neural-network training.","The 2p+1 count gives a concrete prescription for the number of features to simulate, mirroring the classical 2d+1 delays in state-space reconstruction and 2r+1 experiments for noiseless systems.","Because the noise can be non-iid, non-Gaussian, and state-dependent, the result covers realistic observational and process noise that breaks standard likelihood-based identifiability arguments.","The rolling-window version extends identification to nonstationary trajectories, including noisily observed differential equations on a fixed horizon, where time averages do not converge.","The same feature-matching principle could be applied to other data structures, such as spatiotemporal fields or networks, wherever a class of random features for the dependence structure is available."],"fun_headline_variants":["2p+1 random features pin down p noisy parameters","Random features unlock noisy model parameters","2p+1 features recover p noisy parameters","Noisy models identified by 2p+1 random features","Random features make noisy dynamics identifiable"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The decisive premise is that embeddings are not just prevalent among all smooth maps but are almost sure under the specific distribution of random Fourier features; the paper asserts this transfer but its proof only covers the space of all C^1 maps, so the identification claim depends on that unproven inheritance.","fun_headline_variants_meta":{"raw":{"variants":["2p+1 random features pin down p noisy parameters","Random features unlock noisy model parameters","2p+1 features recover p noisy parameters","Noisy models identified by 2p+1 random features","Random features make noisy dynamics identifiable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":958,"prompt_tokens":664,"completion_tokens":294,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":224}},"tokens_in":408,"tokens_out":294,"duration_ms":3335,"temperature":1.0,"reasoning_tokens":224,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T21:32:25.509588+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a parametric family satisfying the paper's assumptions and, for a fixed k=2p+1, numerically estimate the probability (over draws of Ω and α) that the feature-expectation map Φ is non-injective on a fine grid covering Θ; if for some family this probability does not approach zero, the almost-sure embedding claim is false. Equivalently, exhibit two parameters θ1≠θ2 such that E_{θ1}cos(Ω·X+α)=E_{θ2}cos(Ω·X+α) for a set of (Ω,α) of positive measure.","supporting_citations":[],"review_version":1}