{"id":"e1185283-6c01-45bc-987b-c40975cb8518","arxiv_id":"2501.12139","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Energy-information tradeoff in a linear lateral predictive coding network produces two discontinuous phase transitions, with a single unit abruptly becoming selective to a non-Gaussian feature at low and high temperatures.","lead":"A linear network model with lateral predictive coding can learn to detect a non-Gaussian signal hidden in Gaussian noise, and this ability switches on and off abruptly as the network changes its balance between energy use and information preservation. The result suggests that sudden jumps in learning or perception could emerge from a simple energy-information tradeoff.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed transition temperatures are computed from an E(S) curve whose global optimality is not certified; the paper's own local-minima counts show this is a live quantitative risk.","rationale":"I agree with the reader's weakest_assumption. The analytical framework is internally consistent: Eq. (4) follows from the Jacobian of the linear map, Eq. (15) follows from the conditional Gaussian statistics, and the free-energy arithmetic at the two transition temperatures checks out. The paper also honestly reports the existence of local minima, which is an important piece of evidence rather than a concealment. However, the load-bearing quantitative object is the global-minimum curve E(S). The reported local-minimum energy gaps (e.g., 27.55 vs 27.4955 at S=-1.5) are comparable to the free-energy differences that determine the transition temperatures, so unverified global optimality directly threatens the exact locations and even the existence of the intermediate non-detecting phase. This is an internal verification gap, not a disagreement with external consensus, and it does not require rejecting the qualitative picture. Because the reader already made this the basis of a CONDITIONAL verdict, my stress-test does not change the verdict: the paper remains CONDITIONAL pending independent confirmation of the global minima.","tokens_in":32691,"tokens_out":10448,"duration_ms":115732,"concrete_test":"Run a certified global-optimality check on the N=10 version: (1) enumerate the symmetric block families in Fig. S2 (δ1/δ2/δ3/α1/α2/β/γ) and solve their stationarity conditions exactly, giving a lower bound for E(S) within that class; (2) repeat the annealing at S∈{-1.5, 0, 7.1} for N=36 with at least 1000 random initial W, ε=0.005 cooling from κ=10 to 10^10, and record the lowest E by Q class. If any reported E(S) is beaten by an enumerated matrix or a random-restart run with a different Q, recompute F=E-TS and the crossing temperatures; if the crossings at 0.8320 and 1.1283 vanish or shift by more than about 1%, the discontinuous-transition claim is not supported by the data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—two discontinuous transitions at T=0.8320 and T=1.1283 for N=36, and the analogous two-feature transitions at T=0.8316, 0.9935, 1.2612—is not derived from an exact free-energy calculation but from the curve E(S) produced by entropy-clamped annealing (Sec. III B). The paper's own Fig. 1 shows the algorithm frequently stops in local minima: at S=-1.5, 71% of 600 runs report E≈27.55 instead of the claimed global 27.4955, and at S=0 about 40% of runs report high-Q matrices at E≈29.15 instead of 28.7235. Each such local minimum is a different structural phase (one selective unit vs several partially selective units), i.e. exactly the order parameter used to define the transitions. Since the competing free-energy minima are balanced to four decimal places at T=0.8320 and T=1.1283, a systematic miss of the true E(S) by a few hundredths on one branch—an amount smaller than the observed local-minimum gaps—would move the crossings and could eliminate the intermediate non-detecting phase. The annealing is also always started from a single initial matrix, so disconnected basins are not guaranteed to be explored. The analytical expressions (Eq. 15, Eq. 4) are internally consistent, but the E(S) curve has no error bars and no independent optimization (e.g. random-restart basin hopping, exact small-N enumeration) is supplied; global optimality is therefore assumed, not demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies a linear lateral predictive coding (LPC) model with an L1-norm energetic cost E (Eq. 3) and an entropy measure S = -ln det(I+W) (Eq. 4), optimized through the free energy F = E - T S (Eq. 5). For inputs consisting of one non-Gaussian feature embedded in Gaussian background with identity covariance, the authors derive conditional output statistics and the analytical energy expression Eq. (15), then use entropy-clamped annealing to construct the minimum-energy curve E(S). From this curve they infer two discontinuous phase transitions at T = 0.8320 and T = 1.1283 for N = 36, p0 = 0.7: a low-temperature selective phase, an intermediate S ≈ 0 non-selective phase, and a high-temperature selective phase. The same construction is applied to two non-Gaussian features with angle θ = π/4 (N = 16, p0 = 0.6), yielding three reported transitions at T = 0.8316, 0.9935, and 1.2612. The paper also reports a spectrum analysis showing that the complex eigenvalues of the optimal (I+W) lie approximately on a semicircle.","tokens_in":32960,"tokens_out":10643,"duration_ms":113760,"significance":"If the numerical determination is reliable, the result is significant: it demonstrates in a transparent linear model how non-Gaussian structure can be detected through L1-norm energy minimization, with feature detection emerging spontaneously from an energy-information tradeoff rather than from an imposed sparsity penalty. The analytical parts—the conditional single-unit statistics, the energy expression Eq. (15), and the entropy calculation—are internally consistent, and I found no error in the free-energy arithmetic at the reported crossings. The paper is also commendably candid about the existence of local minima. The main weakness is that the central quantitative claim is built on an E(S) curve whose global optimality is not certified; the paper's own numbers show that the annealing frequently stops in structurally different local minima, and the reported transition temperatures are obtained by comparing free-energy branches whose differences are comparable to the observed local-minimum gaps.","major_comments":[{"comment":"The global optimality of the E(S) curve is not established, and this is load-bearing because the transition temperatures are read from crossings of minima of F = E - T S. The paper's own data show that the annealing frequently terminates in local minima that differ in exactly the order parameter used to define phases: at S = 0 about 40% of runs find high-Q matrices with E ≈ 29.15 instead of the claimed global E = 28.7235, and at S = -1.5 about 71% of runs find E ≈ 27.55 with Q ≈ 0.565 instead of E = 27.4955 with Q ≈ 0.8387. A systematic miss of a few hundredths in E on one branch—an order of magnitude smaller than the gaps visible in Fig. 1—could shift the crossings at T = 0.8320 and T = 1.1283 or even eliminate the intermediate non-detecting phase. The manuscript should provide independent certification of the global minima, for example exact enumeration or branch-and-bound for a small N, multi-start basin hopping with random initial matrices, and/or explicit error bars or bounds on the E(S) values used in Fig. 3.","section":"Sec. III.B, Fig. 1, Fig. 3"},{"comment":"The transition temperatures are not accompanied by any sensitivity or error analysis. The two degenerate free-energy minima at T = 1.1283 (S = 7.10 and S = 0) and at T = 0.8320 (S = -1.12 and S = 0) are claimed to be global minima of F, but Fig. 3(d) shows no uncertainty, and the text gives no stopping criterion or optimization tolerance beyond the local-minimum counts of Fig. 1. Because the competing branches are balanced to a precision comparable to the observed local-minimum gaps, the authors should report how T_c changes when E(S) on one branch is perturbed by, say, ±0.01, and should state explicitly whether the reported transitions survive under such perturbations. The same requirement applies to the two-feature transitions reported in Fig. 6, which inherit the same annealing-based E(S) construction.","section":"Sec. III.C, Fig. 3(d); Sec. IV, Fig. 6"},{"comment":"For the two-feature setup with θ ≠ π/2, the input covariance implied by Eq. (16) is not the identity: the feature directions contribute variance 2cos^2(θ/2) along φ1 and 2sin^2(θ/2) along φ2, while the background contributes unit variance. The paper does not discuss whether this makes the task partially accessible to second-order statistics. This does not invalidate the energy-information tradeoff argument, but the text should state the covariance explicitly and clarify what aspect of the two-feature task is genuinely not solvable by PCA when the features are non-orthogonal.","section":"Sec. IV, Eq. (16)"}],"minor_comments":[{"comment":"The phrase 'Energy–information tradeoff induce' should be 'Energy–information tradeoff induces' for subject-verb agreement.","section":"Abstract"},{"comment":"The text says 'typically 10 6 trials'; this should be typeset as 10^6.","section":"Sec. III.B"},{"comment":"The statement that ⟨x_i^2⟩ = (π/2)^{1/2}T appears dimensionally inconsistent with E = N T and with the preceding derivation, which gives ⟨x_i^2⟩ = π T^2 / 2. Please check and correct this expression.","section":"Appendix C, after Eq. (C12)"},{"comment":"The phrase 'magnitudes sqrt(|λ|^2)' should be simply 'magnitudes |λ|'.","section":"Fig. 5(d) caption"},{"comment":"The label 'Rank' is used for the sorted run index in Fig. 1 and for the unit index in Fig. 2; consider using distinct labels such as 'run index' and 'unit index' to avoid confusion.","section":"Fig. 1 and Fig. 2"},{"comment":"The supplementary material is cited as Ref. [30] without a formal title or identifier; please provide a stable citation or DOI for the supplement.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the readership, and the analytical framework is sound. The main risk is that the central quantitative claim rests on an un-certified E(S) curve from entropy-clamped annealing. I would support revision if the authors provide exact small-N checks or a clear sensitivity analysis showing that the reported transition temperatures are robust to small changes in the branch energies. The local-minima counts reported in Fig. 1 are honest but they actually strengthen the concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's the read.\n\nThe genuinely new thing is the L1-norm energy term. With L2 energy, the cost only sees the correlation matrix, so non-Gaussian features are invisible. Replacing it with L1 makes the objective sensitive to higher-order moments, and the paper shows that minimizing E - T S spontaneously produces a single unit that detects the hidden non-Gaussian feature at both low and high T, with an undetecting phase in between. That is a real result and a clean way to connect predictive coding to sparse coding and ICA.\n\nThe analytical core is in good shape. The energy expression Eq. (15), the conditional output statistics, and the free-energy arithmetic at the stated crossings are internally consistent. The paper also reports local-minima bands honestly, which is more than most optimization papers do. Self-citation to Ref. [13] is appropriate: the entropy definition and annealing method come from there. The qualitative picture is robust across different feature directions, p0 values, Laplace and power-law priors, and the two-feature case.\n\nNow the soft spot. The quantitative claim lives entirely in the E(S) curve from entropy-clamped annealing. The paper's own Fig. 1 shows the algorithm frequently landing in local minima whose energy gaps are 0.05 or more—at S=0, the global minimum is 28.72 but 40% of runs sit at 29.15 with a qualitatively different matrix. The free-energy crossings are balanced to four decimal places. A miss of a few hundredths on one branch, smaller than the observed local-minimum gaps, could shift those transition temperatures and potentially shrink or eliminate the intermediate undetecting phase. There are no error bars on E(S) and no independent optimization (exhaustive small-N enumeration, basin hopping, or similar). So the exact values T=0.8320 and T=1.1283 are not established; the existence of the two transitions is plausible but not proven.\n\nIs that load-bearing? For the paper's central claim, partially. If the annealing is missing a lower-energy detecting branch at S=0, the middle phase could vanish. The authors argue the branches are genuine because they see consistent bands and because E(S) is monotonic, but that is not a certificate. This makes the paper a solid conditional result rather than a definitive one.\n\nWho this is for: statistical physicists working on neural representation, and anyone interested in energy-information tradeoffs as an optimization principle. A serious referee can push on the global optimality question and the paper will survive or not depending on that.\n\nRecommendation: send it to peer review. The framework and the qualitative phenomenon are worth referee time, and the numerical gap is addressable in revision.","headline":"A solid analytical framework for energy-entropy tradeoff in lateral predictive coding, but the exact discontinuity temperatures rest on an annealing search that demonstrably misses global minima.","tokens_in":33527,"tokens_out":1888,"would_cite":true,"duration_ms":20929,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["82B26","92B20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Minimizing a free energy in lateral predictive coding makes feature detection appear and disappear through two discontinuous phase transitions.","keywords":["lateral predictive coding","phase transition","feature detection","energy-information tradeoff","L1-norm energy","non-Gaussian features","free energy minimization","entropy"],"falsifier":"Evaluate the same $E(S)$ optimization on a small system ($N=10$) with a certified global-search method that cannot be trapped, and compare the resulting free-energy crossings to $T=0.8320$ and $T=1.1283$; if the global minimum at intermediate $S$ is not the weak homogeneous solution with $S\\approx0$ and $Q=1/\\sqrt{N}$, or if the low-energy feature-detecting branch near $S=-1.12$ is only a local minimum, the discontinuous transitions are artifacts of the annealing algorithm. A concrete spot to check is the paper's own $S=-1.5$ data, where 71% of runs return $E\\approx27.55$ instead of the claimed global $E\\approx27.4955$.","tokens_in":32413,"feed_emoji":"🧠","tokens_out":9967,"duration_ms":96490,"temperature":0.7,"pith_summary":"This paper tries to establish that a single optimization principle — minimizing a free energy that balances transmission cost against information robustness — is enough to make a lateral predictive coding (LPC) network detect non-Gaussian features that are invisible to principal component analysis. The cost is the mean absolute ($L_1$-norm) output of the units, and information robustness is measured by the entropy difference $S = -\\ln \\det(I+W)$ of the output distribution. For one hidden non-Gaussian feature among Gaussian backgrounds with the same variance, the optimal weight matrix comes in three distinct types, and the switches between them at two temperatures are discontinuous. At both low and high temperatures one unit snaps into selective response to the feature (overlap $Q \\approx 0.83$–$0.97$), while at intermediate temperatures the optimal matrix is weak and homogeneous with no selectivity. The same mechanism separates two non-Gaussian features into two distinct units, even when the two features are not orthogonal.","feed_headline":"Two sharp transitions turn on feature detection in predictive coding","feed_subtitle":"At low and high temperature one unit locks onto the hidden feature; in between it vanishes.","key_machinery":"The load-bearing object is the thermodynamic free energy $F = E - T S$, with $S = -\\ln \\det(I+W)$ measuring how much the linear map $\\vec{x} = (I+W)^{-1}\\vec{s}$ expands output entropy (information robustness) and $E$ the mean $L_1$-norm of internal states (metabolic cost). The argument is carried by the order parameter $Q = \\max_i |\\mu_i| / \\sqrt{\\sum_j \\mu_j^2}$, where $\\mu_i$ is the mean response of unit $i$ to the non-Gaussian feature; $Q \\approx 1$ means one unit is selective and $Q = 1/\\sqrt{N}$ means all units respond equally and weakly. Because the energy curve $E(S)$ is non-convex in the middle entropy range, the free energy has two coexisting minima at the transition temperatures. The $L_1$ norm is essential: it couples the energy to higher moments of the input, which lets the non-Gaussian feature be detected even though the input correlation matrix is the identity and carries no directional information.","core_discovery":"The central claim is that in the linear LPC model with steady-state output $\\vec{x} = (I+W)^{-1}\\vec{s}$ and $L_1$-norm energy $E$, the global minimum of $F = E - T S$ undergoes two discontinuous phase transitions in the temperature-like parameter $T$. At $T = 1.1283$ (for $N=36$, $p_0=0.7$) the free-energy landscape has two degenerate minima, one at $S \\approx 7.10$ with $Q \\approx 0.97$ and one at $S = 0$ with $Q = 0.1667 = 1/\\sqrt{36}$; below $T = 0.8320$ another degenerate pair appears near $S \\approx -1.12$ with $Q \\approx 0.83$. Feature detection is therefore feasible at high and low temperature but impossible in between. The paper extends the same optimization to two non-orthogonal non-Gaussian features and reports three discontinuous transitions at $T \\approx 0.8316$, $0.9935$, and $1.2612$, after which the two features are represented by two separate single units. The same qualitative scenario is reported for Laplace and power-law feature distributions and for random feature directions.","pith_inferences":["If a biologically plausible local learning rule can be shown to descend the same free energy, the model predicts sudden, insight-like jumps in single-neuron selectivity during learning, because the feature-detecting and feature-blind solutions are separated by an energy barrier rather than connected continuously.","A natural next experiment is to scan the fraction $p_g$ of Gaussian trials in the coefficient $a_1$; the paper notes this direction but does not explore it, and one would expect a critical $p_g$ beyond which the $L_1$ energy can no longer discriminate the non-Gaussian feature.","For artificial recurrent networks with lateral connections, replacing the usual squared internal states by an $L_1$ cost on activations should reproduce the same discontinuous assignment of hidden units to non-Gaussian latent sources, giving a practical signature of the transition outside the biological setting."],"forward_implications":["At $T<0.8320$ and $T>1.1283$ for the $N=36$, $p_0=0.7$ ensemble, the global minimum of $F=E-TS$ is a matrix in which one unit responds selectively to the non-Gaussian feature, with overlap $Q\\approx0.83$–$0.97$.","Between those temperatures, the optimal matrix has very weak, homogeneous lateral weights, $S\\approx0$, and $Q=1/\\sqrt{N}$, so feature detection fails; the switches at $T=0.8320$ and $T=1.1283$ are discontinuous.","For two non-Gaussian features with angle $\\theta=\\pi/4$, $N=16$, and $p_0=0.6$, the same tradeoff produces discontinuous transitions at $T\\approx0.8316$, $0.9935$, and $1.2612$, after which two different units each detect one feature — even when the features are not orthogonal.","The $L_1$-norm energy, rather than the usual $L_2$-norm, is what enables detection: an $L_2$-norm energy depends only on the input correlation matrix, which is exactly the same for the Gaussian background and the non-Gaussian feature.","The qualitative result persists for the continuous Laplace distribution, a power-law feature distribution, and randomly oriented feature directions, so the mechanism is not an artifact of the discrete three-point prior used in the main runs."],"supporting_citations":[{"why":"supplies the original lateral predictive coding idea and the recursive dynamics that the model optimizes.","marker":"[1]"},{"why":"supplies the formulation of non-symmetric lateral predictive coding interactions that the synaptic matrix $W$ is built from.","marker":"[7]"},{"why":"provides the natural-image-statistics framing and the non-Gaussian feature background that define the detection task.","marker":"[12]"},{"why":"supplies the free-energy construction, the entropy definition, and the entropy-clamped annealing algorithm used throughout.","marker":"[13]"},{"why":"frames blind separation of sources, the task that the two-feature part extends toward.","marker":"[16]"},{"why":"motivates entropy maximization as the information-robustness side of the tradeoff.","marker":"[25]"},{"why":"defines the setting of a single non-Gaussian feature hidden in Gaussian background that the LPC model is tested on.","marker":"[29]"}],"fun_headline_variants":["Discontinuous transitions control feature detection in predictive coding","Two sharp transitions switch feature detection on and off","Energy-information tradeoff triggers abrupt feature detection","Lateral predictive coding flips feature detection at two critical temperatures","Predictive coding: feature detection appears only at low and high temperature"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the annealing search finds the true global energy minimum at every fixed entropy $S$; since roughly 71% of runs at $S=-1.5$ land in a local minimum with $E\\approx27.55$ rather than the claimed global $27.4955$, the phase-transition temperatures depend on an optimizer whose global optimality is not certified.","fun_headline_variants_meta":{"raw":{"variants":["Discontinuous transitions control feature detection in predictive coding","Two sharp transitions switch feature detection on and off","Energy-information tradeoff triggers abrupt feature detection","Lateral predictive coding flips feature detection at two critical temperatures","Predictive coding: feature detection appears only at low and high temperature"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1592,"prompt_tokens":1019,"completion_tokens":573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":505}},"tokens_in":635,"tokens_out":573,"duration_ms":6666,"temperature":1.0,"reasoning_tokens":505,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:28:39.628731+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Evaluate the same $E(S)$ optimization on a small system ($N=10$) with a certified global-search method that cannot be trapped, and compare the resulting free-energy crossings to $T=0.8320$ and $T=1.1283$; if the global minimum at intermediate $S$ is not the weak homogeneous solution with $S\\approx0$ and $Q=1/\\sqrt{N}$, or if the low-energy feature-detecting branch near $S=-1.12$ is only a local minimum, the discontinuous transitions are artifacts of the annealing algorithm. A concrete spot to check is the paper's own $S=-1.5$ data, where 71% of runs return $E\\approx27.55$ instead of the claimed global $E\\approx27.4955$.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the original lateral predictive coding idea and the recursive dynamics that the model optimizes."},{"cited_title":"Predictive coding is a consequence of energy efficiency in recurrent neural networks,","cited_arxiv_id":null,"evidence_quote":"supplies the formulation of non-symmetric lateral predictive coding interactions that the synaptic matrix $W$ is built from."},{"cited_title":"Efficient coding and energy efficiency are promoted by balanced excita- tory and inhibitory synaptic currents in neuronal net- work,","cited_arxiv_id":null,"evidence_quote":"provides the natural-image-statistics framing and the non-Gaussian feature background that define the detection task."},{"cited_title":"Co-emergence of multi-scale cortical activities of irregular firing, oscil- lations and avalanches achieves cost-efficient information capacity,","cited_arxiv_id":null,"evidence_quote":"supplies the free-energy construction, the entropy definition, and the entropy-clamped annealing algorithm used throughout."},{"cited_title":"Energy–information trade-off induces continuous and discontinuous phase transitions in lateral predictive cod- ing,","cited_arxiv_id":null,"evidence_quote":"frames blind separation of sources, the task that the two-feature part extends toward."},{"cited_title":"Evolutionary transitions in learning and cognition,","cited_arxiv_id":null,"evidence_quote":"motivates entropy maximization as the information-robustness side of the tradeoff."},{"cited_title":"Phase transitions in pareto optimal complex networks,","cited_arxiv_id":null,"evidence_quote":"defines the setting of a single non-Gaussian feature hidden in Gaussian background that the LPC model is tested on."}],"review_version":1}