{"id":"6ebc961d-d2f2-41ae-8cf8-7e7406c60c88","arxiv_id":"2504.20556","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For a fixed artificial-noise power budget, the optimal allocation that minimizes side-channel information leakage follows a dual water-filling rule for Gaussian leakage points and an I-MMSE-based rule for arbitrary input distributions.","lead":"This paper finds the smartest way to add artificial noise to a cryptographic chip so that an attacker measuring power or electromagnetic traces learns as little as possible about the secret, for a fixed energy budget. It provides closed-form formulas for splitting a noise power budget across leakage points, and shows the method beats uniform noise in simulations and on an AES power-trace dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 5's SNR-equalizing rule is only optimal if all subchannels share the same input distribution; the paper never states this, and the proof's monotonicity argument does not justify it.","rationale":"The reader identified the same weakest assumption, and I agree. The paper's other main results, Theorem 1 and Theorem 4, are standard convex/KKT arguments and appear internally consistent under their stated conditions. The issue is localized to the max-MI extension: the proof's inference from 'each I_i is increasing' to 'max MI minimization = max SNR minimization' is invalid unless the I_i are identical as functions of ρ. If the authors intended a common normalized input distribution for every subchannel, the theorem becomes correct and the fix is a one-sentence qualification plus a corrected proof; if they intended heterogeneous distributions, the theorem is false and the max-MI contribution for arbitrary inputs needs substantial revision. Either way the current statement overreaches, so the conditional verdict is appropriate. The numerical AES study uses Gaussian Xi, so it does not test the heterogeneous-distribution regime and cannot rescue Theorem 5 as written.","tokens_in":18884,"tokens_out":12955,"duration_ms":133214,"concrete_test":"Set m=2, P1=P2=1, Z1=Z2=0.1, N0=1. Let S1 be standard Gaussian and S2 equiprobable ±1, both unit variance. Compute I_Gauss(ρ)=0.5 log(1+ρ) and I_binary(ρ) from the exact BPSK mutual information integral. Solve (35) numerically over N1∈[0,1] with N2=1-N1, e.g., by a fine grid or scipy.optimize.minimize_scalar on max(I_Gauss(1/(N1+0.1)), I_binary(1/(1-N1+0.1))). Theorem 5 predicts the minimizer is N1=N2=0.5. If a different split yields strictly smaller max MI, the theorem's SNR-equalizing allocation is not optimal for heterogeneous input distributions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing defect is in Theorem 5 (Section V-D, proof in Appendix F). The proof replaces objective max_i I_i(P_i/(N_i+Z_i)) with max_i P_i/(N_i+Z_i), claiming that because I-MMSE implies each I_i is strictly increasing in SNR, minimizing max MI is equivalent to minimizing max SNR. Strict monotonicity is not enough when the per-point input distributions differ: it only gives order preservation inside each subchannel. The functions I_i(ρ) can be different, so the minimizer of the max-MI problem equalizes the values I_i(ρ_i) across active subchannels, not the SNRs ρ_i. Formula (36) forces ρ_i = κ for every active subchannel and is correct only when all subchannels share the same MI-versus-SNR function, i.e., the same normalized input distribution. Section V says the input follows 'an arbitrary distribution' but never states that all subchannels share that distribution. Under the natural reading with heterogeneous distributions, the theorem is false: e.g., with equal P_i and Z_i, one Gaussian and one binary subchannel, (36) gives equal noise powers, but the Gaussian subchannel dominates the max MI and the optimum shifts noise to the Gaussian subchannel until the I_i values match.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an information-theoretic framework for allocating artificial noise power across side-channel leakage points. The model is Y_i = X_i + Z_i + N_i with Gaussian physical and artificial noise, a fixed total artificial-noise budget, and subchannel powers P_i. For Gaussian inputs, the authors derive closed-form allocations that minimize the total mutual information (a dual water-filling rule, Theorem 1) and the maximum mutual information (an SNR-equalizing rule, Theorem 2). For arbitrary input distributions, they give convexity conditions via the I-MMSE relation (Theorem 3), a KKT-based solution for the total mutual information objective (Theorem 4), and claim that the maximum mutual information problem has the same SNR-equalizing solution as the Gaussian case, independent of the input distribution (Theorem 5). Numerical experiments compare the proposed allocations with uniform noise allocation and include an AES-128 case study on the SPERO dataset.","tokens_in":19102,"tokens_out":23329,"duration_ms":238360,"significance":"If the results hold as stated, the paper makes a useful contribution to side-channel countermeasure design: it gives closed-form, non-uniform noise allocations with a clear dual-water-filling interpretation, leverages the I-MMSE relation to obtain stationarity conditions for non-Gaussian inputs, and provides explicit convexity conditions. Strengths include the careful KKT derivations in Appendices A-E, the closed-form nature of the Gaussian solutions, and the demonstration that non-uniform allocation can substantially outperform uniform allocation. The main weakness is that the treatment of 'arbitrary input distributions' in Section V silently presumes that all subchannels share the same input distribution; without this assumption, the central maximum-mutual-information claim is not correct. The paper is therefore promising but needs a substantial clarification and, in places, re-statement.","major_comments":[{"comment":"The proof of Theorem 5 asserts that because each I_i is strictly increasing in SNR, minimizing max_i I_i(ρ_i) is equivalent to minimizing max_i ρ_i. This equivalence is valid only when all subchannels have the same MI-versus-SNR function I(ρ), i.e., the same input distribution (up to power scaling). The manuscript never states this common-distribution assumption; Section V says only that 'the input signals follow an arbitrary distribution.' Under the natural heterogeneous reading, the theorem is false: for m=2 with P_i=1, Z_i=0, N_0=0.1, one Gaussian and one binary subchannel, formula (36) gives N_1=N_2=0.05 and max MI = 0.5 ln(1+20) ≈ 1.522 nats, whereas N_1=0.1, N_2=0 gives max(0.5 ln(1+10), ln 2) ≈ 1.199 nats, a strictly better allocation. Please add the common-input-distribution assumption explicitly and either restrict Theorem 5 to that case or provide the correct heterogeneous generalization.","section":"Section V-D, Theorem 5 and Appendix F"},{"comment":"The same implicit common-distribution assumption underlies Theorem 4. The notation mmse(P_i/(N_i+Z_i)) and I(ρ_i) in (31)-(32) presumes that every subchannel has the same MMSE function; if subchannels have different input distributions, the stationarity condition must be P_i/(N_i+Z_i)^2 · mmse_i(P_i/(N_i+Z_i)) = ν, with per-subchannel mmse_i. As written, the theorem's claim for 'arbitrary input distributions' is too strong. Please state in the model that all subchannels share a common normalized input distribution, or derive the per-subchannel version.","section":"Section V, Theorems 3-4"}],"minor_comments":[{"comment":"The sentence 'the second follows from the chain rule' is not accurate; the inequality I(X^m;Y^m) ≤ sum_i I(X_i;Y_i) follows from the parallel-channel structure (conditional independence of Y_i given X_i) together with subadditivity of entropy. Please correct the justification.","section":"Equation (13)"},{"comment":"The displayed second derivative is off by a factor of P; the correct expression is (1/(2(N+Z)^2))(1 - J - ρ dJ/dρ). Since P>0, the sign condition (C2) is unaffected, but the formula should be corrected.","section":"Appendix C, Eq. (58)"},{"comment":"Figures 2 and 3 appear identical, as do Figures 4 and 5; if this is not intentional, the correct panels for the two input-power distributions should be provided.","section":"Figures 2-5"},{"comment":"Figures 2-6 report no error bars or number of trials, and the percentage savings are stated as precise numbers despite the random sampling of P_i; please provide a fixed seed, or confidence intervals over multiple trials, so the results are reproducible.","section":"Section VI-A"},{"comment":"In the AES case study, the allocation rules from Theorems 1 and 2 are derived for Gaussian inputs, but the real trace intermediate values are not necessarily Gaussian; a sentence explaining why this approximation is appropriate would strengthen the empirical claim.","section":"Section VI-B"}],"recommendation":"major_revision","confidential_remarks":"The theoretical content is largely sound once the common-input-distribution assumption is made explicit; the main risk is the overbroad claim in Theorem 5. I would also check whether the duplicated figures in Section VI are a production error before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The Gaussian-case results (Theorems 1 and 2) are correct and clean, and the dual water-filling interpretation is a nice way to present them. The arbitrary-input max-MI theorem (Theorem 5) is overclaimed, though: as written it only holds when every subchannel has the same normalized input distribution, and the paper never states that condition.\n\nThe problem is well posed, the I-MMSE machinery is used legitimately, and the convexity conditions in Theorem 3 are a real addition. The numerical comparisons against uniform allocation show consistent gains, and the AES-128 SPERO case study gives a concrete, reproducible-feeling check of the method.\n\nThe trouble is in Appendix F. It says minimizing max MI is equivalent to minimizing max SNR because MI is increasing in SNR. That equivalence is only valid when all subchannels share the same MI-vs-SNR function. The paper's own side-channel model allows each leakage point to have its own input distribution, and once distributions differ, equalizing MI values does not imply equal SNRs. One Gaussian and one binary subchannel at the same P and Z already breaks (36): the Gaussian dominates max MI, so the optimum shifts noise to the Gaussian until the two MI values match, rather than forcing equal SNR. The same issue underlies the injectivity argument after the theorem. This isn't cosmetic; it is the whole basis for the claimed distribution-independence.\n\nSmaller issues: the figures have no error bars or repeated-trial information, no code or data is released, and the abstract's wording ('arbitrary input distributions') is stronger than what the theorem actually delivers. The Gaussian part is unaffected and stands.\n\nBottom line: the paper deserves a serious referee. The Gaussian results are worth publishing even if Theorem 5 gets restricted. A conscientious referee should push for an explicit common-distribution assumption, or a corrected treatment of the heterogeneous case.","headline":"Solid Gaussian-case results, but Theorem 5's max-MI allocation for arbitrary inputs only holds when every subchannel has the same normalized input distribution—an assumption the paper never states.","tokens_in":19690,"tokens_out":5277,"would_cite":false,"duration_ms":51936,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["94A15","94A17","90C25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper derives closed-form optimal artificial-noise allocations for side-channel defense and shows that non-uniform injection beats uniform allocation.","keywords":["side-channel attacks","artificial noise injection","mutual information minimization","I-MMSE relation","Gaussian channel capacity","water-filling power allocation","AES-128 power analysis","key recovery success rate"],"falsifier":"Construct two parallel subchannels with equal signal power and equal physical-noise variance but different input distributions, for example Gaussian and binary, and compare the worst-case allocation $N_i=\\max(0,P_i/\\kappa-Z_i)$ against a brute-force numerical minimization of $\\max_i I_i(P_i/(N_i+Z_i))$ under the same budget; if the brute-force optimum gives a lower maximum mutual information, the claimed distribution-free optimality fails.","tokens_in":18681,"feed_emoji":"🔐","tokens_out":11324,"duration_ms":104967,"temperature":0.7,"pith_summary":"The paper claims that spreading artificial-noise power evenly across a side-channel trace is suboptimal. Viewing each sampled point as an independent Gaussian subchannel $Y_i = X_i + Z_i + N_i$ with signal power $P_i$, physical-noise variance $Z_i$, and injected-noise variance $N_i$, it minimizes either the total mutual information $\\sum_i \\log(1 + P_i/(N_i+Z_i))$ or the worst-case mutual information under a total-noise budget $N_0$. For Gaussian inputs the minimizers are closed-form: a dual water-filling rule for total leakage and an SNR-equalizing rule $N_i^* = \\max(0, P_i/\\kappa - Z_i)$ for worst-case leakage. For arbitrary inputs, convexity conditions are derived from the I-MMSE relation, and the worst-case rule is argued to be input-distribution-free. The practical payoff claimed is that the same security can be achieved with less injected noise, which matters for battery-powered devices.","feed_headline":"Optimal noise placement needs 27 percent less power to foil key recovery","feed_subtitle":"On AES power traces, uneven allocation preserves the key-recovery defense at up to 27 percent lower power.","key_machinery":"The load-bearing object is the I-MMSE identity $\\frac{d}{d\\rho} I(\\rho) = \\frac12\\,\\mathrm{mmse}(\\rho)$, which converts derivatives of mutual information with respect to SNR into MMSE expressions and makes the KKT conditions for noise allocation tractable even when the input distribution is non-Gaussian. Around it sit the convexity conditions of Theorem 3 -- $\\mathrm{mmse}(\\rho)+\\frac{d}{d\\rho}(\\rho\\,\\mathrm{mmse}(\\rho))\\ge 0$, equivalently $\\frac{d}{d\\rho}(\\rho J(Y))\\le 1$ or $\\mathbb{E}[(\\frac{d^2}{dY^2}\\log f(Y))^2]\\le 1$ -- and the data-processing upper bound $I(U;Y^m)\\le \\sum_i I(X_i;Y_i)$ that turns leakage minimization into a sum of independent subchannel problems. The optimality conditions in Theorems 1, 2, 4, and 5 are all KKT stationary points expressed through these objects.","core_discovery":"On the paper's own terms, the central claim is that information leakage through side channels can be treated as a power-allocation problem, and the optimal allocation of artificial noise is non-uniform. In the Gaussian model, minimizing the upper bound $\\sum_i I(X_i;Y_i)$ subject to $\\sum_i N_i \\le N_0$ is convex, and the KKT conditions give the dual water-filling allocation of Theorem 1: inject noise into subchannel $i$ until $1/(N_i^*+Z_i) - 1/(N_i^*+Z_i+P_i) = \\nu$, leaving $N_i^*=0$ when $\\nu \\ge 1/Z_i - 1/(Z_i+P_i)$. For the worst-case objective, Theorem 2 gives $N_i^* = \\max(0, P_i/\\kappa - Z_i)$, which equalizes the SNR $P_i/(N_i+Z_i)$ over the active subchannels. For arbitrary input distributions, Theorem 3 converts convexity into computable MMSE and Fisher-information conditions, Theorem 4 states the KKT condition $P_i/(N_i+Z_i)^2\\,\\mathrm{mmse}(P_i/(N_i+Z_i)) = \\nu$ for total leakage, and Theorem 5 claims the SNR-equalizing worst-case allocation is optimal for any input distribution. The AES-based experiments are offered as evidence that these allocations reduce both mutual information and key-recovery success rates relative to uniform noise at equal power.","pith_inferences":["Editorial inference: if the worst-case rule is genuinely distribution-free, a defender only needs per-point signal and physical-noise power estimates, not a statistical model of the leakage; this is a consequence the authors leave implicit.","Editorial inference: the same KKT machinery could allocate a single noise budget across heterogeneous leakage channels, such as power, electromagnetic, and timing, with each channel carrying its own MMSE function.","Editorial inference: the low-SNR argument implies that on very noisy devices the dual water-filling rule applies to nearly any input distribution, which could extend the method beyond cryptographic chips to analog sensors and RF emanation."],"forward_implications":["For Gaussian leakage points, the dual water-filling allocation strictly outperforms uniform allocation for a given budget, so total information leakage is reduced at no additional power cost.","The worst-case rule $N_i^*=\\max(0,P_i/\\kappa - Z_i)$ concentrates noise on the highest-SNR points and equalizes the residual SNR among the points that receive noise.","For arbitrary input distributions satisfying the convexity conditions, computing the MMSE function and using a one-dimensional bisection on the dual variable yields the global optimum for the total-leakage problem.","In the AES-128 experiments, keeping the key-recovery success rate at 0.1 needs roughly 9 percent less noise with the total-leakage rule and 27 percent less with the worst-case rule than uniform allocation.","Even when the total-leakage problem is non-convex, the worst-case problem can still be solved because the reformulated max-SNR problem is convex."],"supporting_citations":[{"why":"Supplies the uniform-allocation baseline and the Gaussian channel-capacity leakage model that the proposed allocation improves upon.","marker":"[3]"},{"why":"Supplies the I-MMSE relation used to differentiate mutual information with respect to SNR for arbitrary input distributions.","marker":"[20]"},{"why":"Supplies the Gaussian channel capacity formula, the water-filling framework, and Fano's inequality used to justify mutual information as a leakage proxy.","marker":"[23]"},{"why":"Supplies mutual information analysis, the attack metric used in the experiments and the worst-case maximum-MI notion.","marker":"[16]"},{"why":"Supplies the arbitrary-input MMSE results and the zero-mean reduction for subchannel power allocation.","marker":"[26]"},{"why":"Supplies the practical power-overhead figures for AES noise injection that motivate the energy-efficiency objective.","marker":"[13]"},{"why":"Supplies the AES-128 power traces used in the key-recovery case study.","marker":"[30]"}],"fun_headline_variants":["Optimal noise cuts SCA defense power by 27 percent","Nonuniform noise allocation saves 27% power in SCA defense","Water-filling noise injection cuts power 27% for SCA defense","Mutual info minimization yields power-efficient SCA shield","SNR-equalizing noise: 27% less power, same SCA defense"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that all protected side-channel points have the same mutual-information-versus-SNR curve, so equalizing mutual information is equivalent to equalizing SNR; if the per-point input distributions differ, the distribution-free worst-case formula may not be optimal.","fun_headline_variants_meta":{"raw":{"variants":["Optimal noise cuts SCA defense power by 27 percent","Nonuniform noise allocation saves 27% power in SCA defense","Water-filling noise injection cuts power 27% for SCA defense","Mutual info minimization yields power-efficient SCA shield","SNR-equalizing noise: 27% less power, same SCA defense"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001081,"raw_usage":{"total_tokens":4572,"prompt_tokens":1048,"completion_tokens":3524,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":3432}},"tokens_in":664,"tokens_out":3524,"duration_ms":23299,"temperature":1.0,"reasoning_tokens":3432,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:27:08.435717+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct two parallel subchannels with equal signal power and equal physical-noise variance but different input distributions, for example Gaussian and binary, and compare the worst-case allocation $N_i=\\max(0,P_i/\\kappa-Z_i)$ against a brute-force numerical minimization of $\\max_i I_i(P_i/(N_i+Z_i))$ under the same budget; if the brute-force optimum gives a lower maximum mutual information, the claimed distribution-free optimality fails.","supporting_citations":[{"cited_title":"Optima l energy efﬁcient design of artiﬁcial noise to prevent side-channel attacks,","cited_arxiv_id":null,"evidence_quote":"Supplies the uniform-allocation baseline and the Gaussian channel-capacity leakage model that the proposed allocation improves upon."},{"cited_title":"Mutual information an d minimum mean-square error in Gaussian channels,","cited_arxiv_id":null,"evidence_quote":"Supplies the I-MMSE relation used to differentiate mutual information with respect to SNR for arbitrary input distributions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian channel capacity formula, the water-filling framework, and Fano's inequality used to justify mutual information as a leakage proxy."},{"cited_title":"Mutu al information analysis: A generic side-channel distinguisher,","cited_arxiv_id":null,"evidence_quote":"Supplies mutual information analysis, the attack metric used in the experiments and the worst-case maximum-MI notion."},{"cited_title":"Optimum power al location for parallel Gaussian channels with arbitrary input distri butions,","cited_arxiv_id":null,"evidence_quote":"Supplies the arbitrary-input MMSE results and the zero-mean reduction for subchannel power allocation."},{"cited_title":"ASNI: Attenuated signature noise injection for low-overh ead power side-channel attack immunity,","cited_arxiv_id":null,"evidence_quote":"Supplies the practical power-overhead figures for AES noise injection that motivate the energy-efficiency objective."},{"cited_title":"SPERO: Simultaneous Power/EM Side-channel Dataset Using Real-time and Oscilloscope Setups","cited_arxiv_id":"2405.06571","evidence_quote":"Supplies the AES-128 power traces used in the key-recovery case study."}],"review_version":1}