{"id":"e84ce65a-4334-4151-912a-da8fd698636a","arxiv_id":"2607.24080","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On-site calibration of the RIS beam response model, via delay-domain multipath rejection and gradient-based fitting, reduces the simulated positioning error floor in RIS-aided localization.","lead":"This paper proposes a calibration framework that estimates a realistic 3D beam model for reconfigurable intelligent surfaces from on-site measurements, reducing the positioning error floor in RIS-aided localization. A two-stage algorithm separates the RIS-reflected signal from multipath, then fits beam parameters with gradient descent; validation uses measured 3D beam patterns.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation may be circular: §II.B says Eq. (11) is the generative model for the numerical results, conflicting with §V's measured-pattern claim; if literal, the reported BRS/ALB gains are self-consistency checks, not real-hardware evidence.","rationale":"I read the paper in good faith: the framework is coherent, the measured-pattern dataset is valuable, and the ALB metric is a reasonable way to quantify mismatch. The central claim is plausible. The load-bearing condition is that the numerical evaluation actually tests the estimated model against independently generated true beams. The paper's own sentence in §II.B ('We use the model in (11) as the generative beam model in the numerical results') directly contradicts §V.B's measured-pattern description. If the former is literal, the headline BRS/ALB results are circular. If the latter is literal, the headline BRS is still largely computed on the calibration grid; only the coarser 2°/4° runs provide partial held-out evidence, and they are not reported on strictly held-out directions. This is not evidence of misconduct; it is an unresolved validation gap. The concrete test above—trace the ground truth and, if measured, compute BRS on a held-out angular split—would settle it. Because the underlying method may well work, the appropriate verdict is unchanged: CONDITIONAL until the ground-truth source and out-of-sample performance are clarified.","tokens_in":17389,"tokens_out":21260,"duration_ms":200154,"concrete_test":"One check: obtain the simulation code/data (or ask the authors) and trace the exact construction of Fig. 6(d) and the ALB ground truth: is the beam pattern the interpolated measured anechoic pattern, or an evaluation of Eq. (11) with a chosen g(φ)? If it is Eq. (11), the validation is circular and the central claim is not supported. If it is measured, then as a second stage of the same check, re-run Step 2 on only the even-indexed azimuth/elevation calibration grid and compute BRS on the odd-indexed held-out grid; if held-out BRS is substantially below 88.5%, the reported number is an in-sample artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript never ties the validity of its headline numbers to a clearly identified ground truth. Section II.B, immediately after Eq. (11), says 'We use the model in (11) as the generative beam model in the numerical results.' If this is literal, then the ground-truth beams used for BRS (Fig. 6) and for the ALB simulations (Fig. 8) are synthesized from the very separable scalar-element-pattern model that Step 2 estimates; the 88.5% BRS and the 0.52→0.74 ALB improvement would then be in-sample fits of a model to its own generator, with no independent evidence about real RIS hardware. Section V.B instead states that measured anechoic 3D patterns were interpolated into the simulations; the two statements are never reconciled. Even taking the measured-pattern interpretation, the 1° BRS grid largely coincides with the CA sampling grid (az/el [-50,50]×[-50,10]), so the headline 88.5% is not an out-of-sample measure of how well model (11) predicts unsampled directions. The visible low-power sidelobe mismatch in Fig. 6(d) is exactly where the scalar-element-pattern assumption would fail. The central claim therefore rests on an evaluation whose ground-truth source and train/test separation are unresolved.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an on-site calibration framework for RIS beam responses in RIS-aided positioning systems. The method is two-stage: (i) a delay-domain sparse-recovery stage extracts the RIS-reflected channel response from signals collected by a calibration agent, using noncoherent aggregation across codewords, delay refinement, and geometry-assisted path identification; (ii) a gradient-based alternating estimator fits the separable beam model b(φ)=g(φ)W^H a(φ) of Eq. (11) to the extracted responses, estimating the effective codebook W and the diagonal gain-pattern matrix Γ. The authors validate the approach by integrating measured 3D beam patterns from a 16×16 RIS prototype under 66 phase configurations into simulations, reporting an average beam response similarity (BRS) of 88.5% for the calibrated model versus 43.7% for the ideal model, and an increase in the probability that the absolute lower bound (ALB) is below 0.5 m from 0.52 to 0.74. The positioning-oriented formulation and the use of measured patterns are valuable, but the validation as written has unresolved circularity and in-sample-evaluation concerns that bear directly on the headline claims.","tokens_in":17703,"tokens_out":7459,"duration_ms":67203,"significance":"If the reported gains are genuine, the contribution is significant: a practical, on-site calibration procedure that works under the RIS's default codebook, explicitly accounts for multipath, and improves the positioning error floor is of clear interest to the RIS-aided localization community. The gradient derivations in Eqs. (36)–(40) are algebraically consistent, the Step 1 design of aggregating delay-domain power over codewords is sensible, and the measurement campaign under 66 phase states is a genuine empirical asset. However, the paper does not ship code, proofs of identifiability/convergence, or a clearly separated training/test evaluation. The central validation evidence is currently compromised by the ambiguity between the generative-model statement in §II.B and the measured-pattern statement in §V.B, and by the fact that the BRS evaluation grid essentially coincides with the calibration sampling grid. These issues must be resolved before the headline numbers can be accepted as evidence about real-hardware behavior.","major_comments":[{"comment":"Immediately after Eq. (11), the paper states 'We use the model in (11) as the generative beam model in the numerical results.' Section V.B instead says that measured complex 3D beam patterns were incorporated into the simulations by replacing the ideal model. These two statements describe different ground truths. If (11) is literally the generative model, then the ground-truth beams used for the BRS in Fig. 6 and the ALB in Fig. 8 are synthesized from the same separable scalar-element-pattern model that Step 2 estimates; the 88.5% BRS and the 0.52→0.74 ALB improvement would then be in-sample fits of the model to its own generator, not evidence about real RIS hardware. If the simulations use measured patterns, that must be stated unambiguously and the train/test split must be specified. The manuscript cannot be interpreted as it stands, and the headline claims rest on this unresolved ambi","section":"§II.B and §V.B (ground-truth inconsistency)"},{"comment":"Equation (45) defines BRS on the S calibration samples, while footnote 4 says the reported BRS is evaluated on a dense 1-degree grid over [−50,50]° azimuth and [−50,20]° elevation. The calibration sampling step is 1 degree over [−50,50]°×[−50,10]°, so the evaluation grid essentially coincides with the fitting grid (with only the 10–20° elevation strip being novel). Thus the BRS largely measures in-sample fit, not generalization to unsampled directions. The visible low-power sidelobe mismatch in Fig. 6(d) is exactly where the scalar-element-pattern assumption of Eq. (11) would fail, and this is hidden by evaluating on the fitting grid. The authors should report BRS on directions not used for calibration, e.g., by holding out a subset of CA positions or using an evaluation grid finer than the calibration grid, and should separately report the match in low-power sidelobe regions.","section":"Eq. (45) and footnote 4 (in-sample BRS)"},{"comment":"The optimization in Eq. (33) is a non-convex factorization problem. The counting condition in Eq. (14) is only necessary and does not account for the unit-norm constraints or the fact that W and Γ are not jointly identifiable without further conditions. No identifiability analysis or convergence guarantee is provided for the alternating gradient/least-squares scheme; the gradient update in Eq. (36) with an unspecified learning rate has no guarantee of reaching a global or even local optimum. Since the central claim is that the calibrated model is accurate enough to reduce the positioning error floor, the paper should either provide identifiability and convergence results for the alternating scheme, or at least include a sensitivity analysis over random initializations and learning rates to demonstrate that the reported BRS/ALB values are not initialization- or step-size-dependent.","section":"§III.B, Eq. (33) (identifiability and convergence)"},{"comment":"The algorithm depends on several hyperparameters that are not specified in Table II or the text: the learning rate l_r in Eq. (36), the regularization constant ϵ_0 in Eq. (40), the stopping thresholds I_max and ϵ_res in Eq. (24), the delay-grid size N_τ in Eq. (19), the path-identification tolerance ε in Eq. (27), and the number of epochs N_ep in Section III-C. Without these values, the reported BRS and ALB curves cannot be reproduced. At minimum, the authors should list the values used and report sensitivity to the most critical parameters (especially l_r, N_τ, and ε).","section":"Table II and §III (unspecified hyperparameters)"}],"minor_comments":[{"comment":"The notation [\\bar B]_{g,:} is used for the ground-truth beam response, but the definition of \\bar B is not made explicit in Section IV.A. Clarify that \\bar B is the measured/interpolated ground-truth matrix (or the synthetic matrix if the generative model is intended).","section":"Eq. (45)"},{"comment":"Table I lists the measured elevation range as [−50°, 0°], while Table II and Section V.B use an elevation coverage of [−50°, 10°]. Reconcile these ranges and state which one is used for calibration sampling and which for BRS evaluation.","section":"Table I vs Table II"},{"comment":"The ALB in Eq. (52) is defined as the norm of the bias term only. Please state explicitly that this is a lower bound on RMSE due to model mismatch and not the full MSE or MCRB.","section":"Eq. (52)"},{"comment":"The Hadamard division in Eq. (15) assumes the pilot matrix X has no zero entries. State the pilot-sequence assumption or use a regularized division.","section":"Eq. (15)"},{"comment":"In the notation list, 'a◦b' is said to denote outer product, while Eq. (10) uses '⊙' for element-wise product. Please double-check the symbols so that the element-wise model in Eq. (10) is unambiguous.","section":"Notation"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unresolved contradiction between the generative-model statement in §II.B and the measured-pattern validation in §V.B. If the authors confirm that all headline numbers come from measured patterns and provide a proper out-of-sample BRS evaluation, the paper could be a solid contribution. The novelty relative to the authors' prior work [7] should also be clarified: the present paper adds multipath handling, a 3D model, and measured 3D patterns, but the beam-model core is an extension of [7]."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Friend,\n\nThe calibration framework in this paper deserves a serious look, but the validation as written has a load-bearing contradiction that has to be resolved before the headline numbers mean anything.\n\nWhat's genuinely good: the two-stage structure is sensible — noncoherent aggregation across codewords followed by delay refinement and geometry-based path identification, then alternating gradient descent and closed-form least squares for a separable 3D beam model. The complexity analysis is transparent. The authors also put in real measurement work: anechoic 3D patterns for a 16x16 RIS under 66 phase configurations, which is more evidence than most papers of this type provide. The ALB derivation is standard and properly isolates the bias term that causes the error floor.\n\nThe soft spots, in proportion:\n\nFirst and most important: §II.B says \"We use the model in (11) as the generative beam model in the numerical results,\" while §V.B says the simulations use measured beam patterns to generate received signals. These cannot both be true. If (11) is the generator, the 88.5% BRS and the 0.52-to-0.74 ALB improvement are self-consistency checks, not hardware validation. If the measured patterns are the generator, then evaluating BRS on a 1° grid that matches the calibration sampling grid is still mostly in-sample, and the low-power sidelobe mismatch in Fig. 6(d) is exactly where the separable scalar-element-pattern model would break. The paper never reconciles these. This is the central issue.\n\nSecond, hyperparameters are unspecified: learning rate, max epochs, thresholds, delay-grid sizes, and the RCS values. No code or data release either, so reproducibility is limited.\n\nThird, the non-convex factorization in (33) has no identifiability or convergence guarantees beyond the unit-norm constraint, which removes one scale ambiguity but leaves the question open.\n\nThe ALB simulation does use a different UE geometry than the calibration positions, so that part has some out-of-sample flavor, but it's still a simulation over the same fitted model.\n\nRecommendation: send it to peer review but ask the authors to state explicitly what ground truth generated the results, report BRS on an independent angular grid (e.g., a sparse calibration grid interpolated to a dense grid, or a held-out angular sector), and release at least the hyperparameters and preferably the measured patterns. The core idea is plausible and worth developing; the current evidence doesn't yet support the headline gains.\n\nRegards.","headline":"The two-stage calibration framework is sensible and the measured data are a real asset, but the validation has a load-bearing contradiction about whether the headline numbers come from measured patterns or the paper's own generative model.","tokens_in":18219,"tokens_out":5576,"would_cite":false,"duration_ms":49375,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On-site RIS beam calibration can cut positioning error floor by replacing ideal beam models with a realistic model fit from a moving agent's measurements.","keywords":["RIS-aided positioning","beam calibration","reconfigurable intelligent surface","beam model mismatch","positioning error floor","delay-domain sparse recovery","absolute lower bound","3D beam pattern measurement"],"falsifier":"Measure a real RIS with strong per-element mutual coupling, fit the proposed model on a 1-degree angular grid, then evaluate the beam-response similarity on a held-out dense angular grid at 0.2-degree resolution; if the average similarity on held-out angles falls well below the reported 88.5% (or the ALB improvement vanishes), the separable model is the limiting factor rather than the estimation algorithm.","tokens_in":17254,"feed_emoji":"📡","tokens_out":4626,"duration_ms":36996,"temperature":0.7,"pith_summary":"This paper tries to establish that a practical 3D reconfigurable-intelligent-surface (RIS) beam response can be estimated on site, in the deployment environment, well enough to remove most of the positioning error caused by using ideal beam models. The authors argue that hardware impairments, mutual coupling, and non-ideal phase tuning distort real RIS beams, and that these distortions can be captured by a product-form beam model whose parameters are fitted from signals received by a calibration agent at known locations. They propose a two-stage algorithm: first extract the RIS-reflected path from the received signal while suppressing line-of-sight and multipath components, then fit the model with alternating gradient-based optimization. Using measured 3D beam patterns from a 16x16 RIS prototype under 66 phase configurations, they report that calibration raises average beam-response similarity from 43.7% (ideal model) to 88.5%, and raises the probability that the absolute lower bound on positioning error stays below 0.5 m from 0.52 to 0.74. If correct, this would let deployed RIS-aided positioning systems operate near their information-theoretic accuracy without anechoic-chamber calibration.","feed_headline":"On-site calibration raises RIS beam match from 44% to 88.5%","feed_subtitle":"Two-stage estimator lifts chance of sub-0.5 m positioning error from 0.52 to 0.74.","key_machinery":"The load-bearing object is the separable product-form beam model b(phi)=g(phi) W^H a(phi), where a(phi) is the array steering vector, W is the effective codebook matrix absorbing mutual coupling and non-ideal phase tuning, and g(phi) is a common scalar element pattern. For estimation, the model is written as B=W^H A(Phi) Gamma, with Gamma a diagonal matrix absorbing the RIS-path channel gain and element-pattern response. The argument runs through two stages: a delay-domain sparse-recovery stage that uses noncoherent aggregation over codewords, high-resolution delay refinement, and known geometry to isolate the RIS-reflected path from the line-of-sight path and multipath; then an alternating","core_discovery":"The central claim is that the mismatch between ideal and true RIS beams—not noise or geometry—is the dominant source of the positioning error floor, and that this mismatch can be largely removed by a calibration procedure that estimates the practical beam response directly. Concretely, the paper shows that fitting the model b(phi)=g(phi) W^H a(phi) (with the RIS-path gain absorbed into a diagonal matrix Gamma) to on-site measurements yields a beam representation that agrees with ground truth at 88.5% average beam-response similarity, compared with 43.7% for the ideal model. When the calibrated model is inserted into the positioning estimator, the probability that the absolute lower bound (AL","pith_inferences":["The paper evaluates beam-response similarity on the same angular sampling grid used for fitting; a stronger test would hold out angles or randomize CA positions to measure generalization to off-grid directions, where low-power sidelobes are hardest to reproduce.","If per-element mutual coupling or edge effects break the scalar-element-pattern assumption, the product-form model cannot represent angle-dependent element responses; a testable extension is to fit separate element patterns per row/column or add a residual correction term and compare BRS on dense off-grid measurements.","The measured BRS already shows residual discrepancies in low-power sidelobes, so the method's practical gain for positioning may be concentrated in the main lobe and strong sidelobes; applications relying on weak reflections (e.g., long-range sensing) may need denser or higher-fidelity calibration.","A natural next step, noted implicitly by the far-field scope, is near-field calibration: at close range the plane-wave steering vector fails, and the fitted model would need a near-field correction; the same two-stage extraction and alternating fitting could be adapted to that regime."],"forward_implications":["Positioning estimators can replace the ideal RIS beam model with the calibrated model, removing the systematic bias that otherwise caps accuracy.","The calibration works with the RIS's default codebook and standard OFDM signals, so no dedicated phase-tuning settings are needed during calibration.","Reducing the angular sampling step from 1 deg to 2 or 4 deg retains most of the positioning benefit (ALB below 0.5 m with probability around 0.72).","Wider signal bandwidth improves the accuracy of extracting the RIS-reflected path, which feeds directly into the calibrated beam model.","The method scales linearly with the number of calibration samples, codewords, and RIS elements, making it applicable to large RIS deployments."],"fun_headline_variants":["On-site calibration lifts RIS beam match from 44% to 88.5%","Calibrated RIS beams raise sub-0.5m positioning odds to 0.74","RIS beam mismatch reduced: calibration yields 88.5% match","Practical beam calibration doubles RIS match, eases error floor","From 44% to 88.5%: on-site calibration sharpens RIS beams"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The true beam is assumed to factor as a common scalar element pattern times a codebook-weighted steering vector, so any per-element coupling or edge effect that breaks this separability will leave the calibrated model unable to represent the low-power sidelobes.","fun_headline_variants_meta":{"raw":{"variants":["On-site calibration lifts RIS beam match from 44% to 88.5%","Calibrated RIS beams raise sub-0.5m positioning odds to 0.74","RIS beam mismatch reduced: calibration yields 88.5% match","Practical beam calibration doubles RIS match, eases error floor","From 44% to 88.5%: on-site calibration sharpens RIS beams"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000438,"raw_usage":{"total_tokens":2085,"prompt_tokens":789,"completion_tokens":1296,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1192}},"tokens_in":533,"tokens_out":1296,"duration_ms":10026,"temperature":1.0,"reasoning_tokens":1192,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:03:45.325931+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure a real RIS with strong per-element mutual coupling, fit the proposed model on a 1-degree angular grid, then evaluate the beam-response similarity on a held-out dense angular grid at 0.2-degree resolution; if the average similarity on held-out angles falls well below the reported 88.5% (or the ALB improvement vanishes), the separable model is the limiting factor rather than the estimation algorithm.","supporting_citations":[],"review_version":1}