{"id":"94786a2b-204c-4fae-8950-efff4ff362c6","arxiv_id":"2608.09676","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Synthetic RF sensing samples can fail measurement consistency even when they pass task checks, and a calibrated audit (RFCheck) can expose and reduce this failure.","lead":"This paper shows that synthetic Wi-Fi and radar sensing samples can pass standard checks (classifier acceptance, summary distances) yet still violate the physical structure of real measurements. It introduces RFCheck, a calibration-based audit that flags such samples and helps repair or select safer synthetic data before augmentation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Measurement tests are validated only against perturbations built from the same physical intuitions; without an independent ground-truth check, 'measurement inconsistency' may just mean deviation on hand-picked statistics.","rationale":"RFCheck's contribution is to make a specific failure mode measurable. The entire empirical argument—controlled perturbations, label-fixed retention, repair/correction—uses the TD/FD/RD scores both as the definition of the failure and as the evaluation metric. Controlled perturbations show the tests respond to the perturbations, but because the perturbations are generated from the same physical premises as the tests, this is in-sample validation, not independent construct validation. The correction/repair results are even more directly circular: the objective minimized (L_meas, Eqs. 12-13) is the same quantity reported as 'flagged ratio.' The only non-circular evidence is task-performance preservation, which is necessary but not sufficient to establish that the audit measures measurement consistency rather than arbitrary distributional distance. The FMCW subject-disjoint result—real samples flagged at elevated rates—reinforces this: the score cannot distinguish 'synthetic pipeline violation' from 'real but different acquisition condition.' This does not refute the paper; the controlled experiments and honest limitations make the failure mode plausible. But the claim as stated in the abstract is stronger than what the current validation establishes. The reader's CONDITIONAL verdict is appropriate; my concern adds a specific condition: independent ground-truth validation of the measurement tests. Therefore verdict_should_be is UNCHANGED, and agreement with the reader's weakest assumption is partial.","tokens_in":15846,"tokens_out":11654,"duration_ms":117322,"concrete_test":"Construct an independent ground-truth validation from the observation model in Eq. (5). Generate synthetic CSI samples by sampling H(f_k) with known measurement-chain violations—e.g., nonzero energy on unoccupied subcarriers before occupied-tone embedding, or a missing RF phase-noise term—and generate a matched set of samples from the same model without violations. Run the published RFCheck TD/FD pipeline on both sets and compute the AUC for separating violated from non-violated samples. If the AUC is not near 1, the tests are not faithful to the stated measurement model, and the paper's flags cannot be interpreted as evidence of measurement inconsistency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that synthetic RF samples can 'fail measurement consistency' rests on the construct validity of the TD/FD (and FMCW range-Doppler) tests. This is the least secure link. The tests are validated mainly by showing they respond to controlled perturbations (TD-canonical, FD-canonical, Local break) and by showing that correction reduces the same scores that appear in the training objectives (Eqs. 12, 13, 17, 18). Both validations are partially circular: the perturbations are designed from the same finite-bandwidth/phase-continuity intuitions encoded in the tests, and the reduction metrics are the losses being optimized. The paper does not compare the audit's flags against an independent, physically grounded definition of measurement inconsistency, such as energy in unoccupied subcarriers or phase discontinuity across the occupied-tone embedding in Eq. (5). The FMCW subject-disjoint observation that real held-out samples are also flagged at elevated rates further suggests the score responds to generic distribution shift, not specifically to a violation of the measurement pipeline. If the tests are too narrow, inconsistent synthetic data can pass; if they are too broad, benign real variability is called a failure. The paper's own limitations paragraph treats the thresholds as benchmark-specific, but the validity of the tests themselves is the load-bearing assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper identifies a failure mode in synthetic RF sensing data: a synthetic sample may pass common task-facing checks (selected summary statistics, label consistency) while deviating from the measurement behavior of real samples produced by the same acquisition and preprocessing pipeline. The authors propose RFCheck, a calibrated audit that computes representation-specific measurement tests (time-domain precursor leakage and frequency-domain local continuation for CSI; range leakage, range-Doppler continuity, and receive-chain temporal residual for FMCW), calibrates per-test thresholds on held-out real data, and flags samples whose joint score exceeds the calibrated range. They then use the audit for retention (low-risk vs. high-risk candidate selection), repair (direct score-guided optimization), and correction (a learned module mapping proposals toward repair-reference targets). Experiments on Widar, ARIL, WiMANS, and M-Gesture show that aggregate summaries and label filters can miss the detected failures, that measurement risk separates downstream behavior under a fixed label-consistency budget, and that correction can produce low-flagged, class-balanced candidate sets in some settings while remaining ineffective for weak proposals and for preserving FMCW motion structure.","tokens_in":16110,"tokens_out":3543,"duration_ms":35233,"significance":"If the central claim holds, RFCheck is a useful complement to existing synthetic-data validation: it provides sample-level, representation-specific, calibration-based audit scores rather than aggregate distribution distances, and it is instantiated in two different RF sensing modalities. The paper is careful to hedge several claims: thresholds are benchmark-specific, the joint score is a conservative union-style detector rather than an exact alpha-level test, and correction is proposal-dependent. The controlled perturbation experiments, label-fixed retention study, and five-seed correction runs are valuable evidence. The main weakness is that the construct validity of the measurement tests is established only against perturbations built from the same physical intuitions encoded in the tests, and the correction experiments are partly circular if the repair-reference targets overlap with the calibration set. These issues are load-bearing for the central claim and need to be addressed with additional experiments and explicit split guarantees.","major_comments":[{"comment":"The paper's central claim that synthetic samples 'fail measurement consistency' rests on the construct validity of the TD/FD (and FMCW range-Doppler) tests. As presented, the tests are validated mainly by showing that they respond to controlled perturbations (TD-canonical, FD-canonical, Local break) and that correction reduces the same scores that appear in the training objectives (Eqs. 12, 13, 17, 18). Both validations are partially circular: the perturbations are designed from the same finite-bandwidth/phase-continuity intuitions encoded in the tests, and the reduction metrics are the losses being optimized. I ask for an independent, physically grounded check of measurement inconsistency, for example energy in unoccupied subcarriers or phase discontinuity across the occupied-tone embedding in Eq. (5), and for a demonstration that the tests reject perturbations designed from a different physical mechanism. The subject-disjoint FMCW observation that real held-out samples are also flagged at elevated rates further suggests that the score may respond to generic distribution shift; the paper should provide a criterion that separates 'measurement inconsistency' from 'domain shift'.","section":"Section III.A and Section V.A"},{"comment":"The finite-sample false-exceedance theorem (Theorem 1) applies to a future real sample exchangeable with the calibration samples. Synthetic samples are not exchangeable with real calibration samples, so the calibrated threshold is not a controlled probability statement for synthetic data; exceedance is only a descriptive benchmark. The paper acknowledges this in the text ('benchmark-specific empirical reference'), but the central claim requires showing that threshold exceedance is a meaningful signal of a pipeline violation rather than an artifact of limited calibration sample size or non-representative calibration. Please report the calibration set size, bootstrap confidence intervals for the thresholds, and the held-out real exceedance rate under the same split used for synthetic evaluation. If held-out real exceedance is much larger than alpha under subject-disjoint folds (as in FMCW), the paper should state explicitly where the benchmark-specific reference ends and the measurement-consistency claim begins.","section":"Section III.A, Eqs. (1)-(4)"},{"comment":"The correction module is trained to map proposals toward repair-reference targets x_ref, and the repair reference is described as an 'offline intervention' using real data. It is not stated whether the real samples used as x_ref are disjoint from the calibration set D_m^cal, from the real samples used to construct controlled perturbations, and from the real training/test splits used for task evaluation. If x_ref samples overlap with the calibration set, then the reduction in RFCheck score after correction is partly by construction, because the target is in-distribution for the very thresholds being audited. Please state the exact split relationships, and if overlap exists, rerun the correction experiment with x_ref drawn from a separate real partition that is disjoint from calibration and task evaluation. This is necessary to interpret the 'full direct repair' and 'held-out correction selected' rows in Table I.","section":"Section IV.C, Eqs. (15)-(16), and Table I"},{"comment":"The FMCW correction-selection row reports a trajectory gap of 0.9475 and a fallback ratio of 0.7238, which the paper interprets as 'correction cannot recover coherent range-Doppler motion.' This is an honest limitation, but it undercuts the general claim stated in the abstract and conclusion that correction 'reduces measurement violations while preserving task behavior.' The FMCW evidence supports measurement-risk reduction only, not task-preserving correction. Please either temper the cross-representation wording to reflect that task-preserving correction is demonstrated only for CSI, or provide an FMCW configuration where task behavior is preserved without a large fallback ratio.","section":"Table II and Section V.E"}],"minor_comments":[{"comment":"The union bound in Eq. (4) is correct but loose; since the paper already interprets the joint score conservatively, a sentence noting that the bound is not tight and that exact joint control is not claimed would be helpful.","section":"Section III.A, Eq. (4)"},{"comment":"The held-out correction row reports 'no selected sample is flagged' after selecting by the RFCheck score; this is expected by the selection rule. Please clarify that the row supports feasibility of the screening workflow rather than serving as an independent test of the audit.","section":"Section IV.C and Table I"},{"comment":"The caption says 'bars report flagged ratios' but the bar panel is small; consider reporting the numeric flagged ratios in the text or a table so the reader can verify the claim that tails are captured.","section":"Section V.A, Fig. 3 caption"},{"comment":"The figure legend uses 'JPI' while the text and Table I use 'joint score'; unify the terminology to avoid confusion.","section":"Section V.C, Fig. 6"},{"comment":"The paper does not mention code or data release. Given the reproducibility value of the calibration workflow and the controlled perturbation constructions, a statement about availability would be appropriate.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central idea is timely and the experimental structure is generally sound, with honest hedging about benchmark-specific thresholds and proposal-dependent correction. The main risk is construct validity: the measurement tests are validated against perturbations and objectives built from the same physical intuitions, and the split hygiene between calibration, repair-reference, and task evaluation is not fully documented. I recommend asking the authors to provide an independent ground-truth check of measurement inconsistency, explicit split guarantees, and a calibration-size sensitivity analysis before acceptance. If these are provided, the paper could become a strong contribution to synthetic RF sensing validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing to know: this paper has a real and useful observation — synthetic RF samples can pass label-consistency checks and summary statistics while still violating the measurement structure of the real pipeline, and a calibration-based audit can expose this. The CSI experiments make that point convincingly. The bigger claim, that a physical audit is the right validation layer, is plausible but rests on a construct-validity assumption the paper doesn't fully close.\n\nWhat's new: the failure mode itself is not previously documented in RF sensing, and RFCheck is a coherent, simple tool. The finite-sample calibration argument is standard exchangeability, honestly labeled, and the paper correctly notes the joint score is a conservative union bound rather than an exact alpha-level test. Credit also for the label-fixed retention study: holding label pass constant and showing low-risk vs high-risk candidates behave differently downstream is the right way to isolate measurement risk from label error. The repair/correction experiments are a good-faith test of actionability, and the paper explicitly states their limits.\n\nSoft spots, in order of softness. First, the measurement tests (TD precursor leakage, FD continuation, range-Doppler) are validated mostly against perturbations built from the same physical intuitions. There's no independent ground-truth check — e.g., energy in unoccupied subcarriers or phase discontinuity across the occupied-tone embedding from Eq. (5). So 'measurement inconsistency' is, strictly, 'deviation on hand-picked statistics.' That's not fatal, because the statistics are reasonable and the paper is careful to call them benchmark-specific references, but it means the audit's coverage is unknown. Second, the intervention results have wide confidence intervals, especially worst-class F1, and the FMCW extension shows a large fallback ratio and trajectory gap. The paper is honest about this, but it means the correction claims are supported only in the 'refined proposal pool' regime. Third, no code or data release, and hyperparameters like lambda weights and margin tau are not reported; reproducibility is incomplete.\n\nThe paper deserves a serious referee. The central finding matters for anyone using synthetic data in wireless sensing, and the CSI evidence is strong enough to warrant scrutiny. I'd want the referees to push on construct validity and ask for code/data. My own verdict is conditionally positive: the audit is a useful tool, not a universal validator.\n\nRecommendation: send to peer review, but require the authors to either provide an independent check of the measurement tests or temper the 'measurement consistency' language. It's not a desk reject.","headline":"A useful, honestly-scoped audit for synthetic RF data, with a real failure mode demonstrated on CSI, held back by construct-validity and reproducibility gaps.","tokens_in":16641,"tokens_out":2001,"would_cite":true,"duration_ms":17984,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A calibrated audit catches synthetic RF data that pass task checks but violate real measurement structure.","keywords":["RF sensing","synthetic data","measurement consistency","channel state information","calibrated measurement audit","data augmentation","FMCW radar","synthetic shortcuts"],"falsifier":"Compute the RFCheck score on a large held-out set of real samples that were not used for calibration; if the flagged ratio substantially exceeds the nominal level (for example, far above 5 percent at $\\alpha=0.05$), the finite-sample calibration claim would be violated. A second, more semantic test: take a synthetic pool whose samples all pass RFCheck, train on it, and check whether the trained model relies on features absent from real measurements; if such a model performs as well on real test data as a model trained on real data alone, the claim that measurement consistency is a necessary precondition for useful augmentation would be weakened.","tokens_in":15610,"feed_emoji":"📡","tokens_out":5636,"duration_ms":45972,"temperature":0.7,"pith_summary":"This paper aims to establish a specific failure mode for synthetic radio-frequency (RF) sensing data: a synthetic sample can look right to a classifier and close under selected summary statistics while still violating the structure that real measurements impose, because it was not produced by the same acquisition and preprocessing pipeline. The paper argues that this measurement inconsistency matters because task models trained on such samples can learn synthetic artifacts instead of sensing behavior, biasing model selection. To make the failure observable, the paper introduces RFCheck, a calibrated measurement audit that compares each synthetic candidate against held-out real samples from the same pipeline and flags candidates whose representation-specific measurement responses exceed the calibrated real range. Experiments on Wi-Fi channel-state information and millimeter-wave FMCW radar gestures show that the audit exposes failures that aggregate distances and label-consistency checks miss, and that the detected failures can be reduced by score-guided repair and correction when the proposal pool already contains task-relevant structure.","feed_headline":"Calibrated audit catches synthetic RF data that pass task checks","feed_subtitle":"Synthetic samples that pass label checks can still break real measurement structure; a calibrated audit catches and repairs them.","key_machinery":"The central object is the measurement-consistency score $S_m(x)=\\max_j M_{m,j}(x)/\\gamma_{m,j}^{(\\alpha_j)}$, where $M_{m,j}$ is the $j$-th representation-specific measurement test for representation $m$ and $\\gamma_{m,j}^{(\\alpha_j)}$ is the empirical $(1-\\alpha_j)$ quantile of test scores computed on held-out real calibration samples from the same pipeline. The max aggregation ensures that a strong abnormality in one test is not averaged away by the others. A finite-sample result (Theorem 1) bounds the probability that a future real sample exceeds the order-statistic threshold by $(n+1-k)/(n+1)\\le \\alpha$ under exchangeability, and the joint score is controlled by a union bound. For CSI the tests are time-domain precursor leakage and local frequency-domain continuation; for FMCW radar they are range leakage, range-Doppler continuation, and receive-chain temporal residual. This score is what carries the argument: it turns 'measurement consistency' into a calibrated, sample-level decision that can be used to rank, retain, repair, and correct synthetic candidates.","core_discovery":"On its own terms, the paper claims that measurement consistency is an independent validation axis for synthetic RF sensing data, distinct from label correctness and task accuracy. Concretely, a synthetic sample may pass common task-facing checks—selected summary distances and label-consistency acceptance—while deviating in localized ways from the empirical measurement behavior of real samples collected and processed by the same sensing pipeline. The paper instantiates RFCheck to make this deviation measurable: it defines measurement tests tied to the representation (precursor-like leakage in the delay domain and local continuation in the frequency domain for CSI; range leakage, range-Doppler continuation, and receive-chain temporal residuals for FMCW), calibrates each test threshold as an empirical quantile on held-out real samples, and flags any sample whose max-normalized joint score exceeds 1. In the reported experiments, label-consistent low-score and high-score candidates behave differently downstream, a repair reference lowers the flagged ratio to 10.83 percent while preserving mean task performance, and correction on a held-out proposal pool yields a class-balanced set with no flagged samples under a fixed training budget. The paper interprets the threshold as a benchmark-specific empirical reference, not a universal measurement boundary.","pith_inferences":["A straightforward extension would apply the same calibration workflow to other RF representations (ultra-wideband, mmWave MIMO channel tensors, or radar point clouds) by defining tests that match their measurement chains; the paper does not test these.","The union-bound interpretation suggests that if a user wants a joint false-alarm rate near $\\alpha$, the per-branch levels can be set to $\\alpha/J$, but the paper deliberately keeps per-branch calibration as a conservative empirical reference and does not tune the joint level.","If measurement consistency is a distinct axis, then synthetic-data releases could be accompanied by calibrated consistency certificates computed against a stated pipeline, which would make shortcuts harder to hide; the paper does not propose such a certification protocol.","A testable extension would check whether selecting on the calibrated score before augmentation reduces shortcut reliance in safety-oriented tasks, where worst-class behavior matters; the paper's worst-class results are suggestive but seed-level intervals are wide."],"forward_implications":["Aggregate summary distances (such as RBF-MMD on magnitude and phase-difference summaries) do not provide sample-level measurement localization, so they can leave localized violations undetected.","Under the same label-consistency acceptance rule and class-balanced budget, low-risk synthetic candidates improve worst-class F1 over high-risk candidates by about 0.063, showing that measurement risk is an independent selection axis beyond label correctness.","A score-guided repair reference lowers the flagged ratio to 10.83 percent while preserving mean task performance, so the detected failure is reducible, not just observable.","Correction of refined proposals reduces measurement risk consistently across five runs, and on a held-out pool of 1140 candidates the calibrated selection produces a class-balanced, label-consistent set of 120 with no flagged samples.","The same calibration principle transfers from CSI to FMCW millimeter-wave gesture sensing by replacing the measurement tests, indicating the failure mode is not specific to one RF representation."],"supporting_citations":[{"why":"Recent synthetic RF generation baseline that this paper contrasts with, supplying the generative-samples context.","marker":"[8]"},{"why":"The kernel two-sample test used as the selected-summary comparison that the paper shows to be insufficient for measurement consistency.","marker":"[18]"},{"why":"FMCW millimeter-wave gesture benchmark that provides the cross-representation validation setting and its range-Doppler structure.","marker":"[20]"},{"why":"CSI phase and measurement-chain model that motivates the delay-domain and frequency-domain tests.","marker":"[23]"},{"why":"Primary in-domain Wi-Fi CSI benchmark used for the controlled perturbations, retention, repair, and correction studies.","marker":"[24]"},{"why":"External Wi-Fi dataset used to check that the calibration principle transfers under different preprocessing.","marker":"[25]"},{"why":"External multi-user Wi-Fi benchmark used for a compact validation check of selection variants.","marker":"[26]"}],"fun_headline_variants":["RFCheck flags synthetic RF data that pass task checks","Synthetic RF data sneak past task checks, audit catches them","Measurement audit reveals synthetic RF data that fool task tests","RFCheck: beyond task checks to catch synthetic RF mismatches","Synthetic RF data pass task checks but fail measurement audit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The audit's thresholds are empirical quantiles computed on held-out real calibration samples from the same pipeline, so the whole argument assumes those calibration samples are representative enough that exceeding the threshold means true measurement inconsistency rather than calibration noise or benign distribution shift.","fun_headline_variants_meta":{"raw":{"variants":["RFCheck flags synthetic RF data that pass task checks","Synthetic RF data sneak past task checks, audit catches them","Measurement audit reveals synthetic RF data that fool task tests","RFCheck: beyond task checks to catch synthetic RF mismatches","Synthetic RF data pass task checks but fail measurement audit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":3029,"prompt_tokens":1059,"completion_tokens":1970,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":675,"completion_tokens_details":{"reasoning_tokens":1889}},"tokens_in":675,"tokens_out":1970,"duration_ms":11333,"temperature":1.0,"reasoning_tokens":1889,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:52:00.212955+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the RFCheck score on a large held-out set of real samples that were not used for calibration; if the flagged ratio substantially exceeds the nominal level (for example, far above 5 percent at $\\alpha=0.05$), the finite-sample calibration claim would be violated. A second, more semantic test: take a synthetic pool whose samples all pass RFCheck, train on it, and check whether the trained model relies on features absent from real measurements; if such a model performs as well on real test data as a model trained on real data alone, the claim that measurement consistency is a necessary precondition for useful augmentation would be weakened.","supporting_citations":[{"cited_title":"RF- Diffusion: Radio signal generation via time-frequency diffusion,","cited_arxiv_id":null,"evidence_quote":"Recent synthetic RF generation baseline that this paper contrasts with, supplying the generative-samples context."},{"cited_title":"M-gesture: Person-independent real-time in-air gesture recognition using commodity millimeter wave radar,","cited_arxiv_id":null,"evidence_quote":"FMCW millimeter-wave gesture benchmark that provides the cross-representation validation setting and its range-Doppler structure."},{"cited_title":"Spotfi: Decimeter level localization using wifi,","cited_arxiv_id":null,"evidence_quote":"CSI phase and measurement-chain model that motivates the delay-domain and frequency-domain tests."},{"cited_title":"Widar3.0: Zero-effort cross-domain gesture recognition with Wi-Fi,","cited_arxiv_id":null,"evidence_quote":"Primary in-domain Wi-Fi CSI benchmark used for the controlled perturbations, retention, repair, and correction studies."},{"cited_title":"Joint activity recognition and indoor localization with WiFi fingerprints,","cited_arxiv_id":null,"evidence_quote":"External Wi-Fi dataset used to check that the calibration principle transfers under different preprocessing."},{"cited_title":"WiMANS: A benchmark dataset for WiFi-based multi-user activity sensing,","cited_arxiv_id":null,"evidence_quote":"External multi-user Wi-Fi benchmark used for a compact validation check of selection variants."}],"review_version":1}