{"id":"2552a924-e679-4485-bcfb-c452d6b7d158","arxiv_id":"2412.08095","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A transformer network with MUSIC-based region gating estimates angle and time of arrival from 5G uplink channel data more accurately than CNN and MUSIC baselines under array impairments.","lead":"A 5G base station can find an indoor device by training a neural network to read phase differences in the signal across its antennas and frequency channels. The authors' transformer-based network, guided by an older spectral-search method, reports better angle and distance accuracy than convolutional networks in simulations and a real parking garage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never shows that synthetic training data reproduce the angle-dependent phase errors that motivate the method, and Eq. (5) models only constant per-antenna phase errors, so the claimed calibration gains may be artifacts of simulator error statistics.","rationale":"The load-bearing condition is the fidelity of the synthetic impairment model to the real angle-dependent array errors. The reader's weakest assumption pointed to simulator fidelity; I agree but sharpen it: Eq. (5) suggests the training model is even incapable of angle-dependent errors unless the simulator departs from the stated model. This is an internal inconsistency, not merely missing evidence. It does not force rejection because the field experiment is separate and could still support a weaker empirical claim, but it makes the stated interpretation 'accurately calibrating the irregular angle-dependent array error' unsupported. The one check would settle it. If the diagnostic fails, the verdict should be REJECT or UNVERDICTED for the calibration claim; if it passes, CONDITIONAL remains appropriate. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":7695,"tokens_out":6195,"duration_ms":63405,"concrete_test":"Obtain the simulator [9] (or the authors' generated dataset) and compute the per-antenna relative phase error as a function of AoA for the 20,000 synthetic training channels, then overlay these curves on Fig. 2. If the synthetic phase errors are approximately flat in AoA, the training set does not reproduce the angle-dependent impairment and the calibration claim collapses. If they do reproduce Fig. 2, retrain the proposed model with constant per-antenna phase errors only; if the two-expert/MUSIC gain over a single transformer persists, the reported improvement is not caused by angle-dependent calibration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the claim that the framework 'accurately calibrates the irregular angle-dependent array error,' the 20,000 synthetic training samples (Sec. 4) must contain array phase errors whose angle dependence matches the anechoic-chamber measurements (Sec. 2.2, Fig. 2), and that same impairment structure must appear in the field-test pico-cell gNB. The manuscript does not establish this. The signal model in Eq. (5) represents the impairment as Gamma ⊙ a, where Gamma is a diagonal matrix whose entries are 'the coefficient of the l-th radio frequency (RF) channel'; that is a per-antenna constant. The paper never states that Gamma depends on AoA or that the link-level simulator [9] injects the measured angle-dependent error pattern. Section 4 describes the simulator only as 'previously presented in our prior work [9]' and selects an indoor scene 'to align with the settings of our field test'; no comparison of simulated versus measured phase errors is provided. Without such a comparison, the two-region expert split and MUSIC gate (Secs. 3.2-3.3) may be fitting simulator-specific error statistics rather than calibrating real angle-dependent impairment. The field CDFs (Fig. 7) could also be explained by a generic ML mapping and do not isolate the calibration mechanism. If the simulated errors are flat in AoA, the motivation for the two-expert design disappears and the reported gains cannot be attributed to the proposed architecture.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a data-and-model-driven framework for joint angle-of-arrival (AoA) and time-of-arrival (ToA) estimation in monostatic 5G NR positioning under array impairments. The received uplink SRS is reshaped into a subcarrier-by-antenna CSI matrix, processed by a transformer network, and the angular space is partitioned into large-angle and small-angle regions, each served by an independently trained transformer expert. A MUSIC-based spectral search is used to decide which expert is applied at inference. The authors report that the proposed framework outperforms CNN and MUSIC baselines in simulated data and in a parking-lot field test with a commercial pico-cell gNB, and they claim that the method calibrates irregular angle-dependent array errors.","tokens_in":7891,"tokens_out":3581,"duration_ms":36445,"significance":"If the claims hold, the paper offers a practical, training-based alternative to explicit array calibration for monostatic 5G positioning, using standard SRS signals and a commercial gNB. The inclusion of anechoic-chamber measurements and an independent field test is a genuine strength and makes the central empirical claim credible in principle. However, the current manuscript does not supply enough evidence to substantiate the load-bearing assertion that the gains come from calibrating angle-dependent array errors: the signal model in Eq. (5) is a per-antenna constant, the simulator that generates all training data is not described, and no simulated-versus-measured error comparison is provided. The paper's contribution would be significant if these gaps are closed, but as written the evidence is under-supported.","major_comments":[{"comment":"The impairment model is inconsistent with the paper's motivating measurement. Eq. (5) defines Gamma as a diagonal matrix whose l-th diagonal entry is 'the coefficient of the l-th radio frequency (RF) channel', which is a per-antenna constant independent of AoA, whereas Sec. 2.2 and Fig. 2 report phase errors that vary with AoA and differ across antennas. The paper never states that Gamma depends on theta_d, nor that the simulator injects an angle-dependent error pattern. Without a model that couples Gamma to AoA, the theoretical development does not support the central claim of calibrating irregular angle-dependent array error.","section":"Sec. 2.1, Eq. (5) and Sec. 2.2"},{"comment":"The 20,000 synthetic training samples are generated by a link-level simulator [9] described only as 'previously presented in our prior work', and the indoor scene is chosen 'to align with the settings of our field test'. The manuscript provides no comparison of simulated array phase errors with the anechoic-chamber measurements in Fig. 2, and no description of how impairments are injected into the simulator. Since the two-expert architecture is motivated by those measured angle-dependent errors, the simulation results in Figs. 4-5 cannot be interpreted as validating the calibration mechanism unless a simulated-versus-measured phase-error comparison is supplied.","section":"Sec. 4, simulator description"},{"comment":"The covariance estimate is written as R = YY^H, where Y is the M-by-N (subcarrier-by-antenna) received matrix, which yields an M-by-M subcarrier covariance; yet the text states that a spectral peak search is performed 'solely within the angle dimension', which requires an N-by-N antenna covariance such as Y^H Y. This mismatch between the stated covariance and the angle-only MUSIC search makes the decision mechanism ambiguous, and the correct expression and dimensions should be clarified.","section":"Sec. 3.3, MUSIC region decision"},{"comment":"The field-test results in Fig. 7 are difficult to assess because the paper does not state how ground-truth AoA/ToA values are derived from the AGV trajectory, how many measurement points are used, what the SNR conditions were, or whether the deployed pico-cell gNB was characterized in the same anechoic chamber whose data appear in Fig. 2. In addition, no ablation isolates the MUSIC region decision: the comparison between transformer variants in Figs. 4-5 does not separate the benefit of the angle-region split from the benefit of the MUSIC selector, and a single transformer trained on the full angle range is not shown in Fig. 7. These omissions leave open the possibility that the reported gains are generic ML fitting rather than calibration of angle-dependent impairments.","section":"Sec. 4, field-test evaluation and ablations"}],"minor_comments":[{"comment":"The abstract and conclusion describe the network as 'transform-based', but the method is 'transformer-based'; the terminology should be unified.","section":"Abstract and Sec. 5"},{"comment":"The ordering of elements in the steering vector in Eq. (4) is not obviously consistent with the rearranged matrix in Eq. (6); the indexing convention should be clarified.","section":"Eq. (4) and Eq. (6)"},{"comment":"The phase error in Fig. 2 is measured relative to the first antenna, but the definition of this relative phase is never given; please state how the phase error is extracted from the CSI.","section":"Sec. 2.2, Fig. 2"},{"comment":"The text says '90% of the distance estimation errors falling within a range of 4 m', but Fig. 7(b) appears to plot ToA error; clarify whether distance error is ToA converted to meters and how the conversion is performed.","section":"Sec. 4, Fig. 7(b)"},{"comment":"The attention heatmap is discussed qualitatively; a quantitative metric, such as attention entropy or correlation with the phase-difference structure, would better support the claim that the model captures global features.","section":"Sec. 4, Fig. 8"},{"comment":"Reference [9] is cited as 'in press' with a 2023 DOI; verify the final publication status and update the citation accordingly.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a six-page workshop paper, and some of the under-specification likely stems from length constraints, but the missing simulator description and the Eq. (5)-versus-Sec. 2.2 mismatch are central to the claimed contribution and should be addressed before publication. I do not see a circularity problem in the deep-learning sense: training on synthetic data and testing on independent field measurements is a legitimate evaluation; the issue is that the synthetic data generation is not shown to reproduce the measured impairment characteristics. The requested additions are within the scope of a revision, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a plausible engineering paper from a 5G positioning group, and the field test is real, but the central calibration claim is under-supported because the training simulator's angle-dependent error model is never compared with the anechoic chamber measurements. I'd send it to a serious referee, but I'd expect major revision.\n\nWhat's actually new: combining a transformer with two angular-region experts and a MUSIC-based gate for joint AoA/ToA from 5G uplink CSI is something I haven't seen in the literature. The anechoic chamber phase-error measurement on a commercial 4-antenna pico-cell gNB is a useful empirical contribution, and the underground parking-lot field test with an AGV is a genuine effort to validate outside simulation. The attention visualization is a nice sanity check that the model isn't just copying nearest-neighbor structure.\n\nWhat's soft: the load-bearing claim is that the framework 'accurately calibrates the irregular angle-dependent array error.' But Eq. (5) models Gamma as a diagonal matrix of per-antenna RF coefficients, which is constant in angle. The paper never reconciles that with the measured angle-dependent behavior in Fig. 2. The link-level simulator [9] is cited but not described beyond saying it follows 3GPP TR 38.901 and that an indoor scene was chosen 'to align with the field test.' No comparison of simulated vs. measured phase errors is shown. If the simulator injects a simpler or flatter impairment, the two-region expert split could just be fitting simulator quirks, not the real physical effect. The field CDFs (Fig. 7) show broad improvement for all ML methods over MUSIC, which is consistent with a generic mapping benefit, not necessarily the calibration mechanism. There are also no error bars, no sample sizes listed beyond 'a collection of 500 data points' for simulations, and no ablation that isolates the MUSIC gate from the parallel-training split. No code or data release, so independent verification is impossible from the preprint.\n\nTo be fair, this is a 6-page workshop paper, and some of these gaps are space artifacts. The core direction is sound, and I don't think any of this is fabricated. But the paper's own framing goes beyond what the presented evidence supports.\n\nWho benefits: people working on 5G-based indoor positioning, particularly practical array-imperfection modeling. The anechoic measurement and the field-test methodology are worth citing even if the transformer results need more support.\n\nRecommendation: send to peer review, but require a comparison of simulated vs. measured phase-error statistics, an explicit angle-dependent impairment model (or a justification for the diagonal form), error bars, and an ablation of the MUSIC gate. Deserves a serious referee.","headline":"Plausible 5G positioning framework with a real field test, but the calibration claim is under-supported because the training simulator's impairment model is never checked against the measured angle-dependent errors.","tokens_in":8524,"tokens_out":2115,"would_cite":true,"duration_ms":19626,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A transformer network with MUSIC-based region selection can jointly estimate angle and time of arrival on a monostatic 5G base station, cutting positioning error below CNN and MUSIC baselines even when the antenna array has irregular…","keywords":["5G NR positioning","angle of arrival","time of arrival","transformer network","MUSIC","array calibration","monostatic sensing","channel state information"],"falsifier":"Measure the phase-error pattern of a second commercial gNB of the same model in the anechoic chamber, train the framework on the first gNB's error statistics (or on the simulator's error statistics), and test on the second gNB's field data; if the RMSE gain over the MUSIC baseline mostly disappears, the claimed calibration is specific to the simulated or chamber-measured error statistics rather than a general calibration capability.","tokens_in":7401,"feed_emoji":"📡","tokens_out":4678,"duration_ms":43163,"temperature":0.7,"pith_summary":"This paper tries to establish that a transformer-based neural network, trained separately on large- and small-angle regions and switched by a MUSIC-based decision rule, can jointly estimate angle of arrival (AoA) and time of arrival (ToA) on a monostatic 5G base station despite irregular angle-dependent phase errors in the antenna array. The authors argue that because the phase difference between any two subcarriers and antennas encodes target position, a transformer that attends globally is better suited than a CNN that only extracts local features. Anechoic chamber measurements reveal that phase errors vary with arrival angle, so the angular space is split and two expert transformers are trained in parallel. Simulated and underground-parking-lot field tests show lower AoA, ToA, and positioning errors than CNN and MUSIC baselines, with 90% of field positioning errors under 3 meters for the transformer-based methods.","feed_headline":"Transformer beats CNN and MUSIC on 5G localization with array errors","feed_subtitle":"A two-expert transformer, chosen by a MUSIC decision rule, keeps 90% of field positioning errors under 3 meters.","key_machinery":"The machinery is the self-attention operation, defined by the query, key, and value matrices in equations (7)–(9), applied to blocks of the reshaped CSI matrix $\\mathbf{Y}\\in\\mathbb{C}^{M\\times N}$, where rows and columns respectively encode the antenna and subcarrier phase relations of the rearranged steering-vector matrix $\\tilde{\\mathbf{A}}(\\theta_d,\\tau_d)$. Because the phase difference between any two blocks encodes AoA and ToA information, the self-attention map computes global correlations across all antenna–subcarrier blocks, which the paper claims is what lets the network fit the irregular angle-dependent array errors. Two transformers trained on disjoint angular regions, plus a MUSIC spectral-peak search in the angle dimension only (treating the subcarrier dimension as a frequency snapshot to estimate the covariance matrix from a single time snapshot), provide the region decision that selects the appropriate expert.","core_discovery":"The central claim is that the proposed data-and-model-driven framework—a dual-channel transformer whose input CSI matrix is reshaped so that rows carry antenna phase relations and columns carry subcarrier phase relations, with two parallel experts for $[-60^\\circ,-45^\\circ)\\cup(45^\\circ,60^\\circ]$ and $[-45^\\circ,45^\\circ]$, and a MUSIC pseudospectrum peak search in angle alone to choose the expert—accurately calibrates irregular angle-dependent array errors and improves joint AoA/ToA positioning. The paper reports that the transformer with parallel training achieves lower RMSE than the CNN and MUSIC baselines in simulations across SNR from $-10$ dB to $30$ dB, and in the field test 90% of cases have angle error within $5^\\circ$, distance error within $4$ m, and positioning error within $3$ m except for the 1D-CNN. The self-attention heatmap shows darker striped patterns in certain columns, indicating that the model emphasizes specific blocks' contributions to angle and distance estimation, a global correlation capability the authors say CNNs lack.","pith_inferences":["The two-expert, model-supervised design suggests a general recipe for other ill-calibrated sensing systems: use a classical model-based method for the well-behaved dimension (here, subcarrier/frequency) and train a learned expert per region for the ill-behaved dimension (here, angle).","A stronger test of the calibration claim would remove the simulator from the loop entirely: train the framework on the anechoic chamber phase-error measurements alone and test on the same commercial gNB in the field, which would isolate the framework's ability to generalize from real measured impairments.","Because the paper assumes a single-path signal model, the framework's behavior under strong multipath or multiple simultaneous users is untested; a natural extension would be to train the experts on clustered delay-angle channels and check whether the MUSIC region decision remains reliable."],"forward_implications":["In the presence of irregular angle-dependent array errors, joint AoA/ToA estimation with the proposed transformer framework yields lower RMSE than CNN and MUSIC baselines across SNR.","Splitting the angular space into large- and small-angle regions and training parallel experts improves AoA accuracy, especially within $\\pm45^\\circ$, where the 1D-CNN even performs worse than MUSIC.","The MUSIC-based region decision is accurate enough to select the right expert while keeping computational complexity low, because the pseudospectrum search is performed only over the angle dimension.","The framework can calibrate array errors directly from CSI, suggesting that monostatic 5G positioning could be improved without dedicated calibration hardware, using the existing uplink sounding reference signal."],"supporting_citations":[{"why":"Supplies the link-level simulator that generates the 20,000 synthetic training samples, with indoor-scene settings aligned to the field test.","marker":"[9]"},{"why":"Provides the CNN-based spatial-spectrum baseline that the proposed transformer is compared against.","marker":"[5]"},{"why":"Provides the 1D-CNN time-frequency baseline whose poorer angle estimation in small-angle regions is used to motivate the parallel-training approach.","marker":"[6]"},{"why":"Defines the 5G NR uplink positioning framework, including the UL-SRS that the signal model in Section 2.1 is based on.","marker":"[7]"},{"why":"Provides error bounds for uplink and downlink 3D localization in 5G systems, framing the accuracy targets and the monostatic configuration.","marker":"[8]"}],"fun_headline_variants":["Transformer-MUSIC fusion sharpens 5G indoor positioning","Hybrid transformer nails 5G positioning despite array errors","Data-model fusion trumps CNN in 5G angle and time estimation","Two-expert transformer outdoes CNN and MUSIC on 5G localization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the 20,000 synthetic training samples from the link-level simulator faithfully reproduce the angle-dependent array phase errors measured in the anechoic chamber and that those errors match the commercial pico-cell gNB used in the field test.","fun_headline_variants_meta":{"raw":{"variants":["Transformer-MUSIC fusion sharpens 5G indoor positioning","Hybrid transformer nails 5G positioning despite array errors","Data-model fusion trumps CNN in 5G angle and time estimation","Two-expert transformer outdoes CNN and MUSIC on 5G localization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00078,"raw_usage":{"total_tokens":3452,"prompt_tokens":955,"completion_tokens":2497,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":2422}},"tokens_in":571,"tokens_out":2497,"duration_ms":16683,"temperature":1.0,"reasoning_tokens":2422,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:13:48.567989+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the phase-error pattern of a second commercial gNB of the same model in the anechoic chamber, train the framework on the first gNB's error statistics (or on the simulator's error statistics), and test on the second gNB's field data; if the RMSE gain over the MUSIC baseline mostly disappears, the claimed calibration is specific to the simulated or chamber-measured error statistics rather than a general calibration capability.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the CNN-based spatial-spectrum baseline that the proposed transformer is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the 5G NR uplink positioning framework, including the UL-SRS that the signal model in Section 2.1 is based on."}],"review_version":1}