{"id":"61ab2924-e179-4d52-a452-9c96ef342604","arxiv_id":"2507.16132","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A meta-RL two-stage optimizer for movable-antenna-aided cell-free DFRC systems is claimed to outperform DRL and fixed antenna baselines under carrier frequency offset.","lead":"This paper proposes a meta-reinforcement learning framework that jointly optimizes movable antenna positions and beamforming in a cell-free radar-communication network affected by carrier frequency offset. The authors report simulations showing faster convergence and higher communication and sensing rates than deep reinforcement learning and fixed antenna baselines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's two-stage decomposition does not solve the max-min problem (19): the worst-case CFO is fixed before resources are re-optimized, with no outer loop, so the claimed worst-case robustness is not established.","rationale":"The paper clearly states the intended problem: maximize worst-case weighted communication and sensing rate under CFO (equation 19). For that claim to hold, the algorithm must produce a resource allocation whose minimum over CFO is at least as large as any other feasible allocation's minimum. The proposed two-stage scheme instead computes a minimizing CFO for an arbitrary starting allocation, fixes that CFO while optimizing resources, and never revisits the CFO minimization. Since the CFO dependence is coupled to the resource variables through the quadratic forms in (22)-(26), there is no reason the fixed CFO remains adversarial after Stage 2. This is exactly the reader's weakest assumption, and I agree with that identification. I also note the Appendix A convergence proof is scoped to the MO subproblem and provides no support for the joint max-min claim. The simulation section contains no code, hyperparameters, seeds, or error bars, so the numerical advantage over DRL and FPA baselines cannot be independently reproduced; however, the decisive issue is the mismatch between formulation and algorithm, not the empirical curves. Given that the central advertised contribution is robustness to the worst-case CFO, and that this is not what the algorithm computes, the reader's REJECT verdict is appropriate. I would not change the verdict. I did not find an equally strong separate objection; the channel-model inconsistencies and missing reproducibility details are real but secondary, and my concrete test targets the load-bearing assumption directly.","tokens_in":20628,"tokens_out":4414,"duration_ms":49725,"concrete_test":"Run a small exhaustive instance to test the Stage-1/Stage-2 coupling: take A=2, U=1, N=M=1, S=1, discretize CFO to a few values (e.g., -1, 0, +1 kHz), beamformer magnitudes to a coarse grid satisfying (19b), and t/r positions to a coarse grid satisfying (19c). Execute Algorithm 1 to obtain a CFO and Algorithm 2 to obtain w*, t*, r*. Then evaluate the WCSR at that output for every CFO grid point to get V_algo = min_{Δf} WCSR(w*, t*, r*, Δf). Next enumerate all resource grid points to compute V_opt = max_{w,t,r} min_{Δf} WCSR(w,t,r,Δf). If V_algo < V_opt, the two-stage algorithm fails to solve (19). A transparent supplementary check is to record whether the minimizing CFO after Stage 2 differs from the Stage-1 CFO; any such difference demonstrates the load-bearing assumption is false.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the proposed MO/PDD + MRL method maximizes the worst-case WCSR in problem (19), i.e., max over z, {pu}, {uu}, {wa}, {ta}, {ra} of min over uA and Δf of βγr + (1-β)Σγc_u. The algorithm does not implement this coupled max-min. Stage 1 solves subproblem (20) for the CFO that minimizes the current WCSR at fixed resources; Stage 2 solves subproblem (21) for resources at that fixed CFO. The paper then sequences Algorithm 1 and Algorithm 2 without any outer iteration returning to (20). This is not a harmless decoupling: the matrices C, c1, c2,u and the objective in (22)-(26) depend on the resource variables, so the CFO that was worst-case in Stage 1 need not be worst-case after Stage 2 changes beamformers and MA positions. The reported value is therefore not the worst-case WCSR of the final allocation, and the robustness claim in the abstract is unsupported. The convergence proof in Appendix A addresses only the inner MO loop for a fixed resource allocation, not the joint problem (19); it cannot close this gap. Baseline comparisons may still show heuristic gains, but they do not validate the stated robust optimization claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a full-duplex cell-free dual-functional radar-communication (CF-DFRC) system with movable antennas (MAs) under carrier frequency offset (CFO). The authors formulate a max-min problem (19) that maximizes the worst-case weighted communication and sensing rate over the CFO vector, jointly optimizing transmit/receive beamforming, MA positions, and uplink power. They propose a two-stage algorithm: Stage 1 uses manifold optimization (MO) with a penalty term to find a worst-case CFO for a fixed resource allocation (subproblem (20)), and Stage 2 uses meta-reinforcement learning (MRL) to optimize the resource variables for that fixed CFO (subproblem (21)). Simulations compare the proposed approach with DRL-based and fixed-position-antenna baselines, claiming faster convergence and higher worst-case WCSR under CFO impairments.","tokens_in":20892,"tokens_out":7327,"duration_ms":72585,"significance":"If the technical claims were correct, the paper would address a relevant and timely problem: integrating MA position optimization with CFO robustness in a distributed full-duplex ISAC architecture. The system model is ambitious, and the simulation study covers convergence, power scaling, CFO range, and user distance, which is useful for conveying the intended operating regime. However, the central robust-optimization claim is not established. The algorithm does not solve the max-min problem as formulated, the transformation leading to the CFO subproblem is mathematically incorrect, the PDD mechanism is not actually implemented, and the convergence proof in Appendix A is invalid. No reproducible code, parameter-free derivations, or falsifiable predictions are provided. These load-bearing issues prevent the paper from supporting its abstract-level claims, despite the plausibility of the general direction.","major_comments":[{"comment":"The proposed two-stage decomposition does not solve the max-min problem (19). Subproblem (20) is solved for a fixed resource allocation, yielding a CFO vector that minimizes the current WCSR; subproblem (21) then optimizes the resource variables for that fixed CFO. Because the matrices C-tilde and C and the scalars c1 and c2,u in (22)-(24) depend on w_a, t_a, r_a, z, and the receive filters, the worst-case CFO for the Stage-1 allocation need not be worst-case for the Stage-2 allocation. Since Algorithm 1 and Algorithm 2 are executed without an outer iteration returning to (20), the final objective value is not the worst-case WCSR of the final resource allocation. The robustness claim in the abstract is therefore not supported by the algorithm or the simulations.","section":"Section IV, Eqs. (20)-(21), Algorithms 1 and 2"},{"comment":"The 'equivalent' transformation from the minimization problem (27) to the maximization problem (28) is mathematically incorrect. Minimizing beta * c1 / (phi^H C-tilde phi + c-bar1) + (1-beta) * sum_u c2,u / (phi^H C-tilde phi + c-bar2,u) is not equivalent to maximizing beta * (phi^H C-tilde phi + c-bar1) / c1 + (1-beta) * sum_u (phi^H C-tilde phi + c-bar2,u) / c2,u; the reciprocal operation changes the optimizer. This invalidates the subsequent derivation of the worst-case CFO subproblem and means that even Stage 1 is not solved as stated.","section":"Section IV.A, Eq. (28)"},{"comment":"The paper refers to a penalty dual decomposition (PDD) approach, but the mechanism is not implemented. The penalty term lambda * ||u_A - phi||^2 in (29) uses a fixed positive constant lambda, and there is no update rule for lambda or for a dual variable associated with the equality constraint (27c). Consequently, the constraint u_A = phi is not enforced at convergence, and Algorithm 1 does not provide a solution to problem (27)/(20). This is a load-bearing gap because Stage 1's output is used as the fixed CFO for Stage 2.","section":"Section IV.A, Eqs. (27)-(31)"},{"comment":"The convergence proof is flawed. The bound in (60) is a constant independent of ||phi(t+1) - phi(t)||, so it does not establish the Lipschitz continuity of L(phi) that the argument requires. Moreover, the proof asserts Lipschitz continuity of the composite gradient process from the separate Lipschitz properties of the projection, the retraction, and L(phi); none of the displayed inequalities analyzes the gradient at the retracted point. The final summations in (70)-(71) also do not follow from the preceding inequalities. The convergence of Algorithm 1 is therefore not proven.","section":"Appendix A, Eq. (60)"},{"comment":"The uplink channel definition is inconsistent. Equation (2) writes h_u(r_a) = bar-h_u^H F_up(t_a) with F_up(t_a) in C^{L_u,a x N}, using the transmit MA positions t_a and transmit dimension N for a receive-side uplink channel. The received signal model in (10) and the SINR expressions in (15)-(16) depend on this quantity, so the dimensions of the channel vector do not match the receive array. This undermines the system model on which problem (19) is built.","section":"Section III.B, Eq. (2)"}],"minor_comments":[{"comment":"The abstract contains an incomplete sentence ('we adopt to jointly optimize') and the final sentence is truncated ('the MA-aided CF-DFRC system exhibits'), which obscures the main claims.","section":"Abstract"},{"comment":"There are several grammatical errors: 'A significant challenge in wideband CF-DFRC is systems' and 'this paper proposes a robust optimization framefore' should be corrected.","section":"Section I"},{"comment":"The MDP and algorithm definitions contain undefined or mislabeled quantities: the minus sign in 'ra <- -ra union B0 union B1' is unexplained, and 'Computing hat-B_pi' is not defined.","section":"Section IV.B, Algorithm 2"},{"comment":"The training loss plotted in Figs. 5 and 6 is not defined in the text; the reward in Figs. 3 and 4 is also not formally defined for the simulations (e.g., whether it is averaged over seeds or episodes).","section":"Figures 3-6"},{"comment":"The caption of Fig. 7 reads 'WCSR versus episodes with Transmit power', but the horizontal axis is transmit power; the caption should be corrected.","section":"Fig. 7"},{"comment":"The notation in (10)-(14) is confusing: the definition of X_a and the stacking into y are not fully explained, and the same symbol D is used for different quantities in different equations.","section":"Equations (10)-(14)"}],"recommendation":"reject","confidential_remarks":"The manuscript is not ready for publication. The core robust-optimization claim is not validated by the algorithm, the algebraic transformation in Eq. (28) is a fundamental error, and the PDD and convergence arguments are incomplete. There is also a notable self-citation pattern: ref. [27] is authored by overlapping researchers and is used to motivate the CFO model, and the novelty of the CFO treatment relative to [27] is not clarified. The numerical evaluation would benefit from standard deviations across random seeds and a precise definition of the training loss; however, given the technical gaps, these additions alone would not suffice for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's what you should know: this paper does not actually solve the problem it states. The central max-min problem (19) is addressed by a two-stage scheme: first find a worst-case CFO at fixed resource allocation, then optimize resources at that CFO, and stop. There is no outer loop. The CFO that is worst for the initial allocation need not be worst after beamforming and MA positions move, so the reported WCSR is not a worst-case value. This is the load-bearing claim of the abstract, and it's unsupported. The convergence proof in Appendix A doesn't rescue it: it only treats the inner MO iterate for fixed resources, and the key bound (60) is a constant bound, not a Lipschitz inequality.\n\nThat's the bad news. The good news: the specific problem—joint MA positioning, beamforming, and CFO robustness in a full-duplex cell-free DFRC system—is new in this combination. The system model is careful: field-response channels for SI, IAI, and sensing, full-duplex operation, and an OFDM CFO formulation. The simulation section compares three baselines and shows the proposed MRL method converging faster and to higher reward; those plots are plausible, though they lack error bars and hyperparameters. The MA-vs-FPA comparison is a sensible sanity check.\n\nThe other soft spots are less severe but real. The channel equations have several inconsistencies: (2) writes hu(ra) on the left but the right side depends on ta; the SI and IAI model in (3) and (4) has jumbled subscripts; 'ABS' is never defined; the introduction promises a CRLB analysis that never appears. The RL section is hard to follow—the meta-reward equations (47)-(58) mix notation (π, B0, F0, etc.) and the reader is left to guess the architecture. No code or seeds are provided, so the numerical claims are not independently checkable.\n\nWho is this for? A reviewer working on MA-aided ISAC or robust DFRC. It's not a paper to build on directly, but the system model and the MA-for-CFO motivation are worth borrowing. The current version should be rejected, but not desk-rejected: a serious referee can point out the missing outer iteration and push the authors to either prove the sequential heuristic works or abandon the worst-case framing. The ideas are worth a revision.\n\nRecommendation: send to peer review, with the expectation of major revision or rejection unless the authors fix the max-min mismatch.","headline":"The system model is elaborate and the MA-for-CFO idea is fresh, but the algorithm does not actually solve the stated max-min problem, so the worst-case robustness claim is unsupported.","tokens_in":21465,"tokens_out":3281,"would_cite":false,"duration_ms":34075,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Movable antennas plus a two-stage meta-learning optimizer keep worst-case communication-and-sensing rate high under carrier-frequency error in cell-free DFRC networks, beating deep RL and fixed-antenna baselines, this paper claims.","keywords":["meta-reinforcement learning","movable antennas","cell-free dual-functional radar-communication","carrier frequency offset","manifold optimization","penalty dual decomposition","worst-case beamforming","full-duplex integrated sensing and communication"],"falsifier":"Simulate an outer-loop version: after every MA-position and beamforming update, re-solve the worst-case CFO subproblem and iterate both stages until the WCSR stops changing, or evaluate the single-pass final configuration by exhaustive or random sampling over the admissible CFO range $\\Delta f_{\\min} \\le \\Delta f_{a,a'} \\le \\Delta f_{\\max}$. If the sampled worst-case WCSR of the single-pass solution is noticeably below the reported value, or if the outer-loop version achieves a materially higher WCSR, then the claimed worst-case robustness is not established.","tokens_in":20413,"feed_emoji":"📡","tokens_out":13840,"duration_ms":113313,"temperature":0.7,"pith_summary":"Carrier frequency offset (CFO) between distributed access points degrades both the communication capacity and the sensing accuracy of wideband cell-free dual-functional radar-communication (CF-DFRC) systems. This paper argues that movable antennas (MAs), whose positions can be adapted to the channel, give the system enough spatial flexibility to absorb much of the CFO damage, provided the antenna positions and beamforming are re-optimized together. To do that optimization, the paper proposes a two-stage algorithm: a manifold-optimization and penalty-dual-decomposition stage finds the worst-case CFO, and a meta-reinforcement-learning stage tunes MA positions and beamformers in a data-driven way for fast adaptation to changing channels. The paper's central claim is that this combined scheme significantly outperforms conventional deep RL and fixed-position-antenna baselines in weighted communication-and-sensing rate under CFO impairments, and nearly closes the gap to the CFO-free case. A sympathetic reader would care because CFO robustness is a practical blocker for wideband spectrum-sharing 6G networks, and the paper offers a path where the antenna hardware itself is the compensating mechanism.","feed_headline":"Movable antennas plus meta-learning beat deep RL under CFO","feed_subtitle":"A two-stage optimizer keeps radar-plus-communication rate high even when frequency offsets distort wideband signals.","key_machinery":"The load-bearing object is the two-stage decomposition of the max-min problem (19), which maximizes the worst-case weighted communication and sensing rate (WCSR) over the CFO vector. Stage one (subproblem (20)) holds the MA positions and beamformers fixed and searches for the constant-modulus CFO vector $\\mathbf{u}_A$ that minimizes the weighted sum of rates; fractional programming recasts the fractional SINRs via a quadratic transform, the constant-modulus constraint is treated as a complex circle manifold, and a Riemannian conjugate-gradient method with retraction and Armijo line search, closed by a convex CVX step, produces the worst-case CFO. Stage two (subproblem (21)) fixes that CFO and lets a DDPG-based meta-reinforcement-learning agent jointly choose transmit and receive MA positions, beamforming vectors, and powers, using the WCSR as reward and projection operators to enforce power and position constraints. The meta-learning layer is what the paper credits for fast adaptation: an exploration policy is trained to generate rollouts that improve an exploitation policy, with the improvement measured by a meta-reward, so the agent adapts quickly when the channel changes.","core_discovery":"On the paper's own terms, the central discovery is that the worst-case weighted communication and sensing rate (WCSR) of a full-duplex, MA-aided CF-DFRC system under CFO can be effectively maximized by decoupling the max-min problem into a worst-case CFO subproblem and a joint MA-position and beamforming subproblem, then solving the former with manifold optimization plus penalty dual decomposition and the latter with a meta-reinforcement-learning policy that adapts across dynamic environments. The paper derives the CFO-corrupted received-signal model for both the uplink communication stream and the sensing echo, asserts that CFO inflates the Cramér–Rao lower bound on target position estimation, and reports simulations in which the proposed MRL approach converges faster and reaches higher rewards than conventional DRL, with MA-enabled schemes outperforming fixed-position antennas over a range of transmit powers, CFO intervals, and target distances.","pith_inferences":["My inference: the margin over the baselines may be partly an artifact of the single-pass two-stage decomposition — because the worst-case CFO is computed only once with the antennas frozen, a re-optimized antenna configuration could face a different worst-case CFO, so the reported worst-case WCSR may be optimistic; adding an outer iteration between stages is a direct test.","My inference: the meta-learning advantage should grow as the channel becomes less stationary, since adaptation value rises with task diversity; a benchmark sweeping channel coherence time and MA movement speed would isolate that mechanism.","My inference: the complex-circle-manifold trick for the constant-modulus CFO vector transfers to other phase-error-dominated wideband problems, such as RIS phase-shift optimization or asynchronous massive-MIMO ISAC, where the same manifold machinery and outer-loop caveat would apply.","Editorial note on the manuscript: the state-space definition cites [47] and [48] for including MA positions in the agent's state, but the reference list ends at [37]; the state-design claim lacks the cited support as printed, though the missing references do not touch the two-stage algorithm's core mechanism."],"forward_implications":["If the central claim is right, wideband CF-DFRC deployments can tolerate imperfect inter-AP synchronization and still deliver weighted communication-plus-sensing rates close to those of a perfectly synchronized system.","The MRL agent needs fewer episodes than conventional DRL to reach a given WCSR, which is what makes re-optimizing antenna positions in real time feasible as wireless environments change.","Widening the CFO variation range degrades every scheme, but the proposed one degrades more slowly, so the MA-plus-MRL pairing acts as a robustness layer rather than a complete cure.","Shrinking the movable-antenna region shrinks the performance gain, which ties the claimed benefit directly to the spatial degrees of freedom the antennas are allowed to exploit."],"supporting_citations":[{"why":"Supplies the CFO-induced inter-carrier interference analysis that motivates the CFO signal model in (1), the impairment the whole paper is built around.","marker":"[4]"},{"why":"Introduces position-flexible antenna architectures and motivates using them as the hardware remedy for CFO-induced channel degradation.","marker":"[6]"},{"why":"Provides the alternating-optimization template that the Stage-1 penalty-dual-decomposition and manifold-based solution is explicitly modeled on.","marker":"[19]"},{"why":"The movable-antenna ISAC beamforming design whose joint-optimization setting this paper extends to full-duplex, cell-free operation under CFO.","marker":"[20]"},{"why":"Supplies the field-response channel theory connecting MA positions to per-path phase differences, used for every channel model in Section III.","marker":"[30]"},{"why":"Gives the line-of-sight sensing channel model used to compute the radar SINR and the sensing rate term.","marker":"[31]"},{"why":"Provides the Riemannian conjugate-gradient framework, retraction, and Armijo line search that carry the worst-case CFO subproblem.","marker":"[36]"},{"why":"The convex-optimization toolbox (CVX) that solves the convex CFO-offset refinement in subproblem (43).","marker":"[37]"}],"fun_headline_variants":["Meta-RL beats deep RL for movable-antenna DFRC with CFO","Adaptive meta-RL positions movable antennas to fight CFO","Movable antennas + meta-learning: CFO mitigation in DFRC","MRL outshines DRL for MA-aided DFRC under CFO"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the worst-case carrier frequency offset found in stage one, with the antennas and beamformers frozen, is still the worst case after stage two moves the antennas and re-optimizes the beamformers — the two stages run once each, with no outer loop re-checking the CFO, and max-min problems generally give no such guarantee.","fun_headline_variants_meta":{"raw":{"variants":["Meta-RL beats deep RL for movable-antenna DFRC with CFO","Adaptive meta-RL positions movable antennas to fight CFO","Movable antennas + meta-learning: CFO mitigation in DFRC","MRL outshines DRL for MA-aided DFRC under CFO"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00086,"raw_usage":{"total_tokens":3764,"prompt_tokens":1012,"completion_tokens":2752,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":2673}},"tokens_in":628,"tokens_out":2752,"duration_ms":22752,"temperature":1.0,"reasoning_tokens":2673,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:17:14.560927+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an outer-loop version: after every MA-position and beamforming update, re-solve the worst-case CFO subproblem and iterate both stages until the WCSR stops changing, or evaluate the single-pass final configuration by exhaustive or random sampling over the admissible CFO range $\\Delta f_{\\min} \\le \\Delta f_{a,a'} \\le \\Delta f_{\\max}$. If the sampled worst-case WCSR of the single-pass solution is noticeably below the reported value, or if the outer-loop version achieves a materially higher WCSR, then the claimed worst-case robustness is not established.","supporting_citations":[{"cited_title":"Analysis of new and existing methods of reducing intercarrier interference due to carrier frequency offset in OFDM,","cited_arxiv_id":null,"evidence_quote":"Supplies the CFO-induced inter-carrier interference analysis that motivates the CFO signal model in (1), the impairment the whole paper is built around."},{"cited_title":"AI- empowered fluid antenna systems: Opportunities, challenges, and future directions,","cited_arxiv_id":null,"evidence_quote":"Introduces position-flexible antenna architectures and motivates using them as the hardware remedy for CFO-induced channel degradation."},{"cited_title":"User localization and environment mapping with the assistance of ris,","cited_arxiv_id":null,"evidence_quote":"Provides the alternating-optimization template that the Stage-1 penalty-dual-decomposition and manifold-based solution is explicitly modeled on."},{"cited_title":"Movable-antenna array enhanced beam- forming: Achieving full array gain with null steering,","cited_arxiv_id":null,"evidence_quote":"Supplies the field-response channel theory connecting MA positions to per-path phase differences, used for every channel model in Section III."},{"cited_title":"Cram ´er-rao bound optimization for joint radar-communication beamforming,","cited_arxiv_id":null,"evidence_quote":"Gives the line-of-sight sensing channel model used to compute the radar SINR and the sensing rate term."},{"cited_title":"Optimization on manifolds: Methods and applications,","cited_arxiv_id":null,"evidence_quote":"Provides the Riemannian conjugate-gradient framework, retraction, and Armijo line search that carry the worst-case CFO subproblem."}],"review_version":1}