{"id":"465e478b-80c8-43c1-a4e7-a9245f41f601","arxiv_id":"2411.13970","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Soft Actor-Critic agent jointly plans a UAV's trajectory and steers its directional movable antenna to shorten data collection time in backscatter sensor networks.","lead":"This paper proposes a deep reinforcement learning system that flies a drone with a steerable directional antenna to collect data from tiny backscatter sensors, cutting collection time and energy. A generalist might read it because backscatter IoT devices are ultra-low-power, and smarter drones could make remote sensor data gathering faster and cheaper.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The MA's orientation never enters the channel model; the MA/FPA comparison in the simulations is a fixed 10 dBi vs 5 dBi gain gap, not a test of movability.","rationale":"The reader's weakest assumption correctly identifies the gain mismatch between the MA and the FPA baseline. I agree with that concern, but I believe the problem is more structural: the model never couples the MA orientation to the channel, so the advertised mechanism of aiming the main lobe to improve channel gain is absent from the simulation. The MA-vs-FPA difference in the link budget is a fixed scalar gain offset, and the MA's orientation only incurs reorientation time and energy. An equal-gain rerun would settle whether the claimed advantage is due to movability or simply to a 5 dB gain gap. If the concern lands, the current simulation results do not support the central claim, so I recommend rejecting the current evidence while noting that a corrected comparison could make the paper viable.","tokens_in":9126,"tokens_out":11935,"duration_ms":123733,"concrete_test":"Rerun the Fig. 3 and Fig. 4 simulations with the MA's GTR set to 5 dBi, equal to the FPA baseline, while keeping the MA reorientation time/energy model and all other settings unchanged. If the MA+SAC curves no longer dominate FPA+SAC, the headline improvement is produced by the unspecified 10 dBi versus 5 dBi gain gap, not by antenna movability; if they still dominate, the gain confound is not the sole explanation and the claim would be strengthened.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim is that a movable antenna improves data collection by precisely aiming its main lobe at each backscatter device. Yet in the only place where antenna gain enters the link budget, Eqs. (4)-(5) of Section II-B, the reader gain GTR is a fixed scalar, and no formula, pattern, or parameter makes it depend on the orientation angles (θ,φ). Equations (1)-(2) define the angles that would point the main lobe at a BD, but those angles are used only in the mechanical time/energy model tMA, Eq. (14). The communication rate Rk(t) is therefore identical for a correctly aimed and a misaimed MA. Combined with Section IV's baseline description, where the FPA omni gain is 5 dBi while Table I gives GTR=10 dBi with no separate MA gain entry, the simulation compares a 10 dBi link against a 5 dBi link, not a steerable antenna against a fixed one. If the gains were matched, the MA would only add reorientation overhead, so the claimed advantage is not attributable to movability as advertised.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a deep reinforcement learning (SAC) approach for a UAV equipped with a movable directional antenna (MA) collecting data from backscatter devices (BDs). The system model includes a probability-based LoS/NLoS path loss model, a backscatter data rate expression, a UAV propulsion and MA reorientation energy/time model, and an MDP formulation whose actions are the UAV's movement and the MA's initial orientation. Simulations compare the MA-equipped UAV using SAC against an FPA-equipped UAV and an AC baseline, reporting reductions in data collection time and energy. The central advertised mechanism is that steering the MA's main lobe toward each BD increases channel gain, thereby reducing collection time.","tokens_in":9328,"tokens_out":3082,"duration_ms":32757,"significance":"If the central mechanism were correctly modeled and supported, the paper would address a timely combination of movable antennas, UAV data collection, and backscatter communications, and the SAC-based formulation could be a useful baseline for future work. The paper also contributes a reasonable MDP design with a compact state space based on azimuth angles and distances. However, the manuscript as written does not establish the claimed benefit. The antenna orientation never enters the channel model, so the simulations do not actually test 'movability'; the reported gain may simply reflect a fixed 5 dB gain difference between the MA and FPA. In addition, the path loss model has a dimensional inconsistency that affects all numerical results. The paper is therefore a promising idea whose current evidence is not yet convincing.","major_comments":[{"comment":"The MA's orientation angles (θ,φ) are defined in Eqs. (1)-(2) to point the main lobe at a BD, but the link budget in Eqs. (4)-(5) uses a fixed scalar gain GTR=10 dBi (Table I) and contains no antenna pattern or dependence on (θ,φ). Consequently, the data rate Rk(t) in Eq. (10) is identical for a correctly aimed and a misaimed antenna. The only place the angles matter is the mechanical reorientation time in Eq. (14). Since the FPA baseline is described as a 5 dBi omni-directional antenna while GTR=10 dBi, the simulation compares a 10 dBi link against a 5 dBi link, not a steerable antenna against a fixed one. This confound is load-bearing because the abstract and introduction attribute the improvement to 'precisely aiming its main lobe at each BD.' The authors must introduce a gain model that depends on the pointing error (e.g., a main-lobe pattern G(θ,φ) with a specified peak gain and beamwidth) and then compare the MA against an FPA with the same peak gain, or otherwise explicitly separate the gain advantage from the steering advantage.","section":"Section II-B, Eqs. (3)-(5) and Section II-C, Eq. (10)"},{"comment":"The path loss quantities LLoS,k(t) and LNLoS,k(t) in Eqs. (4)-(5) are expressed in dB (20log10(...) plus losses in dB), and Lk(t) in Eq. (3) is a weighted sum of these dB values. However, Eq. (10) then uses Lk(t)^{-2} as if Lk(t) were a linear (dimensionless) channel power gain. Weighted averaging of dB losses is also physically invalid; the correct procedure is to average the linear losses and then convert to dB, or to compute the effective path loss in linear domain first. This inconsistency affects every numerical result in Section IV and the rate computation, so the quantitative claims about time and energy reductions are not trustworthy without correction.","section":"Eqs. (3)-(5) and (10)"},{"comment":"The simulation evidence lacks statistical support. The BDs are randomly distributed, but no error bars, confidence intervals, or multiple-seed results are reported for the time and energy comparisons in Figs. 3 and 4. The AC-FPA baseline is explicitly excluded from these figures 'due to convergence challenges,' which removes the control condition needed to show that SAC, rather than the antenna gain difference, drives the improvement. The paper should provide results averaged over many random BD deployments (with standard deviations), include the AC-FPA result where possible (or explain the failure and provide a converged FPA baseline trained with SAC), and specify the MA's peak gain and directivity so the reader can separate the antenna gain contribution from the steering contribution.","section":"Section IV, Figs. 3-5"}],"minor_comments":[{"comment":"The paper contains several typographical and notation issues, such as the inconsistent spacing in 'UA V' (which should be 'UAV' or consistently 'UA V'), the unnumbered 'P1' in Eq. (15), and the use of 'θ' in Eq. (7) while the text refers to elevation angle without defining which angle is used in the LoS probability.","section":"Throughout"},{"comment":"The action space includes θinit(t) and φinit(t), but since the MA is always reoriented to point at each qualified BD during data collection, the initial orientation only affects the mechanical adjustment time. With no antenna pattern in the channel model, this part of the action space reduces to a time-penalty optimization; the authors should clarify what physical benefit is being optimized beyond reducing reorientation time.","section":"Section III-B"},{"comment":"The reward components pf(t), pMA(t), and pc(t) are described only qualitatively. Explicit formulas for these penalties and a discussion of how their scaling relates to the hand-tuned coefficients rbs=50 and rf=500 would improve reproducibility and help readers assess whether the learned policy is actually minimizing the true objective in (15).","section":"Section III-B3"},{"comment":"The training details are insufficient for reproduction: no neural network architectures, learning rates, batch sizes, replay buffer sizes, or entropy target are reported, and the converged SAC curves in Fig. 2 are not quantified with run-to-run variance. Please provide a hyperparameter table or a reference to a public implementation.","section":"Section IV, Fig. 2"}],"recommendation":"major_revision","confidential_remarks":"The central claim is currently not supported because the simulation's MA/FPA comparison is confounded by a fixed gain difference and the path loss model is dimensionally inconsistent. However, both issues are fixable within the scope of the manuscript by adding a proper antenna pattern, correcting the path loss conversion, and rerunning the comparison with gain-matched baselines and statistical reporting. The novelty of applying SAC to MA-equipped UAV backscatter collection is reasonable, but the paper would need a substantial revision before it can be accepted. I would not recommend rejection outright because the core idea is coherent and the MDP formulation is a serviceable starting point."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nTwo things you should know before reading: (1) the DRL/SAC part is textbook-solid and the system model is clearly explained; (2) the advertised \"movable antenna\" benefit is not actually modeled. In the channel model, Section II-B, the reader antenna gain appears as a fixed scalar GTR (Table I: 10 dBi). It never depends on the orientation angles (θ,φ) that equations (1)-(2) are supposed to aim. Those angles only show up later in the mechanical reorientation time/power model, Eqs. (13)-(14). So Rk(t) is the same whether the main lobe points at a BD or away. The FPA baseline, by contrast, uses an omni antenna at 5 dBi. The comparison therefore tests a 10 dBi link against a 5 dBi link while attributing the difference to movability. That is a load-bearing confound.\n\nWhat is genuinely there: a clean monostatic backscatter setup with UAV hovering, a legitimate simplification of the state space to azimuth+distance, and a correct SAC implementation with sensible reward shaping. For the niche of UAV-assisted backscatter collection, this is a reasonable design study. The authors are also transparent about the AC-FPA convergence failure, which is honest, though it weakens the comparison.\n\nSoft spots beyond the gain confound: no error bars or statistical analysis across random BD placements, no MA gain value in Table I, no code/data, and the excluded AC-FPA baseline leaves only favorable comparisons. All of these are fixable. The paper would be more defensible reframed as \"a high-gain directional antenna on a UAV,\" or the antenna pattern must be included and gains matched.\n\nMy take: the core idea is worth a referee's time only if the authors fix the antenna modeling and rerun the experiments. As-is, I would not cite the headline result, and peer review should demand a revision before acceptance.","headline":"A competent SAC/UAV data-collection study whose 'movable antenna' advantage rests on an unmodeled, confounded 10 dBi vs 5 dBi gain gap.","tokens_in":9870,"tokens_out":3501,"would_cite":false,"duration_ms":35532,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A movable antenna on a UAV cuts backscatter data-collection time and energy in simulations.","keywords":["movable antenna","backscatter communication","UAV trajectory optimization","deep reinforcement learning","Soft Actor-Critic","data collection time","wireless sensor networks"],"falsifier":"Run the same SAC-based simulation with the movable antenna's peak gain set equal to the fixed antenna's gain (for example, both at 5 dBi, or both at 10 dBi with the fixed antenna also directional and aimed at each sensor): if total data-collection time does not decrease when only steering freedom is added, the paper's central claim is refuted.","tokens_in":8912,"feed_emoji":"📡","tokens_out":9342,"duration_ms":87667,"temperature":0.7,"pith_summary":"The paper tries to show that a data-collecting drone becomes faster and more energy-efficient when its antenna can be steered. Backscatter sensors communicate by reflecting a reader's radio signal, so the link works best when the drone's directional antenna points its main lobe straight at each sensor; the authors model this with elevation and azimuth angles that are computed to aim at each device. They formulate the mission as minimizing the sum of flight time, communication time, and antenna reorientation time, and solve the joint trajectory-and-pointing problem with a Soft Actor-Critic (SAC) reinforcement learning agent. In simulation with 20 sensors in a 200 m area, the steerable-antenna drone with SAC finishes in 49 seconds and flies 365 meters, versus 63.7 seconds and 447.9 meters for the same antenna with an actor-critic baseline, and the fixed-antenna drone must nearly visit each sensor. If true, this means steerable antennas, not just smarter flight paths, are a practical lever for backscatter sensor networks.","feed_headline":"Steerable drone antenna cuts backscatter data-collection time","feed_subtitle":"Simulations show a movable antenna beats fixed antennas on both time and energy for sensor data runs.","key_machinery":"The load-bearing mechanism is the movable antenna (MA), a directional antenna whose main-lobe orientation $(\\theta,\\phi)$ is changed in elevation and azimuth, with $\\theta_k$ and $\\varphi_k$ computed from the drone's position and each sensor's location via (1)-(2). This steering raises the effective channel gain in the LoS/NLoS path-loss model, allowing data collection from longer range, and the reorientation cost appears both in the objective through $t_{MA}$ and in the reward through $p_{MA}$, so the agent explicitly trades hover time against turn time. The second mechanism is Soft Actor-Critic (SAC), an off-policy maximum-entropy reinforcement learning algorithm with twin Q-networks and an adaptive entropy coefficient $\\alpha$, which provides stable learning over the action space $\\{a_f, a_r, \\theta_{init}, \\varphi_{init}\\}$ that jointly controls the drone's movement and the antenna's initial pointing.","core_discovery":"On its own terms, the paper's central discovery is that the antenna's reconfigurable orientation is what makes efficient data collection possible: the movable antenna's elevation angle $\\theta_k$ and azimuth angle $\\varphi_k$ are set by equations (1)-(2) to point directly at each backscatter device, increasing the reader antenna gain in the path-loss model and hence the data rate $R_k(t)=B\\log_2(1+\\xi_k P_c L_k(t)^{-2}/N_0)$. With that gain, the drone can hover at one position and rotate the antenna among all sensors whose received powers pass the sensitivity thresholds, collecting their data sequentially, so the trajectory problem changes from 'fly to each sensor' to 'choose hover points and initial antenna orientations.' The SAC-based policy learns this behavior, and the comparison to a fixed-position omnidirectional antenna (5 dBi) and to an actor-critic baseline is the evidence offered for the discovery. The claim is empirical rather than analytical: no performance guarantee is proven, but the simulated mission shows the mechanism working.","pith_inferences":["A testable implication the authors leave implicit is that the advantage may come from higher directional gain rather than from steerability itself: the fixed antenna is set to 5 dBi omnidirectional while the movable antenna's peak gain is never stated, so a matched-gain comparison is the clean way to isolate the steering benefit.","The same SAC-based positioning-and-pointing policy should transfer to any steerable antenna hardware, such as a gimballed dish or phased array, which would make the contribution a control method rather than a new antenna result.","Extending the state to include channel-state information or other drones' positions would let the approach handle multi-drone missions, where pointing decisions also manage co-channel interference.","A real-world test with backscatter tags would be needed to confirm whether the probabilistic LoS channel model overstates the gain in dense urban environments; without it, the 49-second result is a simulation benchmark rather than a measured one."],"forward_implications":["If the paper is right, a single drone can serve more backscatter sensors per mission because rotating a steerable antenna replaces multiple close fly-bys.","Mission time becomes less sensitive to the number of sensors and the size of the area, since the flight component is the one that shrinks.","Energy consumption falls with flight time, and the modeled antenna reorientation energy $P_{MA}$ is small enough that the net budget still improves.","The SAC training recipe is practical: it converges in roughly 3 million steps, while the actor-critic baseline needs about 10 million and fails to converge for fixed antennas."],"supporting_citations":[{"why":"introduces movable antennas for wireless communication, the technology the paper applies to UAVs","marker":"[9]"},{"why":"applies a movable antenna to UAV placement and configuration, the direct scenario this paper extends to data collection with reorientation","marker":"[10]"},{"why":"provides the hierarchical DRL backscatter data-collection baseline with fixed antennas that the movable-antenna approach outperforms","marker":"[7]"},{"why":"gives the complete backscatter-radio link budget that grounds the received-power and data-rate model","marker":"[14]"},{"why":"supplies the probabilistic LoS/NLoS path-loss model for UAV-ground links","marker":"[15]"},{"why":"defines Soft Actor-Critic, the RL algorithm used for the joint trajectory and antenna-orientation optimization","marker":"[16]"}],"fun_headline_variants":["Movable antenna UAV cuts sensor data collection time","SAC-trained drone steers antenna to slash data run time","UAV with steerable antenna reduces backscatter collection time","Deep RL optimizes UAV antenna orientation for faster sensor reads","Movable antenna on drone beats fixed one for backscatter networks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the movable antenna's directional gain, not just its freedom to reorient, produces the improvement: the baseline fixed antenna is a 5 dBi omnidirectional antenna while the movable antenna's peak gain is left unspecified, so if both antennas had equal peak gain, the reported time and energy savings could disappear.","fun_headline_variants_meta":{"raw":{"variants":["Movable antenna UAV cuts sensor data collection time","SAC-trained drone steers antenna to slash data run time","UAV with steerable antenna reduces backscatter collection time","Deep RL optimizes UAV antenna orientation for faster sensor reads","Movable antenna on drone beats fixed one for backscatter networks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1376,"prompt_tokens":981,"completion_tokens":395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":597,"tokens_out":395,"duration_ms":4490,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:40:24.033057+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same SAC-based simulation with the movable antenna's peak gain set equal to the fixed antenna's gain (for example, both at 5 dBi, or both at 10 dBi with the fixed antenna also directional and aimed at each sensor): if total data-collection time does not decrease when only steering freedom is added, the paper's central claim is refuted.","supporting_citations":[{"cited_title":"Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,","cited_arxiv_id":null,"evidence_quote":"defines Soft Actor-Critic, the RL algorithm used for the joint trajectory and antenna-orientation optimization"},{"cited_title":"Hierarchical deep reinforcement learning for backscattering data collection with multiple UA Vs,","cited_arxiv_id":null,"evidence_quote":"provides the hierarchical DRL backscatter data-collection baseline with fixed antennas that the movable-antenna approach outperforms"},{"cited_title":"Complete link budgets for backscatter- radio and RFID systems,","cited_arxiv_id":null,"evidence_quote":"gives the complete backscatter-radio link budget that grounds the received-power and data-rate model"},{"cited_title":"Optimal lap altitude for maximum coverage,","cited_arxiv_id":null,"evidence_quote":"supplies the probabilistic LoS/NLoS path-loss model for UAV-ground links"}],"review_version":1}