REVIEW 4 major objections 6 minor 60 references
Multi-AAV-enabled Distributed Beamforming in Low-Altitude Wireless Networking for AoI-Sensitive IoT Data Forwarding
T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that AAV swarms using distributed beamforming can cut IoT data age without return flights, and that a modified soft actor-critic algorithm (SAC-TLS) jointly learning trajectories and schedules achieves this in simulation.
desk verdict A competent, incremental DRL-for-UAV-relay paper; the A2A broadcast assumption is the main unverified lever, but the work is deserving of a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The virtual antenna array (VAA) is the central physical mechanism: AAVs jointly transmit to the base station with an SNR given by the squared sum of per-AAV amplitude terms, (sum over j of sqrt(P_j(t) g0 d_{j,BS}^{-alpha}))^2 / sigma^2, which defines the air-to-ground rate in Eq. (8). This coherent beamforming gain is what removes the need for return flights. The second mechanism is the SAC-TLS algorithm, which stacks temporal sequence input, layer-normalized GRU gates, and squeeze-and-excitation channel attention on top of soft actor-critic, and uses dynamic proximity-based action mapping (DPAM) to replace risky binary communication decisions with deterministic proximity-based ones.
What would settle it
A simulation or field test that replaces the lossless air-to-air broadcast with a finite-rate, lossy channel, or that adds random oscillator phase offsets to the beamforming sum in Eq. (8), would settle it: if the average AoI no longer falls by the reported 17.3%, the gain depends on those idealizations.
Extended reading notes
Core claim
The central claim is that age of information in AAV-relayed IoT networks can be minimized by having AAVs collectively form a virtual antenna array for the air-to-ground hop, so that sensors' data reaches a remote base station without requiring the AAVs to physically return to it. The paper formulates a non-convex, mixed-integer optimization that jointly chooses AAV hover trajectories and sensor communication schedules to minimize time-averaged AoI and AAV energy consumption. To solve it, the paper proposes SAC-TLS, which augments soft actor-critic with temporal sequence state input, a layer-normalized gated recurrent unit to capture long-term dependencies, and a squeeze-and-excitation block
Load-bearing premise
The system assumes every AAV can reliably broadcast all collected data to every other AAV within each time slot and that distributed beamforming achieves perfect phase alignment, so the only bottlenecks are the ground-to-air collection and the air-to-ground link.
Editorial extensions
If this is right
- If the system works as simulated, a swarm of single-antenna AAVs can serve as a long-range relay to a distant base station without periodic return flights, eliminating the main source of AoI spikes in drone-relayed IoT.
- Jointly optimizing trajectories and communication schedules in one DRL policy outperforms treating them separately; the ablation results indicate that temporal sequence input, LNGRU, and the SE block each contribute to the improvement.
- The learned policy scales with swarm and network size: more AAVs reduce average AoI at the cost of higher energy, more sensors reduce average AoI, and SAC-TLS keeps the best trade-off in each tested configuration.
- SAC-TLS retains near-original performance after 90% structured pruning (about 20% reward loss under tuned regularization), suggesting the trained policy can run on resource-constrained onboard hardware.
- Faster convergence and higher cumulative reward imply the method operates as an online controller rather than an offline planner, useful in environments with changing data arrivals and channel conditions.
Reading between the lines
- If the lossless air-to-air broadcast assumption is relaxed, the time-slot budget must explicitly allocate air-to-air transmission, and the policy would need to trade collection breadth against intra-swarm delivery; the current DPAM scheduler would likely need an extra air-to-air-aware reward term.
- A hardware consequence of the coherent-beamforming model is that oscillator synchronization and positioning accuracy become first-order constraints; sub-wavelength phase errors would degrade the virtual-antenna gain and convert the reported gains into ordinary multi-hop diversity gains.
- The same SAC-TLS architecture (temporal sequence input, recurrent normalization, channel attention) could apply to other time-critical mobile relay settings, such as ground robots or mixed aerial-terrestrial relays, whenever age of information is the metric.
- A testable extension is to replace the fixed SNR model with measured channel data and compare SAC-TLS's learned trajectories against a lower-bound oracle that knows future data arrivals, isolating how much of the gain comes from prediction versus exploration.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies an IoT data-forwarding system in which a swarm of AAVs forms a virtual antenna array (VAA) to relay sensor data to a remote base station, with the goal of minimizing time-average age of information (AoI) and AAV energy consumption. The problem is formulated as a non-convex mixed-integer program over AAV trajectories and communication schedules, and a deep reinforcement learning algorithm, SAC-TLS, is proposed by augmenting soft actor-critic with temporal sequence input, layer-normalized GRU (LNGRU), and squeeze-and-excitation (SE) blocks. Simulation experiments compare SAC-TLS with greedy, TD3, PPO, TQC, and SAC baselines, and report a 17.3% reduction in average AoI, a 24.5% increase in cumulative reward, 1.6x faster convergence, ablation results, and pruning performance under 90% compression.
Significance. The topic is timely and the proposed algorithmic components (temporal sequence modeling, LNGRU, SE blocks) are plausible enhancements for a dynamic UAV/IoT setting. The paper also provides a complexity analysis and an ablation/pruning study. If the reported gains were supported under realistic communication assumptions, the work would be a useful step toward multi-AAV distributed-beamforming AoI optimization. However, the quantitative claims are currently conditional on a non-standard AoI update, an unverified lossless A2A broadcast assumption, and a simplified MDP that removes the scheduling variable from the learned action space. Thus, the significance is not yet established at the level claimed.
major comments (4)
- [Section 3.3, Eq. (10)] The AoI update A_i(t+1)=(1-Q(t))(A_i(t)+1) when SN i is scheduled is not the conventional AoI definition. Standard AoI after successful delivery resets to the age of the delivered update (typically 1), and partial progress does not continuously scale down the age. Here, Q(t)<1 multiplicatively reduces the age even if only a fraction of the data is forwarded, and Q(t)=1 gives A_i(t+1)=0, which is inconsistent with data generated at the current slot. This non-standard metric is exactly the term minimized in the reward (Eq. (20)), so the reported 17.3% AoI gain in Section 6 may be an artifact of the metric itself. The authors should derive Eq. (10) from a packet/freshness model or cite a precedent; otherwise, they should adopt a standard AoI definition and rerun the evaluation.
- [Section 3.2, Eqs. (7) and (8)] The A2A phase is modeled with a finite broadcast rate and a duration delta_A2A(t)=max_j S_j/R_j^{A2A}, but the text then assumes that all collected data can be reliably broadcast within each time slot. No simulation parameters for the A2A link (B_j', P_j', rho0, alpha) or sensor data sizes D_i are reported, and no verification is given that delta_G2A + delta_A2A + delta_A2G + delta_move <= 1 for the trained policy. In parallel, Eq. (8) assumes perfect coherent beamforming with no synchronization overhead or phase error. Both idealizations sit directly upstream of Q(t) in Eq. (9) and the AoI update Eq. (10). If these phases are not negligible or phase alignment is imperfect, the A2G capacity shrinks and the claimed AoI improvement may vanish. The authors should report the A2A link budget, validate the slot-time feasibility, and include a sensitivity analysis with respect to synchronizatio
- [Section 5.1, DPAM vs. Section 4.1] Problem P1 in Section 4 optimizes the binary communication schedule Phi, but the MDP simplification with DPAM removes the beta_{i,j}(t) decision from the action space and instead sets it deterministically by whether SNs are within the AAV communication radius. This is a heuristic, not a learned scheduling policy. The paper therefore does not solve P1 as stated: only trajectories are optimized, and the schedule is an input rule. The authors should either reformulate P1 to include the DPAM mapping as a constraint, or characterize the suboptimality of the proximity-based schedule. This is load-bearing for the 'joint trajectory and communication scheduling' contribution claimed in the Introduction and Section 4.
- [Section 6.3, Figs. 5 and 6] The claimed trend that the time-average AoI decreases as the number of SNs increases is counterintuitive and is not explained by the model. With fixed AAV resources and total data volume, adding SNs should generally increase per-SN waiting time. The paper's explanation that 'more SNs allow the system to gather and transmit more frequent updates' does not follow from Eq. (10). Additionally, Figs. 4-8 show single-run curves without error bars or seed statistics, so the convergence and comparison claims are not statistically grounded. Please provide multiple-seed means/confidence intervals and re-examine the mechanism behind Fig. 6.
minor comments (6)
- [Throughout] Repeated grammar issues: 'a AAV' should be 'an AAV'; 'characteristize' in Section 3.2 should be 'characterize'; Table 2 has 'Y min, X max' in the Y-border row.
- [Section 5.3] The sentence 'the actor network combines LNGRU and SE block to generate action at based on the state sequence' appears twice in the same paragraph, once before the detailed description; please remove the duplicate.
- [Section 5.2, Eq. (23)] Eq. (23) defines the temperature alpha as an average log-probability, which is not the standard SAC temperature update. In SAC, alpha is learned by minimizing a different objective. Please clarify or remove this equation, as it is not used in the algorithm.
- [Section 5.3, Eq. (28)] The SE block uses H and W for the feature map, but the LNGRU hidden state h_t is a vector, not a 2D feature map. Please define the dimensions of h_t and how global average pooling is applied in this context.
- [Section 5.4] The quantities K and N_b are used in the training-phase complexity expression before being defined. Please define them explicitly, including the condition for when the replay buffer triggers updates.
- [Section 6.1, Table 3] Several parameters needed to reproduce the simulation are missing: transmit powers P_i and P_j', channel bandwidths B and B_j', reference channel gain rho0, path-loss exponent alpha, and sensor data sizes D_i. A2A parameters are especially important given the concerns in Major Comment 2.
Circularity Check
No significant circularity: the claimed AoI/energy gains are empirical simulation comparisons in which the reward function is the standard RL encoding of the stated objectives; no prediction reduces by construction to a fitted input or to a self-citation.
full rationale
The paper's central claim is that SAC-TLS outperforms baseline DRL and greedy policies in its own simulator. The reward function in Eq. (20) is explicitly composed of the same AoI and energy terms that are later evaluated; this is the normal construction of an RL objective rather than a fitted input being relabeled as a prediction, and all compared algorithms optimize the same MDP/reward. The AoI dynamics in Eq. (10) follow directly from the model assumptions, and no parameter is fitted to the evaluation data and then reported as a prediction. The only self-citation is the prior conference version [1], which is not load-bearing for any derivation. The lossless A2A broadcast assumption (Sec. 3.2) and the coherent-beamforming SNR model (Eq. 8) are idealizations that may affect realism and correctness, but they are inputs to the model, not conclusions that reduce to their own premises. External benchmark verification is not present, but that is an evidentiary limitation, not circularity.
Assumptions & free parameters
free parameters (4)
- Reward normalization weights rho1, rho2, rho3 =
not reported
- Pruning regularization strength lambda =
lg lambda approximately -6.0 (around 1e-6)
- Temporal sequence length n =
not specified
- DPAM communication radius / scheduling threshold =
not specified
assumptions (5)
- domain assumption All G2A and A2A data can be broadcast reliably within each time slot (Section 3.2, before Eq. 7).
- domain assumption Perfect coherent distributed beamforming: SNR at the BS is (sum_j sqrt(P_j g0 d_{j,BS}^{-alpha}))^2 / sigma^2 (Eq. 8), ignoring synchronization and phase errors.
- ad hoc to paper AoI update in Eq. (10): A_i(t+1) = (1-Q(t))(A_i(t)+1) when scheduled, where Q is the fraction of total data forwarded.
- domain assumption Probabilistic LoS model Eq. (2) with parameters m,n, reported in Table 3 as a=5.18 and b=0.43.
- domain assumption Rotary-wing propulsion energy model Eq. (11) with parameters P0, P1, Utip, v0, d0, s, rho from prior literature.
Cite this review
Pith. "Pith review of Multi-AAV-enabled Distributed Beamforming in Low-Altitude Wireless Networking for AoI-Sensitive IoT Data Forwarding." pith.science (2026). https://pith.science/paper/AAJYXK4S
@misc{pith2026250901427,
author = {Pith},
title = {Pith review of: Multi-AAV-enabled Distributed Beamforming in Low-Altitude Wireless Networking for AoI-Sensitive IoT Data Forwarding},
year = {2026},
howpublished = {\url{https://pith.science/paper/AAJYXK4S}},
note = {Machine review of arXiv:2509.01427}
}
read the original abstract
With the rapid development of low-altitude wireless networking, autonomous aerial vehicles (AAVs) have emerged as critical enablers for timely and reliable data delivery, particularly in remote or underserved areas. In this context, the age of information (AoI) has emerged as a critical performance metric for evaluating the freshness and timeliness of transmitted information in Internet of things (IoT) networks. However, conventional AAV-assisted data transmission is fundamentally limited by finite communication coverage ranges, which requires periodic return flights for data relay operations. This propulsion-repositioning cycle inevitably introduces latency spikes that raise the AoI while degrading service reliability. To address these challenges, this paper proposes a AAV-assisted forwarding system based on distributed beamforming to enhance the AoI in IoT. Specifically, AAVs collaborate via distributed beamforming to collect and relay data between the sensor nodes and remote base station. Then, we formulate an optimization problem to minimize the AoI and AAV energy consumption, by jointly optimizing the AAV trajectories and communication schedules. Due to the non-convex nature of the problem and its pronounced temporal variability, we introduce a deep reinforcement learning solution that incorporates temporal sequence input, layer normalization gated recurrent unit, and a squeeze-and-excitation block to capture long-term dependencies, thereby improving decision-making stability and accuracy, and reducing computational complexity. Simulation results demonstrate that the proposed SAC-TLS algorithm outperforms baseline algorithms in terms of convergence, time average AoI, and energy consumption of AAVs.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
AoI-Sensitive Data Forwarding with Distributed Beamforming in UAV-Assisted IoT
Z. Lang, G. Liu, G. Sun, J. Li, Z. Sun, J. Wang, and V . C. M. Leung, “AoI-sensitive data forwarding with distributed beamforming in UAV-assisted IoT,” CoRR, vol. abs/2502.09038, 2025
work page Pith review arXiv 2025
-
[2]
Low-altitude intelligent transportation: System architecture, infrastructure, and key technologies,
C. Huang, S. Fang, H. Wu, Y. Wang, and Y. Yang, “Low-altitude intelligent transportation: System architecture, infrastructure, and key technologies,” J. Ind. Inf. Integr., vol. 42, p. 100694, 2024
work page 2024
-
[3]
Networked ISAC based UAV tracking and handover towards low-Altitude economy,
Y. Feng, C. Zhao, H. Luo, F. Gao, F. Liu, and S. Jin, “Networked ISAC based UAV tracking and handover towards low-Altitude economy,” IEEE Transactions on Wireless Communications, 2025
work page 2025
-
[4]
An online joint optimization approach for qoe maximization in uav- enabled mobile edge computing,
L. He, G. Sun, Z. Sun, P . Wang, J. Li, S. Liang, and D. Niyato, “An online joint optimization approach for qoe maximization in uav- enabled mobile edge computing,” in Proc. IEEE INFOCOM 2024 , 2024, pp. 101–110
work page 2024
-
[5]
Space-Air-Ground integrated wireless networks for 6g: Basics, key technologies, and future trends,
Y. Xiao, Z. Ye, M. Wu, H. Li, M. Xiao, M. Alouini, A. Al-Hourani, and S. Cioni, “Space-Air-Ground integrated wireless networks for 6g: Basics, key technologies, and future trends,” IEEE J. Sel. Areas Commun., vol. 42, no. 12, pp. 3327–3354, 2024
work page 2024
-
[6]
A two time- Scale joint optimization approach for UAV-assisted MEC,
Z. Sun, G. Sun, L. He, F. Mei, S. Liang, and Y. Liu, “A two time- Scale joint optimization approach for UAV-assisted MEC,” inIEEE INFOCOM. IEEE, 2024, pp. 91–100
work page 2024
-
[7]
W. Xie, G. Sun, B. Liu, J. Li, J. Wang, H. Du, D. Niyato, and D. I. Kim, “Joint optimization of uav-carried IRS for urban low altitude mmWave communications with deep reinforcement learning,” CoRR, vol. abs/2501.02787, 2025
work page Pith review arXiv 2025
-
[8]
Multi-objective optimization for multi-uav-assisted mo- bile edge computing,
G. Sun, Y. Wang, Z. Sun, Q. Wu, J. Kang, D. Niyato, and V . C. M. Leung, “Multi-objective optimization for multi-uav-assisted mo- bile edge computing,” IEEE Trans. Mob. Comput. , vol. 23, no. 12, pp. 14 803–14 820, 2024
work page 2024
Show all 60 references
-
[9]
UAV-Enabled secure communications via collaborative beamforming with imperfect eavesdropper information,
G. Sun, X. Zheng, Z. Sun, Q. Wu, J. Li, Y. Liu, and V . C. M. Leung, “UAV-Enabled secure communications via collaborative beamforming with imperfect eavesdropper information,” IEEE Trans. Mob. Comput., vol. 23, no. 4, pp. 3291–3308, 2024
2024
-
[10]
BARGAIN-MATCH: A game theoretical approach for resource allocation and task offloading in vehicular edge computing networks,
Z. Sun, G. Sun, Y. Liu, J. Wang, and D. Cao, “BARGAIN-MATCH: A game theoretical approach for resource allocation and task offloading in vehicular edge computing networks,” IEEE Trans. Mob. Comput., vol. 23, no. 2, pp. 1655–1673, 2024
2024
-
[11]
Rule-guided DRL for UAV-Assisted wireless sensor networks with no-Fly zones safety,
Z. Bai, J. Shi, Z. Li, M. Li, and K. Chen, “Rule-guided DRL for UAV-Assisted wireless sensor networks with no-Fly zones safety,” IEEE Trans. Cogn. Commun. Netw. , vol. 11, no. 2, pp. 1268–1280, 2025
2025
-
[12]
Joint task offloading and resource allocation in aerial-terrestrial UAV networks with edge and fog computing for post-disaster rescue,
G. Sun, L. He, Z. Sun, Q. Wu, S. Liang, J. Li, D. Niyato, and V . C. M. Leung, “Joint task offloading and resource allocation in aerial-terrestrial UAV networks with edge and fog computing for post-disaster rescue,” IEEE Trans. Mob. Comput., vol. 23, no. 9, pp. 8582–8600, 2024
2024
-
[13]
Multi-objective optimization for UAV swarm-assisted IoT with virtual antenna arrays,
J. Li, G. Sun, L. Duan, and Q. Wu, “Multi-objective optimization for UAV swarm-assisted IoT with virtual antenna arrays,” IEEE Trans. Mob. Comput., vol. 23, no. 5, pp. 4890–4907, 2024
2024
-
[14]
Two-way aerial secure communications via distributed collaborative beam- forming under eavesdropper collusion,
J. Li, G. Sun, Q. Wu, S. Liang, P . Wang, and D. Niyato, “Two-way aerial secure communications via distributed collaborative beam- forming under eavesdropper collusion,” in Proc. IEEE INFOCOM
-
[15]
UAV swarm-enabled collaborative secure relay commu- nications with time-domain colluding eavesdropper,
C. Zhang, G. Sun, Q. Wu, J. Li, S. Liang, D. Niyato, and V . C. M. Leung, “UAV swarm-enabled collaborative secure relay commu- nications with time-domain colluding eavesdropper,” IEEE Trans. Mob. Comput., vol. 23, no. 9, pp. 8601–8619, 2024
2024
-
[16]
Multi-Objective optimization approaches for physical layer se- cure communications based on collaborative beamforming in UAV networks,
J. Li, G. Sun, H. Kang, A. Wang, S. Liang, Y. Liu, and Y. Zhang, “Multi-Objective optimization approaches for physical layer se- cure communications based on collaborative beamforming in UAV networks,” IEEE/ACM Trans. Netw., vol. 31, no. 4, pp. 1902–1917, 2023
1902
-
[17]
Uav-enabled collaborative beamforming via multi-agent deep reinforcement learning,
S. Liu, G. Sun, J. Li, S. Liang, Q. Wu, P . Wang, and D. Niyato, “Uav-enabled collaborative beamforming via multi-agent deep reinforcement learning,” IEEE Trans. Mob. Comput., vol. 23, no. 12, pp. 13 015–13 032, 2024
2024
-
[18]
AoI-aware scheduling for air-ground collaborative mobile edge computing,
Z. Qin, Z. Wei, Y. Qu, F. Zhou, H. Wang, D. W. K. Ng, and C.-B. Chae, “AoI-aware scheduling for air-ground collaborative mobile edge computing,” IEEE Transactions on Wireless Communications , vol. 22, no. 5, pp. 2989–3005, 2022
2022
-
[19]
UAV trajectory planning for aoi-minimal data collection in UAV-aided IoT networks by transformer,
B. Zhu, E. Bedeer, H. H. Nguyen, R. Barton, and Z. Gao, “UAV trajectory planning for aoi-minimal data collection in UAV-aided IoT networks by transformer,” IEEE Trans. Wirel. Commun., vol. 22, no. 2, pp. 1343–1358, 2023
2023
-
[20]
A learning-based iterative algorithm for AoI-optimal trajectory planning in UAV-assisted IoT networks,
Z. Huang, H. Chen, B. Gu, S. Gong, Z. Su, and M. Guizani, “A learning-based iterative algorithm for AoI-optimal trajectory planning in UAV-assisted IoT networks,” IEEE Transactions on Wireless Communications, 2025
2025
-
[21]
OH-DRL: An AoI-guaranteed energy-efficient approach for UAV- assisted IoT data collection,
B. Yang, Y. Yu, X. Hao, P . L. Yeoh, J. Zhang, L. Guo, and Y. Li, “OH-DRL: An AoI-guaranteed energy-efficient approach for UAV- assisted IoT data collection,” IEEE Transactions on Wireless Commu- nications, 2025
2025
-
[22]
Reliability-optimal UAV-Assisted mobile edge computing: Joint resource allocation, data transmission scheduling and motion con- trol,
J. Zhou, M. Wang, D. Tian, K. Qu, G. Qu, X. Duan, and X. Shen, “Reliability-optimal UAV-Assisted mobile edge computing: Joint resource allocation, data transmission scheduling and motion con- trol,” IEEE Trans. Mob. Comput., vol. 24, no. 5, pp. 4217–4234, 2025
2025
-
[23]
Age efficient optimization in uav-aided VEC network: A game theory viewpoint,
Z. Han, Y. Yang, W. Wang, L. Zhou, T. N. Nguyen, and C. Su, “Age efficient optimization in uav-aided VEC network: A game theory viewpoint,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 12, pp. 25 287–25 296, 2022
2022
-
[24]
AoI optimization in the UAV-Aided traffic monitoring network under attack: A stackelberg game viewpoint,
Y. Yang, W. Wang, L. Liu, K. Dev, and N. M. F. Qureshi, “AoI optimization in the UAV-Aided traffic monitoring network under attack: A stackelberg game viewpoint,” IEEE Trans. Intell. Transp. Syst., vol. 24, no. 1, pp. 932–941, 2023
2023
-
[25]
Evaluating AoI- centric HARQ protocols for UAV networks,
H. Feng, J. Wang, Z. Fang, J. Chen, and D. Do, “Evaluating AoI- centric HARQ protocols for UAV networks,”IEEE Trans. Commun., vol. 72, no. 1, pp. 288–301, 2024. 15
2024
-
[26]
Interference-Aware online optimization for cellular-connected multiple UAV networks with energy constraints,
C. Zhan, H. Hu, Z. Liu, J. Wang, and R. Fan, “Interference-Aware online optimization for cellular-connected multiple UAV networks with energy constraints,” IEEE Trans. Mob. Comput., vol. 23, no. 12, pp. 13 804–13 820, 2024
2024
-
[27]
Lyapunov-guided deep reinforcement learning for semantic- aware aoi minimization in uav-assisted wireless networks,
Y. Long, S. Gong, S. Sun, G. C. Lee, L. Li, and D. Niyato, “Lyapunov-guided deep reinforcement learning for semantic- aware aoi minimization in uav-assisted wireless networks,” IEEE Transactions on Wireless Communications, 2025
2025
-
[28]
Age of information aware UAV deployment for intelligent transportation systems,
R. Han, Y. Wen, L. Bai, J. Liu, and J. Choi, “Age of information aware UAV deployment for intelligent transportation systems,” IEEE Trans. Intell. Transp. Syst., vol. 23, no. 3, pp. 2705–2715, 2022
2022
-
[29]
AoI-minimal UAV crowdsensing by model-based graph convo- lutional reinforcement learning,
Z. Dai, C. H. Liu, Y. Ye, R. Han, Y. Yuan, G. Wang, and J. Tang, “AoI-minimal UAV crowdsensing by model-based graph convo- lutional reinforcement learning,” in IEEE INFOCOM 2022 - IEEE Conference on Computer Communications, London, United Kingdom, May 2-5, 2022. IEEE, 2022, pp...
2022
-
[30]
Reliable and energy-efficient communications via collaborative beamforming for UAV networks,
X. Zheng, G. Sun, J. Li, S. Liang, Q. Wu, M. Yin, D. Niyato, and V . C. M. Leung, “Reliable and energy-efficient communications via collaborative beamforming for UAV networks,” IEEE Trans. Wirel. Commun., vol. 23, no. 10, pp. 13 235–13 251, 2024
2024
-
[31]
Joint sensing and age of information optimization for energy constrained UAV assisted integrated sensing, calculation and communication,
Z. Liu, X. Liu, W. Yang, and X. Zhang, “Joint sensing and age of information optimization for energy constrained UAV assisted integrated sensing, calculation and communication,” IEEE Trans- actions on Wireless Communications, 2025
2025
-
[32]
AoI- minimal clustering, transmission and trajectory co-design for uav- assisted WPCNs,
X. Liu, H. Liu, K. Zheng, J. Liu, T. Taleb, and N. Shiratori, “AoI- minimal clustering, transmission and trajectory co-design for uav- assisted WPCNs,” IEEE Trans. Veh. Technol., vol. 74, no. 1, pp. 1035– 1051, 2025
2025
-
[33]
Online altitude control and scheduling policy for minimizing aoi in UAV- assisted IoT wireless networks,
M. Samir, C. Assi, S. Sharafeddine, and A. Ghrayeb, “Online altitude control and scheduling policy for minimizing aoi in UAV- assisted IoT wireless networks,” IEEE Trans. Mob. Comput., vol. 21, no. 7, pp. 2493–2505, 2022
2022
-
[34]
Min- imizing the AoI in resource-constrained multi-source relaying systems: Dynamic and learning-based scheduling,
A. Zakeri, M. Moltafet, M. Leinonen, and M. Codreanu, “Min- imizing the AoI in resource-constrained multi-source relaying systems: Dynamic and learning-based scheduling,” IEEE Trans. Wirel. Commun., vol. 23, no. 1, pp. 450–466, 2024
2024
-
[35]
Optimization of informa- tion freshness in multi-RIS cooperative assisted wireless sensor network,
L. Qu, A. Huang, and M. J. Khabbaz, “Optimization of informa- tion freshness in multi-RIS cooperative assisted wireless sensor network,” IEEE Trans. Wirel. Commun., vol. 23, no. 11, pp. 16 332– 16 345, 2024
2024
-
[36]
Up- downlink aoi-driven multi-source data collection in uav-assisted wireless sensor networks,
M. Zhao, Y. Xiao, J. Yao, T. Wang, J. Lee, and T. Q. S. Quek, “Up- downlink aoi-driven multi-source data collection in uav-assisted wireless sensor networks,” IEEE Trans. Wirel. Commun. , vol. 24, no. 2, pp. 1178–1192, 2025
2025
-
[37]
Average AoI minimization with directional charging for wireless-Powered network edge,
Q. Chen, S. Guo, W. Xu, J. Li, T. Shi, H. Gao, and Z. Cai, “Average AoI minimization with directional charging for wireless-Powered network edge,” IEEE Transactions on Mobile Computing, 2025
2025
-
[38]
AoI-sensitive data collection in multi-UAV-assisted wireless sensor networks,
X. Gao, X. Zhu, and L. Zhai, “AoI-sensitive data collection in multi-UAV-assisted wireless sensor networks,” IEEE Trans. Wirel. Commun., vol. 22, no. 8, pp. 5185–5197, 2023
2023
-
[39]
Coalitional formation- based group-buying for UAV-enabled data collection: An auction game approach,
N. Qi, Z. Huang, W. Sun, S. Jin, and X. Su, “Coalitional formation- based group-buying for UAV-enabled data collection: An auction game approach,” IEEE Trans. Mob. Comput. , vol. 22, no. 12, pp. 7420–7437, 2023
2023
-
[40]
Finite block length NOMA MU pairing UAV-enable system: Per- formance analysis and optimization,
T. M. Hoang, B. C. Nguyen, T. T. H. Le, X. N. Tran, and P . T. Hiep, “Finite block length NOMA MU pairing UAV-enable system: Per- formance analysis and optimization,” IEEE Trans. Mob. Comput. , vol. 23, no. 10, pp. 9804–9820, 2024
2024
-
[41]
Tradeoff between age of information and operation time for UAV sensing over multi- cell cellular networks,
C. Zhan, H. Hu, J. Wang, Z. Liu, and S. Mao, “Tradeoff between age of information and operation time for UAV sensing over multi- cell cellular networks,” IEEE Trans. Mob. Comput., vol. 23, no. 4, pp. 2976–2991, 2024
2024
-
[42]
Proactive obsolete packet management based analysis of age of information for LCFS het- erogeneous queueing system,
Y. A. K. Reddy and T. G. Venkatesh, “Proactive obsolete packet management based analysis of age of information for LCFS het- erogeneous queueing system,” IEEE Trans. Mob. Comput. , vol. 24, no. 3, pp. 1513–1529, 2025
2025
-
[43]
Maximizing age-energy efficiency in wireless powered industrial IoE networks: A dual-layer DQN-based approach,
H. Zheng, K. Xiong, M. Sun, H. Wu, Z. Zhong, and X. Shen, “Maximizing age-energy efficiency in wireless powered industrial IoE networks: A dual-layer DQN-based approach,” IEEE Trans. Wirel. Commun., vol. 23, no. 2, pp. 1276–1292, 2024
2024
-
[44]
Intelligent scheduling of uavs and sensors for information age minimization at wireless powered internet of things,
J. Li, X. Wang, J. Wu, and Z. Ning, “Intelligent scheduling of uavs and sensors for information age minimization at wireless powered internet of things,” in Proc. IEEE CSCWD , W. Shen, J. A. Barth `es, J. Luo, T. Qiu, X. Zhou, J. Zhang, H. Zhu, K. Peng, T. Xu, and N. Chen, Eds...
2024
-
[45]
AoI minimiza- tion based on deep reinforcement learning and matching game for IoT information collection in SAGIN,
G. Zhang, X. Wei, X. Tan, Z. Han, and G. Zhang, “AoI minimiza- tion based on deep reinforcement learning and matching game for IoT information collection in SAGIN,” IEEE Transactions on Communications, 2025
2025
-
[46]
LI2: A new learning-based approach to timely monitoring of points-of-interest with UAV,
Z. Huang, W. Wu, K. Wu, H. Yuan, C. Fu, F. Shan, J. Wang, and J. Luo, “LI2: A new learning-based approach to timely monitoring of points-of-interest with UAV,” IEEE Trans. Mob. Comput., vol. 24, no. 1, pp. 45–61, 2025
2025
-
[47]
Distributed age-of-information scheduling with NOMA via deep reinforcement learning,
C. Zhang, Y. Zou, Z. Zhang, D. Yu, J. T. G ´omez, T. Lan, F. Dressler, and X. Cheng, “Distributed age-of-information scheduling with NOMA via deep reinforcement learning,” IEEE Trans. Mob. Com- put., vol. 24, no. 1, pp. 30–44, 2025
2025
-
[48]
Collaborative ground-space communications via evolutionary multi-objective deep reinforcement learning,
J. Li, G. Sun, Q. Wu, D. Niyato, J. Kang, A. Jamalipour, and V . C. M. Leung, “Collaborative ground-space communications via evolutionary multi-objective deep reinforcement learning,” IEEE J. Sel. Areas Commun., vol. 42, no. 12, pp. 3395–3411, 2024
2024
-
[49]
AoI-aware scheduling and trajectory optimization for multi-UAV-assisted wireless networks,
Y. Long, W. Zhang, S. Gong, X. Luo, and D. Niyato, “AoI-aware scheduling and trajectory optimization for multi-UAV-assisted wireless networks,” in Proc. IEEE GLOBECOM , 2022, pp. 2163– 2168
2022
-
[50]
Service continuity based data delivery optimization in satellite-terrestrial networks,
F. Wang, D. Jiang, Z. Wang, and S. Mumtaz, “Service continuity based data delivery optimization in satellite-terrestrial networks,” IEEE Trans. Veh. Technol., vol. 72, no. 10, pp. 13 604–13 617, 2023
2023
-
[51]
Aoi optimization for uav-assisted wireless sensor networks,
A. Sun, C. Sun, J. Du, C. Chen, C. Huang, and J. Sui, “Aoi optimization for uav-assisted wireless sensor networks,” in Proc. IEEE ICC, 2024, pp. 1487–1492
2024
-
[52]
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,
T. Haarnoja, A. Zhou, P . Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” inProc. PMLR ICML, vol. 80, 2018, pp. 1856–1865
2018
-
[53]
Layer normalization,
J. Lei Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” ArXiv e-prints, pp. arXiv–1607, 2016
2016
-
[54]
Squeeze-and-excitation networks,
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. IEEE CVPR, 2018, pp. 7132–7141
2018
-
[55]
Addressing function approximation error in actor-critic methods,
S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” ser. Proceedings of Machine Learning Research, vol. 80. PMLR, 2018, pp. 1582–1591
2018
-
[56]
Proximal policy optimization algorithms,
J. Schulman, F. Wolski, P . Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” CoRR, vol. abs/1707.06347, 2017
2017 arXiv
-
[57]
Con- trolling overestimation bias with truncated mixture of continuous distributional quantile critics,
A. Kuznetsov, P . Shvechikov, A. Grishin, and D. P . Vetrov, “Con- trolling overestimation bias with truncated mixture of continuous distributional quantile critics,” in Proc. PMLR ICML, vol. 119, 2020, pp. 5556–5566
2020
-
[58]
Continuous control with deep rein- forcement learning,
T. P . Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep rein- forcement learning,” in Proc. ICLR, 2016
2016
-
[59]
Compressing deep reinforcement learning networks with a dynamic structured pruning method for autonomous driving,
W. Su, Z. Li, M. Xu, J. Kang, D. Niyato, and S. Xie, “Compressing deep reinforcement learning networks with a dynamic structured pruning method for autonomous driving,” IEEE Trans. Veh. Tech- nol., vol. 73, no. 12, pp. 18 017–18 030, 2024
2024
-
[2024]
IEEE, 2024, pp. 331–340
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.