REVIEW 5 major objections 6 minor 14 references
A Multi-Armed Bandit Framework for Online Optimisation in Green Integrated Terrestrial and Non-Terrestrial Networks
T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that a bandit controller choosing user-cell associations, bandwidth splits, and macro-cell shutdowns hour by hour can keep users satisfied while cutting terrestrial energy consumption.
desk verdict A plausible bandit-based controller for green TN-NTN, but the reported gains are undermined by an evaluation that appears to train on the same snapshots it tests on and by an unfair baseline that wastes 30 MHz of spectrum. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the BCOMD algorithm paired with a four-parameter action space. BCOMD—bandit-feedback constrained online mirror descent—maintains a distribution over arm configurations, receives the cost and constraint violation of only the sampled arm, forms unbiased gradient estimates (loss, constraint, and a bias term), and updates the distribution in the dual space via a mirror map before projecting back to the simplex. The action space is defined by the parameter vector θ = [ε, τν, τRSRP, α], where ε is the fraction of bandwidth given to the satellite tier, τν and τRSRP are shutdown thresholds on load and received signal strength, and α weighs cell load in the pricing association rule Pi(j)=RSRP_ij − αν_j. This mechanism lets the controller balance the Lagrangian cost (sum-throughput minus regularised energy) against the long-run constraint on unsatisfied UEs, with the dual variable λt updated from accumulated violations.
What would settle it
A decisive check is to split the 24-hour data chronologically: train the bandit on the first part of the day and test on a held-out day or held-out hours, then compare its realised cost and constraint violations against the same time-of-day baseline. If the reported 19% throughput gain and 5% energy saving disappear or become negative under a strict train/test split, the online-optimisation claim is not supported.
Extended reading notes
Core claim
The paper's central claim is that jointly optimising three control levers—UE association, bandwidth split between terrestrial and satellite tiers, and shutdown of terrestrial macro base stations—amounts to an online decision problem that can be solved with bandit feedback alone. Each lever setting is an arm; the action is selected from a distribution updated by BCOMD, a constrained online mirror descent algorithm that estimates cost and constraint-violation gradients from the observed cost of the single chosen arm. The authors report that the learned policy keeps the proportion of unsatisfied UEs near zero in low traffic, improves throughput over 3GPP-TN by up to 19% and over 3GPP-NTN by up to 10% in low-traffic hours, and reduces terrestrial energy consumption by about 5%, while in high-traffic periods it trades a slight energy increase for better load balancing and user satisfaction.
Load-bearing premise
The load-bearing premise is that a policy trained and selected on 168,000 pre-collected network snapshots and then used to sample one action per hour faithfully represents online optimisation, and that BCOMD's constraint-satisfaction guarantees, taken from an unpublished analysis, carry over to this non-stationary 24-hour deployment.
Editorial extensions
If this is right
- If the framework works as reported, a network operator can reduce terrestrial energy use by letting satellites absorb traffic from lightly loaded macro cells during off-peak hours.
- The same bandit policy can improve user satisfaction in peak hours by shifting users away from overloaded cells through the load-aware pricing association rule.
- Because the controller needs only the realised cost and constraint violation of the action it took, it can be applied without an explicit traffic model or full network state.
- The reported 19% throughput and 5% energy gains are upper-end figures for low-traffic periods; during high-traffic periods the benefit shifts from energy to load balancing and satisfaction.
- The framework's performance depends on the chosen grid of 875 arm configurations; finer or differently chosen grids would change the achievable trade-off.
Reading between the lines
- A likely implicit consequence is that the approach can be combined with existing 5G sleep-mode mechanisms, since MBS shutdown is already one of the control levers; the bandit would then decide which cells sleep and when.
- The fixed grid of 875 arms could be replaced by continuous or hierarchical action search, but the bandit's regret bound would need re-examination because the current guarantees come from an unpublished analysis and assume constraints that may not hold in a non-stationary 24-hour traffic cycle.
- Because traffic has a strong daily periodicity, a simple time-of-day look-up table trained on the same snapshots might match much of the reported gain; the distinctive test is whether the bandit stays competitive when traffic is shifted, e.g., by a special event or outage.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes an online optimization framework for integrated terrestrial and non-terrestrial networks (TN-NTN), formulated as a multi-armed bandit problem and solved with the BCOMD algorithm. The action space consists of 875 configurations of four parameters: bandwidth split epsilon, load threshold tau_nu, RSRP threshold tau_RSRP, and association weight alpha. The cost function trades off sum log-throughput against terrestrial energy consumption, while the constraint penalizes unsatisfied UEs. A 24-hour system-level simulation is described, comparing the proposed framework against 3GPP-TN and 3GPP-NTN baselines, and reporting lower unsatisfied-UE fractions, up to 19% throughput gains, and 5% energy savings in low-traffic periods. The manuscript does not provide code, data, or confidence intervals, and the evaluation protocol is ambiguous regarding whether the algorithm is run online or trained on snapshots that include the evaluation hours.
Significance. If substantiated, the framework is relevant to green TN-NTN management because it couples load-aware UE association, bandwidth split, and MBS shutdown in a single online decision process, going beyond fixed max-RSRP baselines. The problem is well motivated and the system model follows 3GPP channel models. However, the central empirical claims are currently not supported by the reported evaluation: the protocol does not clearly implement an online loop, the main terrestrial baseline is handicapped by using only 10 MHz of the 40 MHz total bandwidth, and the optimality narrative relies on an unpublished algorithm reference. The paper also lacks statistical uncertainty quantification and reproducible implementation details. These issues are load-bearing for the paper's central claim, so the manuscript requires substantial revision before the empirical conclusions can be accepted.
major comments (5)
- [Section V, first paragraph; Algorithm 1] The evaluation protocol described in Section V ('we collected 7·10^3 snapshots ... Then, we used the learned policy to sample an action for each hour of the day') does not implement the online loop of Algorithm 1, in which x_t is updated after each bandit observation and no separate learned policy exists before the horizon ends. If the policy was trained on all 168,000 snapshots, including the hours used for evaluation, the reported unsatisfied-UE curves and the 19% throughput / 5% energy figures are contaminated by test data. If instead the 24 hourly evaluations are the only online steps, then BCOMD receives only T=24 bandit feedback samples for an action space of n=875 arms, which is far too few to support the regret and constraint-violation guarantees claimed via reference [7]. The authors should run a genuine online experiment with time-separated train/test snapshots, or report per-time-step regret and cumulative constraint violation over a horizon long enough to exercise the algorithm.
- [Section V-B, benchmarks; Table I] The 3GPP-TN benchmark allocates only 10 MHz to the terrestrial tier even though no satellite tier is present, while Table I states total bandwidth W=40 MHz. A terrestrial-only deployment would reasonably use the full 40 MHz, so the TN baseline is handicapped in both throughput and the number of unsatisfied UEs. The reported average throughput gains of 19% (low-traffic) and 1% (high-traffic) over 3GPP-TN are therefore not attributable to the proposed framework alone. The comparison should be rerun with a 40 MHz terrestrial-only baseline, or the authors should cite a 3GPP specification that fixes the TN tier at 10 MHz even without a satellite tier.
- [Section IV-B, Algorithm 1, and reference [7]] The optimality narrative for the framework relies on BCOMD's sublinear regret and constraint-violation guarantees, which are stated to come from the unpublished reference [7]. Because no theorem or proof is reproduced, and because the setting here is non-stationary with a large action space, a reader cannot verify that those guarantees apply to this problem. In addition, the BCOMD hyperparameters eta, mu, gamma, and Omega, as well as the exact functional form and values of zeta used in Eq. (7), are not reported. Without these, the simulation results cannot be reproduced. Please include the relevant conditions from [7] in an appendix and report all hyperparameters.
- [Section III, Eq. (8), and Section V-A] The constraint is defined in Eq. (8) as the fraction of UEs with R_i < rho_i, but the text immediately after says 'if a UE perceives a RSRP lower than a set threshold RSRPmin, it is considered unsatisfied.' These two criteria are not equivalent, and the figures in Section V-A are described in terms of unsatisfied UEs without stating which definition is plotted. The discrepancy affects the meaning of the constraint violation and should be clarified.
- [Section V, Figs. 1-3] All quantitative claims are based on a single 24-hour simulation run, with no confidence intervals, random seeds, or statistical tests. Since the reported gains of 19% and 5% are differences between single trajectories, it is impossible to assess whether they are significant or within simulation noise. Multiple independent runs with uncertainty quantification are needed to support the empirical conclusions.
minor comments (6)
- [Eq. (6)] The indicator function is typeset as `/BD{p_j>0}`; it should be a standard indicator and the energy units should be stated explicitly.
- [Eqs. (9)-(11)] The oracle policy in (9) is defined over x_t although f_t(a_t,x_t) depends on the realized action a_t; the expectation over a_t should be made explicit, as in (11).
- [Eqs. (14)-(16)] The symbol f_t is used both for the cost vector and for the scalar cost f_t(a_t,x_t), which makes the estimator in (16) harder to follow; please distinguish vector and scalar.
- [Fig. 1] The caption says 'satisfied UE proportion' while the y-axis label and text refer to 'unsatisfied UE'; please align the terminology.
- [Table I] The demand parameter lambda_U, the hourly UE counts, and the BCOMD hyperparameters (eta, mu, gamma, Omega, zeta) are not listed, so the simulation cannot be reproduced from the paper alone.
- [References] Reference [7] is marked 'Under Submission'; either provide a citable version or reproduce the required regret and constraint-violation guarantees in an appendix.
Circularity Check
The reported online gains are in-sample evaluations: the policy is 'learned' from the same 168k snapshots against which per-hour actions are scored, making the 19%/5% figures fitted rather than predicted.
-
fitted input called prediction
[Section V 'Simulation Results and Analysis', first paragraph; Algorithm 1]
"Using a custom-built system-level simulator, we collected 7 · 10^3 snapshots of the network for each hour of the day, yielding a total of 168 · 10^3 samples. Then, we used the learned policy to sample an action for each hour of the day to evaluate the resulting performance."
Algorithm 1 is the only learning procedure in the paper and it updates x_t online via bandit feedback; no separate training/evaluation split is described. A policy that can 'sample an action for each hour' before those hours are evaluated can only be the result of training on the pre-collected 168k snapshots of those same hours, especially since T=24 hourly decisions and n=875 arms leave no chance for genuine online learning. The reported throughput, energy, and UE-satisfaction metrics are exactly the components of the cost (7) and constraint (8) that the learned policy was chosen to minimise, so the 19% and 5% figures are in-sample fitted values, not predictions on unseen conditions.
full rationale
The optimization formulation itself is not definitionally circular: the cost function in Eq. (7), the constraint in Eq. (8), and the BCOMD updates in Algorithm 1 are stated from first principles, and there is no direct Eq. X = Eq. Y tautology. The circularity lies in the evaluation protocol: Section V collects 168,000 snapshots, then uses 'the learned policy' to select each hour's action and scores the same hours. Because the policy is fitted to those same snapshots, the headline gains are in-sample evaluations of the objective it was trained to minimize, which is a fitted-input-called-prediction pattern. I did not treat the reliance on unpublished reference [7] as a separate circular step: the BCOMD guarantees are invoked for the 'optimality' narrative, but the numerical results are empirical and do not numerically derive from those guarantees; moreover the reference lists no authors, so a self-citation chain cannot be established from the paper text alone. Score 6 reflects that the central performance claims reduce to a fit on the evaluation data, although the framework components themselves are independently defined.
Assumptions & free parameters
free parameters (2)
- zeta (cost regularisation factor) =
inversely proportional to number of UEs K
- BCOMD hyperparameters eta, mu, gamma, Omega =
not reported
assumptions (5)
- ad hoc to paper BCOMD attains sublinear regret and constraint violation as stated in [7].
- domain assumption Orthogonal TN and NTN bandwidth partitions make inter-tier interference negligible (Eq. 3).
- domain assumption Large-scale channel and energy models follow 3GPP TR 38.901, 38.811, 38.821 and Piovesan et al. [10].
- domain assumption UE data-rate demands are independent exponential random variables with parameter lambda_U.
- domain assumption The 875-arm grid [epsilon, tau_nu, tau_RSRP, alpha] contains a near-optimal configuration for each traffic condition.
Cite this review
Pith. "Pith review of A Multi-Armed Bandit Framework for Online Optimisation in Green Integrated Terrestrial and Non-Terrestrial Networks." pith.science (2026). https://pith.science/paper/XUKDZJBZ
@misc{pith2026250609268,
author = {Pith},
title = {Pith review of: A Multi-Armed Bandit Framework for Online Optimisation in Green Integrated Terrestrial and Non-Terrestrial Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/XUKDZJBZ}},
note = {Machine review of arXiv:2506.09268}
}
read the original abstract
Integrated terrestrial and non-terrestrial network (TN-NTN) architectures offer a promising solution for expanding coverage and improving capacity for the network. While non-terrestrial networks (NTNs) are primarily exploited for these specific reasons, their role in alleviating terrestrial network (TN) load and enabling energy-efficient operation has received comparatively less attention. In light of growing concerns associated with the densification of terrestrial deployments, this work aims to explore the potential of NTNs in supporting a more sustainable network. In this paper, we propose a novel online optimisation framework for integrated TN-NTN architectures, built on a multi-armed bandit (MAB) formulation and leveraging the Bandit-feedback Constrained Online Mirror Descent (BCOMD) algorithm. Our approach adaptively optimises key system parameters--including bandwidth allocation, user equipment (UE) association, and macro base station (MBS) shutdown--to balance network capacity and energy efficiency in real time. Extensive system-level simulations over a 24-hour period show that our framework significantly reduces the proportion of unsatisfied UEs during peak hours and achieves up to 19% throughput gains and 5% energy savings in low-traffic periods, outperforming standard network settings following 3GPP recommendations.
Figures
Reference graph
Works this paper leans on
-
[7]
Adversarial multi-armed bandits with constraints in d ynamic environ- ments,
“Adversarial multi-armed bandits with constraints in d ynamic environ- ments,” Under Submission , 2025
work page 2025
-
[1]
D. L ´ opez-P´ erezet al. , “A survey on 5g radio access network energy efficiency: Massive mimo, lean carrier design, sleep modes, and machine learning,” IEEE Communications Surveys & Tutorials , vol. 24, no. 1, pp. 653–697, 2022
work page 2022
-
[2]
Non-terrestrial networks in the 6g era: Challenges and opportunities,
M. Giordani et al. , “Non-terrestrial networks in the 6g era: Challenges and opportunities,” IEEE Network , vol. 35, no. 2, 2021
work page 2021
-
[3]
Uav communications in integrated terrestrial and non-terrestrial networks,
M. Benzaghta et al. , “Uav communications in integrated terrestrial and non-terrestrial networks,” in 2022 IEEE GLOBECOM , December 2022, pp. 1–6
work page 2022
-
[4]
Distributed pricing-based user association for downlin k heterogeneous cellular networks,
K. Shen et al. , “Distributed pricing-based user association for downlin k heterogeneous cellular networks,” IEEE Journal on Selected Areas in Communications, vol. 32, no. 6, pp. 1100–1113, 2014
work page 2014
-
[5]
H. Alam et al. , “Throughput and coverage trade-off in integrated terrestrial and non-terrestrial networks: an optimizatio n framework,” in IEEE ICC W orkshops, 2023
work page 2023
-
[6]
——, “Optimizing integrated terrestrial and non-terres trial networks performance with traffic-aware resource management,” 2025 . [Online]. Available: https://arxiv.org/abs/2410.06700
work page Pith review arXiv 2025
-
[8]
TR 38.901, Study on channel model for frequ encies from 0.5 to 100 GHz ,
3GPP TSG RAN, “TR 38.901, Study on channel model for frequ encies from 0.5 to 100 GHz ,” V17.0.0, March 2022
work page 2022
Show all 14 references
-
[9]
TR 38.811, Study on New Radio (NR) to support non-ter restrial networks,
——, “TR 38.811, Study on New Radio (NR) to support non-ter restrial networks,” V15.4.0, September 2020
2020
-
[10]
Machine learning and analytical power consumption models for 5g base stations,
N. Piovesan et al., “Machine learning and analytical power consumption models for 5g base stations,” IEEE Communications Magazine , vol. 60, no. 10, pp. 56–62, 2022
2022
-
[11]
TR 38.821, Solutions for NR to support non - terrestrial networks (NTN),
3GPP TSG RAN, “TR 38.821, Solutions for NR to support non - terrestrial networks (NTN),” V16.1.0, May 2021
2021
-
[12]
TR 36.763, Study on NB-IoT / eMTC support for NTN,
——, “TR 36.763, Study on NB-IoT / eMTC support for NTN,” V17.0.0, June 2021
2021
-
[13]
TR 36.814, E-UTRA; Further advancements for E-UTR A phys- ical layer aspects,
——, “TR 36.814, E-UTRA; Further advancements for E-UTR A phys- ical layer aspects,” V9.2.0, March 2017
2017
-
[14]
TR 36.931, E-UTRA; Radio Frequency (RF) requireme nts for LTE Pico Node B,
——, “TR 36.931, E-UTRA; Radio Frequency (RF) requireme nts for LTE Pico Node B,” V17.0.0, March 2022
2022
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.