REVIEW 4 major objections 5 minor 65 references
A Unified QoS-Aware Multiplexing Framework for Next Generation Immersive Communication with Legacy Wireless Applications
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read By learning slicing decisions from queue backlogs, two algorithms keep 6G immersive and legacy traffic on one RAN with sublinear regret and measured throughput and latency gains.
desk verdict Useful slicing framework, but the main regret bound rests on an unproved and generally false linearity assumption—worth referee time, not acceptance as is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a three-level control loop. (1) Virtual queues $G_n(k)$ are constructed from the real packet queues and the delay thresholds $\delta_{e,M}$, so that the long-term average-backlog constraint becomes a queue-stability problem. (2) The Lyapunov drift-minus-utility bound converts the original problem into a per-frame utility maximization that is solved by the PBRA algorithm, a penalty successive upper-bound minimization with block coordinate descent. (3) A non-stochastic contextual-bandit learner, of the LinEXP3 type, chooses the frequency-domain slice partition $\mathcal{F}_l$ at each super-frame, using the current virtual-queue and buffer-queue values (and, in Ad2S-NR, Kalman-extrapolated pathloss means $\hat{\mu}_n(l)$) as the context vector $X_l$.
What would settle it
Take a two-user, two-subchannel version, compute the optimal $F_k^*(\mathcal{F}_l)$ by exhaustive search over all binary allocations for a dense grid of backlog values, and fit $F_k^*$ against the context $X_l$: if the residual grows with $X_l$ or the slope depends on the allocation variables $b_n(t,f)$, the linearity assumption fails and the regret bound of Theorem 1 cannot be invoked. Alternatively, simulate Ad2S for large $L$ under i.i.d. channels and check whether the log-regret versus log-$L$ slope exceeds $2/3$ asymptotically.
Extended reading notes
Core claim
The paper's central claim is that the long-term throughput-maximization problem with mixed QoS constraints can be transformed, via the Lyapunov drift theorem, into a short-term utility maximization in which every slice choice is scored by a backlog-weighted rate sum. It then treats that utility as a non-stochastic reward and chooses slice configurations with a linear contextual-bandit learner whose context is the vector of virtual-queue and buffer-queue states, giving the Ad2S algorithm. The theorem states that with learning rate $\eta = L^{-2/3}(|\mathcal{F}|(1+4N))^{-1/3}(\log|\mathcal{F}|)^{2/3}$ and exploration parameter $\gamma = L^{-1/3}(|\mathcal{F}|(1+4N)\log|\mathcal{F}|)^{1/3}$, the cumulative regret after $L$ super-frames obeys $\mathrm{Regret}_L^{\mathrm{Ad2S}} \le 5 L^{2/3} (|\mathcal{F}|(1+4N)\log|\mathcal{F}|)^{1/3} F_{\max}$, and the non-stationary refinement Ad2S-NR, which adds ME-KF pathloss estimates to the context, satisfies the analogous bound with $1+8N$ in place of $1+4N$. Both bounds are sublinear in $L$, which is what it means for the slicing rule to be asymptotically optimal.
Load-bearing premise
The sublinear regret bound rests on the assumption that the per-frame utility $F_k^*(\mathcal{F}_l)$ produced by the resource-allocation solver is a linear function of the context vector fed to the bandit learner; the paper asserts this linearity rather than proving it, and the affine approximation it derives in the appendix has coefficients $\alpha_n(k)$, $\beta_n(k)$ that depend on the allocation decisions themselves.
Editorial extensions
If this is right
- If the regret bounds hold, an online slicing rule can approach the optimal configuration without knowing the channel statistics in advance, because the sublinear regret guarantees that the time-averaged utility converges to the best fixed slice policy in hindsight.
- Mixed 6G and legacy operation would no longer require hard statistical isolation: the simulations report a 3.86 Mbps throughput gain, a 63.96% latency reduction, and a 24.36% reduction in exploration time against the tested baselines.
- The numerical study reports 92.8% URLLC QoS satisfaction, meaning the probability that the frame's delivered data clears the queue within the delay bound, which is higher than the EXP3, contextual-UCB, and non-adaptive baselines.
- Operators can choose where to sit on the throughput-versus-latency trade-off by tuning the ratio $\omega_Q/\omega_T$ of the Lyapunov weights, with the suggested operating range $6\times 10^{-5}$ to $8\times 10^{-5}$.
Reading between the lines
- Editors' extension: the same virtual-queue-as-context recipe could be applied to other shared infrastructures—edge clouds, fronthaul, or open RAN midhaul—where hard-QoS and elastic flows compete, and the regret analysis would carry over as long as a linear-utility approximation is plausible.
- Editors' extension: a direct stress test is to run Ad2S on a scenario with a deliberately nonlinear utility (e.g., a fifth-order rate model); the regret is expected to degrade to linear, which would confirm that the linearity premise is the load-bearing part of Theorem 1.
- Editors' extension: the ME-KF tracker is replaceable; feeding a learned channel predictor (say, a sequence model) into the context vector would test whether the sublinear bound survives when the extrapolation is no longer provably unbiased.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript proposes a two-timescale radio access network slicing framework for multiplexing legacy eMBB/URLLC traffic with immersive MBBLL traffic. A Lyapunov drift transformation is used to convert the long-term throughput maximization into a per-frame weighted utility problem, a penalty-BCD algorithm (PBRA) is proposed for the frame-level mixed-integer resource allocation, and two contextual-bandit slicing algorithms (Ad2S and Ad2S-NR) are proposed for the super-frame-level slicing decisions, with an ME-KF channel tracker for non-stationary scenarios. The central theoretical claim is Theorem 1, which bounds the cumulative regret of both algorithms by O(L^{2/3}) under stationary and non-stationary channels. The paper also reports extensive numerical experiments showing throughput, latency, and convergence improvements over several baselines, including EXP3, contextual UCB, and a PPO-based scheme.
Significance. If the regret bounds were valid, the paper would provide a useful advance in non-stochastic network slicing with QoS awareness, and the numerical study is unusually thorough, covering multiple baselines, bursty traffic, varying traffic intensities, parameter sweeps, and an RL comparison. The ME-KF tracking component is a plausible engineering contribution, and the complexity analysis in Lemma 4 is useful. However, the theoretical guarantee is the paper's headline contribution, and it depends on a linear realizability condition that is neither proven nor generically true for the stated problem; the numerical results do not repair this gap because the regret plots are not accompanied by any alternative theoretical justification or by confidence intervals that would let the reader assess the variability of the reported gains.
major comments (4)
- [Section VI (Theorem 1), Appendix C, Appendix D, Eq. (35)] The claimed sublinear regret relies on the assertion that F*_k(F_l) is linear in the context vector X_l. This is not established and is generically false: F*_k is defined in Problem 3.1 as the maximum over {b_n(t,f), p(t,f)} of an affine objective, and the maximum of affine functions is convex and piecewise linear, with kinks at points where the optimal allocation changes. The context vector contains nonlinear features such as Q_n^2 and G_n Q_n, but the max operator breaks global linearity regardless of feature choice. Appendix D does not repair the gap: Eq. (35) gives r_n(k) = alpha_n(k) E[hat R_n(l)] + beta_n(k), with alpha_n(k) and beta_n(k) explicitly depending on the allocation variables b_n(t,f) and p(t,f), which are outputs of the optimization and therefore context-dependent. Hence the linear model required by the LinEXP3 regret bound in [45] is misspecified, and Theorem 1's sublinear regret bound does not follow.
- [Appendix A, proof of Lemma 1; Eq. (5), Eq. (15)] The proof drops the max operator in the queue dynamics (5) by asserting that r_n(k) is guaranteed to be less than Q_n(k). No constraint enforcing this inequality appears in Problem 1, Problem 2, or Problem 3.1, and the objective rewards large r_n(k), so the inequality is not guaranteed by the problem formulation. Without it, Eq. (30) and hence the drift bound (15) are not valid, and the claimed equivalence between Problem 1 and Problem 2 is not established. Even if the inequality did hold, minimizing the drift bound in (15) is a surrogate optimization, not an equivalent transformation, absent an optimality-gap analysis.
- [Appendix C, unbiased estimation step; Eq. (20)] The unbiasedness derivation assumes that there exists a true parameter theta_{l,f} such that the per-super-frame reward equals <X_l, theta_{l,f}>. Since the linearity assumption fails as described above, the estimator in Eq. (20) is the least-squares projection of the reward onto the span of X_l, not an unbiased estimator of a true parameter of the linear model. The importance-weighting factor I{...}/pi_l(f|X_l) does not restore the realizability condition required by the regret bound of [45], so the unbiased-estimation hypothesis of that theorem is not satisfied.
- [Algorithm 2, Eq. (19)] The softmax denominator in the policy is written as sum over f' != f of exp(...), omitting the current action f. As written, pi_l(.|X_l) is not a normalized probability distribution over slicing actions, so Algorithm 2 is not the LinEXP3 policy whose regret bound is cited. If this is a typographical error, the denominator should be corrected to sum over all f' in the action set; as it stands, the policy is not implementable as stated and the exploration probability gamma/|F| does not combine with the exponential term to give a valid distribution.
minor comments (5)
- [Section II-A, after Eq. (2)] The word "executively" should be "exclusively": each time-frequency element can be allocated to only one user.
- [Section III, first paragraph] The sentence "we first formulate ... and than discuss" contains a typo: "than" should be "then".
- [Theorem 1, Eq. (26) and Eq. (27)] The statement "approximately 1.3x additional regret" in the paragraph after Theorem 1 is only a numerical illustration; the ratio of the two bounds depends on N and |F| and is not a constant 1.3 in general.
- [Appendix B, Lemma 2 and Eq. (33)] The proof of Lemma 2 would be easier to follow if the notation for the penalty term and the surrogate function distinguished the iteration index of sigma from the block-coordinate inner iterations more clearly.
- [Fig. 4] The labels "1x" and "1.3x" in Fig. 4 are not defined in the caption; the reader has to infer that they refer to the regret overhead ratio claimed in the text.
Circularity Check
Theorem 1's sublinear regret bound rests on asserting—not proving—that the piecewise-defined optimum F*_k is linear in the context; Appendix C makes LinEXP3's central hypothesis true 'according to the definition' of F*.
-
self definitional
[Appendix C.A.1, bullet 3 and 'Unbiased Estimation' paragraph (Proof of Theorem 1)]
"According to the definition of F∗_k(Fl), for k∈[(l−1)×K̄+1,l×K̄], F∗_k(Fl) is linear with both {Gn(k)} and {Cn(k)}. ... E_l[θ̂_l,f] = E_l[ I{Fl=f}/π_l(f|X_l) (X_l^T X_l)^{-1} X_l (1/K̄) Σ F∗_k(Fl) ] = E_l[ E_l[I{Fl=f}/π_l(f|X_l)|X_l] θ_l,f ] = θ_l,f."
F∗_k(Fl) is defined in Problem 2 as the maximum over {bn(t,f),p(t,f)} of a linear objective in the rates rn(k); the maximum of affine functions is convex and piecewise-linear, not globally affine. Linearity therefore does not follow from the definition; the proof asserts it as if it did. The subsequent unbiasedness computation replaces (1/K̄)ΣF∗_k(Fl) by X_l^T θ_l,f, i.e. it assumes the very linear model whose validity is in question. Since the LinEXP3 bound of [45] applies only under a true linear reward, Theorem 1's regret bound rests on a hypothesis made true by assertion ('according to the definition') rather than by the system equations. This is a circular verification of the external theorem's premise.
-
other
[Appendix D, Eq. (35) and the following coefficient definitions]
"rn(k)=αn(k)E_l[R̂_n(l)]+βn(k), where αn(k)=Σ_{(t,f)∈T_k×F} {B, μ_n(l)≥ln(τ|F|/P_tot)/2; B b_n(t,f)p(t,f)e^{2N(0,σ_n^2)}, μ_n(l)<...}, βn(k)=Σ...{B log2(b_n(t,f)p(t,f))+2N(0,σ_n^2)log2(e), 0} are affine non-stochastic coefficients needed to be properly estimated."
αn(k) and βn(k) are defined through b_n(t,f) and p(t,f), which are the outputs of the frame-scale optimization (Algorithm 1) and therefore depend on the queue/virtual-queue context and on the chosen slice Fl. Calling these coefficients 'non-stochastic' does not make them fixed parameters of a true linear reward model: they are functions of the same allocation decisions whose aggregate value F∗_k(Fl) the bandit is trying to predict. The asserted affine representation does not exhibit the fixed θ required by LinEXP3; it defines the 'linear model' in terms of the very optimization outputs. Thus the non-stationary regret bound (27) inherits the same by-assertion verification of the linearity hypothesis.
full rationale
The paper's numerical evaluation is self-contained against external baselines, and the LinEXP3 regret theorem [45] is genuine external support. No load-bearing self-citation chain was found, and no fitted parameter is renamed as a prediction. The circularity is confined to the theoretical bridge between the frame-scale optimization and the bandit regret analysis. Appendix C must verify two hypotheses of [45]: unbiased estimation and linear reward in the context. The unbiasedness algebra is correct only if a fixed θ exists with (1/K̄)ΣF∗_k(Fl)=X_l^T θ; the existence of such θ is exactly the linearity hypothesis. The 'proof' of linearity consists of asserting that F∗_k(Fl), a maximum over allocations of a linear objective, is linear with {Gn(k)} and {Cn(k)} 'according to the definition'. A finite maximum of affine functions is piecewise-linear and generally not affine, so the premise is not derived. Appendix D does not repair this: its affine coefficients αn(k),βn(k) contain b_n(t,f),p(t,f), which are themselves outputs of the optimization and hence context-dependent. The central regret bound therefore reduces, at the critical step, to an assertion that the external theorem's premise holds by construction of the context vector rather than by the equations of the system. This is partial circularity in the theoretical derivation, which is why the score is 6 rather than 0-2; the empirical claims and the system model itself are not circular.
Assumptions & free parameters
free parameters (3)
- omega_Q =
5e-8 to 8e-8
- omega_T =
4e-4 to 1e-3
- tau =
1 dB
assumptions (6)
- standard math Lyapunov drift theorem provides an equivalent reformulation of the long-term throughput maximization as a drift-minus-utility minimization.
- standard math LinEXP3 regret bound from [45] holds when the reward model is linear in the context and the estimator is unbiased.
- ad hoc to paper The reward F*_k(F_l) is linear in the context vector X_l.
- domain assumption Channel magnitudes within each super-frame are i.i.d. and ergodic, with minimum ergodic period K_e < K.
- ad hoc to paper Throughput r_n(k) is always less than queue length Q_n(k), so the max operator in the queue dynamics can be dropped.
- domain assumption The pathloss follows the AR(1) model (13) with Gaussian process and measurement noise.
invented entities (1)
-
Virtual queue G_n(k)
independent evidence
Cite this review
Pith. "Pith review of A Unified QoS-Aware Multiplexing Framework for Next Generation Immersive Communication with Legacy Wireless Applications." pith.science (2026). https://pith.science/paper/XVPE5G6U
@misc{pith2026250421444,
author = {Pith},
title = {Pith review of: A Unified QoS-Aware Multiplexing Framework for Next Generation Immersive Communication with Legacy Wireless Applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVPE5G6U}},
note = {Machine review of arXiv:2504.21444}
}
read the original abstract
Immersive communication, including emerging augmented reality, virtual reality, and holographic telepresence, has been identified as a key service for enabling next-generation wireless applications. To align with legacy wireless applications, such as enhanced mobile broadband or ultra-reliable low-latency communication, network slicing has been widely adopted. However, attempting to statistically isolate the above types of wireless applications through different network slices may lead to throughput degradation and increased queue backlog. To address these challenges, we establish a unified QoS-aware framework that supports immersive communication and legacy wireless applications simultaneously. Based on the Lyapunov drift theorem, we transform the original long-term throughput maximization problem into an equivalent short-term throughput maximization weighted by virtual queue length. Moreover, to cope with the challenges introduced by the interaction between large-timescale network slicing and short-timescale resource allocation, we propose an adaptive adversarial slicing (Ad2S) scheme for networks with invarying channel statistics. To track the network channel variations, we also propose a measurement extrapolation-Kalman filter (ME-KF)-based method and refine our scheme into Ad2S-non-stationary refinement (Ad2S-NR). Through extended numerical examples, we demonstrate that our proposed schemes achieve 3.86 Mbps throughput improvement and 63.96% latency reduction with 24.36% convergence time reduction. Within our framework, the trade-off between total throughput and user service experience can be achieved by tuning systematic parameters.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[45]
Efficient and robust algorithms for adversarial linear contextual bandits,
G. Neu and J. Olkhovskaya, “Efficient and robust algorithms for adversarial linear contextual bandits,” in Proceedings of Thirty Third Conference on Learning Theory . PMLR, Jul. 2020, pp. 3049–3068
work page 2020
-
[1]
C. Gachet, “Recommendation itu-r m.2160-0 (11/2023) - framework and overall objectives of the future development of imt for 2030 and beyond,” Nov. 2023. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 16
work page 2023
-
[2]
Survey on 6G frontiers: trends, applications, require- ments, technologies and future research,
C. D. Alwis, A. Kalla, Q.-V . Pham, P. Kumar, K. Dev, W.-J. Hwang, and M. Liyanage, “Survey on 6G frontiers: trends, applications, require- ments, technologies and future research,” IEEE Open Journal of the Communications Society, vol. 2, p. 51, Apr. 2021
work page 2021
-
[3]
Toward low-latency and ultra-reliable virtual reality,
M. S. Elbamby, C. Perfecto, M. Bennis, and K. Doppler, “Toward low-latency and ultra-reliable virtual reality,” IEEE Network , vol. 32, no. 2, pp. 78–84, Mar. 2018
work page 2018
-
[4]
3GPP work plan - version december 6th 2024,
3GPP, “3GPP work plan - version december 6th 2024,” Tech. Rep., Dec. 2024
work page 2024
-
[5]
Joint spectrum reservation and on-demand request for mobile virtual network operators,
Y . Zhang, S. Bi, and Y .-J. A. Zhang, “Joint spectrum reservation and on-demand request for mobile virtual network operators,” IEEE Trans. Commun., vol. 66, no. 7, pp. 2966–2977, Jul. 2018
work page 2018
-
[6]
Joint scheduling of URLLC and eMBB traffic in 5G wireless networks,
A. Anand, “Joint scheduling of URLLC and eMBB traffic in 5G wireless networks,” IEEE/ACM Trans. Networking , vol. 28, no. 2, p. 14, Feb. 2020
work page 2020
-
[7]
Superposition- based URLLC traffic scheduling in 5G and beyond wireless networks,
M. Almekhlafi, M. A. Arfaoui, C. Assi, and A. Ghrayeb, “Superposition- based URLLC traffic scheduling in 5G and beyond wireless networks,” IEEE Trans. Commun. , vol. 70, no. 9, pp. 6295–6309, Sep. 2022
work page 2022
Show all 65 references
-
[8]
Resource slicing for eMBB and URLLC services in radio access network using hierarchical deep learning,
M. Setayesh, S. Bahrami, and V . W. Wong, “Resource slicing for eMBB and URLLC services in radio access network using hierarchical deep learning,” IEEE Trans. Wireless Commun. , vol. 21, no. 11, pp. 8950–8966, Nov. 2022
2022
-
[9]
Resource allocation and slicing puncture in cellular networks with eMBB and URLLC terminals coexistence,
Y . Zhao, X. Chi, L. Qian, Y . Zhu, and F. Hou, “Resource allocation and slicing puncture in cellular networks with eMBB and URLLC terminals coexistence,” IEEE Internet Things J. , vol. 9, no. 19, pp. 18 431–18 444, Oct. 2022
2022
-
[10]
Temporal characterization and prediction of VR traffic: a network slicing use case,
F. Chiariotti, M. Drago, P. Testolina, M. Lecci, A. Zanella, and M. Zorzi, “Temporal characterization and prediction of VR traffic: a network slicing use case,” IEEE Trans. Mob. Comput. , vol. 23, no. 5, pp. 3890–3908, May 2024
2024
-
[11]
Service multiplexing and revenue maximization in sliced c-RAN incorporated with URLLC and multicast eMBB,
J. Tang, B. Shim, and T. Q. S. Quek, “Service multiplexing and revenue maximization in sliced c-RAN incorporated with URLLC and multicast eMBB,” IEEE J. Select. Areas Commun. , vol. 37, no. 4, pp. 881–895, Feb. 2019
2019
-
[12]
Multicast eMBB and bursty URLLC service multiplexing in a comp-enabled RAN,
P. Yang, X. Xi, Y . Fu, T. Q. S. Quek, X. Cao, and D. Wu, “Multicast eMBB and bursty URLLC service multiplexing in a comp-enabled RAN,” IEEE Trans. Wireless Commun. , vol. 20, no. 5, pp. 3061–3077, May 2021
2021
-
[13]
How should I orchestrate resources of my slices for bursty URLLC service provision?
P. Yang, X. Xi, T. Q. S. Quek, J. Chen, X. Cao, and D. Wu, “How should I orchestrate resources of my slices for bursty URLLC service provision?” IEEE Trans. Commun., vol. 69, no. 2, pp. 1134–1146, Feb. 2021
2021
-
[14]
RAN slicing for massive IoT and bursty URLLC service multiplexing: analysis and optimization,
P. Yang, X. Xi, T. Q. S. Quek, J. Chen, X. Cao, and d. Wu, “RAN slicing for massive IoT and bursty URLLC service multiplexing: analysis and optimization,” IEEE Internet Things J. , vol. 8, no. 18, pp. 14 258–14 275, Sep. 2021
2021
-
[15]
Resource allocation for multi-traffic in cross- modal communications,
L. Wang, A. Yin, X. Jiang, M. Chen, K. Dev, N. M. Faseeh Qureshi, J. Yao, and B. Zheng, “Resource allocation for multi-traffic in cross- modal communications,” IEEE Trans. Netw. Serv. Manage. , vol. 20, no. 1, pp. 60–72, Mar. 2023
2023
-
[16]
Resource Allocation in an Open RAN System Using Network Slicing,
M. Karbalaee Motalleb, V . Shah-Mansouri, S. Parsaeefard, and O. L. Alcaraz L ´opez, “Resource Allocation in an Open RAN System Using Network Slicing,” IEEE Trans. Netw. Serv. Manage., vol. 20, no. 1, pp. 471–485, Mar. 2023
2023
-
[17]
Two-level soft RAN slicing for customized services in 5G-and-beyond wireless communications,
W. Shi, J. Li, P. Yang, Q. Ye, W. Zhuang, X. Shen, and X. Li, “Two-level soft RAN slicing for customized services in 5G-and-beyond wireless communications,” IEEE Trans. Ind. Inf. , vol. 18, no. 6, pp. 4169–4179, Jun. 2022
2022
-
[18]
Dynamic network slicing and resource allocation for heterogeneous wireless services,
J. Kwak, J. Moon, H.-W. Lee, and L. B. Le, “Dynamic network slicing and resource allocation for heterogeneous wireless services,” in 2017 IEEE 28th Annual International Symposium on Personal, Indoor, and Mobile Radio Communications (PIMRC) . Montreal, QC: IEEE, Oct. 2017, pp. 1–5
2017
-
[19]
Deep Reinforcement Learning for Optimization of RAN Slicing Relying on Control- and User-Plane Separation,
H. Tu, L. Zhao, Y . Zhang, G. Zheng, C. Feng, S. Song, and K. Liang, “Deep Reinforcement Learning for Optimization of RAN Slicing Relying on Control- and User-Plane Separation,” IEEE Internet Things J., vol. 11, no. 5, pp. 8485–8498, Mar. 2024
2024
-
[20]
Performance vs. Cost Tradeoff for Network Slicing in Open RAN: An Intelligent Hierarchical Algorithm for Flexible Utility-Control,
G. Zhou, L. Zhao, G. Zheng, S. Song, and K.-C. Chen, “Performance vs. Cost Tradeoff for Network Slicing in Open RAN: An Intelligent Hierarchical Algorithm for Flexible Utility-Control,” IEEE Trans. Veh. Technol., vol. 73, no. 11, pp. 17 697–17 713, Nov. 2024
2024
-
[21]
Intelligent radio access network slicing for service provisioning in 6G: a hierarchical deep reinforcement learning approach,
J. Mei, X. Wang, K. Zheng, G. Boudreau, A. B. Sediq, and H. Abou- Zeid, “Intelligent radio access network slicing for service provisioning in 6G: a hierarchical deep reinforcement learning approach,” IEEE Trans. Commun., vol. 69, no. 9, pp. 6063–6078, Sep. 2021
2021
-
[22]
On estimating the autoregressive coefficients of time-varying fading channels,
J. Vinogradova, G. Fodor, and P. Hammarberg, “On estimating the autoregressive coefficients of time-varying fading channels,” in 2022 IEEE 95th Vehicular Technology Conference: (VTC2022-Spring) , Jun. 2022, pp. 1–5, iSSN: 2577-2465
2022
-
[23]
A unified channel estimation framework for stationary and non-stationary fading environments,
Q. Shi, Y . Liu, S. Zhang, S. Xu, and V . K. N. Lau, “A unified channel estimation framework for stationary and non-stationary fading environments,” IEEE Trans. Commun. , vol. 69, no. 7, pp. 4937–4952, Jul. 2021
2021
-
[24]
Deep Reinforcement Learning for Scalable Dynamic Bandwidth Allocation in RAN Slicing With Highly Mobile Users,
S. Choi, S. Choi, G. Lee, S.-G. Yoon, and S. Bahk, “Deep Reinforcement Learning for Scalable Dynamic Bandwidth Allocation in RAN Slicing With Highly Mobile Users,” IEEE Trans. Veh. Technol., vol. 73, no. 1, pp. 576–590, Jan. 2024
2024
-
[25]
Feeling of presence maximization: mmWave-enabled virtual reality meets deep reinforcement learning,
P. Yang, T. Q. S. Quek, J. Chen, C. You, and X. Cao, “Feeling of presence maximization: mmWave-enabled virtual reality meets deep reinforcement learning,” IEEE Trans. Wireless Commun., vol. 21, no. 11, pp. 10 005–10 019, Nov. 2022
2022
-
[26]
laco: a latency-driven network slicing orchestration in beyond-5G networks,
L. Zanzi, V . Sciancalepore, A. Garcia-Saavedra, H. D. Schotten, and X. Costa-P ´erez, “laco: a latency-driven network slicing orchestration in beyond-5G networks,” IEEE Trans. Wireless Commun. , vol. 20, no. 1, pp. 667–682, Jan. 2021
2021
-
[27]
Learning With Side Information: Elastic Multi-Resource Control for the Open RAN,
X. Zhang, J. Zuo, Z. Huang, Z. Zhou, X. Chen, and C. Joe-Wong, “Learning With Side Information: Elastic Multi-Resource Control for the Open RAN,” IEEE J. Sel. Areas Commun. , vol. 42, no. 2, pp. 295–309, Feb. 2024
2024
-
[28]
Can terahertz provide high-rate reliable low-latency communications for wireless VR?
C. Chaccour, M. N. Soorki, W. Saad, M. Bennis, and P. Popovski, “Can terahertz provide high-rate reliable low-latency communications for wireless VR?” IEEE Internet Things J. , vol. 9, no. 12, pp. 9712–9729, Jun. 2022
2022
-
[29]
Online multi- user scheduling for XR transmissions with hard-latency constraint: performance analysis and practical design,
X. Zhao, Y .-J. A. Zhang, M. Wang, X. Chen, and Y . Li, “Online multi- user scheduling for XR transmissions with hard-latency constraint: performance analysis and practical design,” IEEE Trans. Commun., pp. 1–1, 2024
2024
-
[30]
Meta-reinforcement learning in non-stationary and dynamic environments,
Z. Bing, D. Lerch, K. Huang, and A. Knoll, “Meta-reinforcement learning in non-stationary and dynamic environments,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 3, pp. 3476–3491, Mar. 2023
2023
-
[31]
Little’s law,
J. D. C. Little and S. C. Graves, “Little’s law,” in Building Intuition: Insights From Basic Operations Management Models and Principles , D. Chhajed and T. J. Lowe, Eds. Boston, MA: Springer US, 2008, pp. 81–100
2008
-
[32]
Study on scenarios and requirements for next generation access technologies,
“Study on scenarios and requirements for next generation access technologies,” Apr. 2022, the 3rd Generation Partnership Project, 3GPP document 38.801
2022
-
[33]
NG-RAN; architecture description,
3GPP, “NG-RAN; architecture description,” 3GPP, Technical Specifica- tion TS 38.401, Mar. 2021, release 16
2021
-
[34]
Management and orchestration; concepts, use cases and require- ments,
——, “Management and orchestration; concepts, use cases and require- ments,” 3GPP, Technical Specification TS 28.530, Jun. 2023, release 18
2023
-
[35]
Autoregressive modeling for fading channel simulation,
K. Baddour and N. Beaulieu, “Autoregressive modeling for fading channel simulation,” IEEE Trans. Wireless Commun. , vol. 4, no. 4, pp. 1650–1662, Jul. 2005
2005
-
[36]
Adaptive data-aided time-varying channel tracking for massive mimo systems,
R. Chopra, C. R. Murthy, and K. Appaiah, “Adaptive data-aided time-varying channel tracking for massive mimo systems,” IEEE Trans. Commun., vol. 72, no. 9, pp. 5458–5472, Sep. 2024
2024
-
[37]
J. D. Hamilton, Time series analysis . Princeton, N.J: Princeton University Press, 1994
1994
-
[38]
A nonstationary poisson view of internet traffic,
T. Karagiannis, M. Molle, M. Faloutsos, and A. Broido, “A nonstationary poisson view of internet traffic,” in IEEE INFOCOM 2004, vol. 3. Hong Kong, PR China: IEEE, 2004, pp. 1558–1569
2004
-
[39]
Global optimization advances in mixed-integer nonlinear programming, minlp, and constrained derivative-free optimization, cdfo,
F. Boukouvala, R. Misener, and C. A. Floudas, “Global optimization advances in mixed-integer nonlinear programming, minlp, and constrained derivative-free optimization, cdfo,” Eur. J. Oper. Res. , vol. 252, no. 3, pp. 701–727, Aug. 2016
2016
-
[40]
Neely, Stochastic network optimization with application to com- munication and queueing systems , San Rafael, CA, USA: Morgan & Claypool, 2010
M. Neely, Stochastic network optimization with application to com- munication and queueing systems , San Rafael, CA, USA: Morgan & Claypool, 2010
2010
-
[41]
Dynamic programming,
R. Bellman, “Dynamic programming,” Sci., vol. 153, no. 3731, pp. 4–37, 1966
1966
-
[42]
The nonstochastic multiarmed bandit problem,
P. Auer, N. Cesa-Bianchi, Y . Freund, and R. E. Schapire, “The nonstochastic multiarmed bandit problem,” SIAM J. Comput. , vol. 32, no. 1, pp. 48–77, Jan. 2002
2002
-
[43]
S. P. Boyd and L. Vandenberghe, Convex optimization, version 29 ed. Cambridge New York Melbourne New Delhi Singapore: Cambridge University Press, 2023
2023
-
[44]
Network slicing for service-oriented networks under resource constraints,
N. Zhang, Y .-F. Liu, H. Farmanbar, T.-H. Chang, M. Hong, and Z.-Q. Luo, “Network slicing for service-oriented networks under resource constraints,” IEEE J. Select. Areas Commun. , vol. 35, no. 11, pp. 2512–2521, Nov. 2017. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 17
2017
-
[46]
New York, NY: Wiley, 1987
Anton and Howard, Elementary linear algebra, 5th ed. New York, NY: Wiley, 1987
1987
-
[47]
Evaluation and analysis of the performance of the EXP3 algorithm in stochastic environments,
Y . Seldin, C. Szepesv ´ari, P. Auer, and Y . Abbasi-Yadkori, “Evaluation and analysis of the performance of the EXP3 algorithm in stochastic environments,” in Proceedings of the Tenth European Workshop on Reinforcement Learning . PMLR, Jan. 2013, pp. 103–116, iSSN: 1938-7228
2013
-
[48]
The non-stationary stochastic multi-armed bandit problem,
R. Allesiardo, R. F ´eraud, and O.-A. Maillard, “The non-stationary stochastic multi-armed bandit problem,” International Journal of Data Science and Analytics , vol. 3, no. 4, pp. 267–283, Jun. 2017
2017
-
[49]
Incorporation of time delayed measurements in a discrete-time kalman filter,
T. Larsen, N. Andersen, O. Ravn, and N. Poulsen, “Incorporation of time delayed measurements in a discrete-time kalman filter,” in Proceedings of the 37th IEEE Conference on Decision and Control , vol. 4. Tampa, FL, USA: IEEE, 1998, pp. 3972–3977
1998
-
[50]
Linear estimation,
T.Kailauth, A. H. Sayed, and B. Hassibi, “Linear estimation,” IEEE Trans. Inf. Theory, vol. 51, no. 6, pp. 2236–2240, 2000
2000
-
[51]
Extrapolation of delayed measurements for fusion in a distributed sensor network,
R. A. J. Chagas and J. Waldmann, “Extrapolation of delayed measurements for fusion in a distributed sensor network,” in 2016 24th Mediterranean Conference on Control and Automation (MED) . Athens, Greece: IEEE, Jun. 2016, pp. 1319–1324
2016
-
[52]
Sequential multi-hypothesis testing in multi-armed bandit problems: an approach for asymptotic optimality,
G. R. Prabhu, S. Bhashyam, A. Gopalan, and R. Sundaresan, “Sequential multi-hypothesis testing in multi-armed bandit problems: an approach for asymptotic optimality,” IEEE Trans. Inf. Theory , vol. 68, no. 7, pp. 4790–4817, Jul. 2022
2022
-
[53]
Study on channel model for frequencies from 0.5 to 100 ghz,
3GPP, “Study on channel model for frequencies from 0.5 to 100 ghz,” Apr. 2024, the 3rd Generation Partnership Project, 3GPP document 38.901
2024
-
[54]
System architecture for the 5G system (5gs),
——, “System architecture for the 5G system (5gs),” Jun. 2024, the 3rd Generation Partnership Project, 3GPP document 23.501
2024
-
[55]
A contextual-bandit approach to personalized news article recommendation,
L. Li, W. Chu, J. Langford, and R. E. Schapire, “A contextual-bandit approach to personalized news article recommendation,” in Proceedings of the 19th international conference on World wide web . Raleigh North Carolina USA: ACM, Apr. 2010, pp. 661–670
2010
-
[56]
Stochastic multi-armed-bandit problem with non-stationary rewards,
O. Besbes, Y . Gur, and A. Zeevi, “Stochastic multi-armed-bandit problem with non-stationary rewards,” in Advances in Neural Information Processing Systems , vol. 27. Curran Associates, Inc., 2014
2014
-
[57]
Poulkov, Ed., Future Access Enablers for Ubiquitous and Intelligent Infrastructures, ser
V . Poulkov, Ed., Future Access Enablers for Ubiquitous and Intelligent Infrastructures, ser. Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering. Cham: Springer International Publishing, 2019, vol. 283
2019
-
[58]
Lyapunov Drift-Plus-Penalty Optimization for Queues With Finite Capacity,
L. Bracciale and P. Loreti, “Lyapunov Drift-Plus-Penalty Optimization for Queues With Finite Capacity,” IEEE Commun. Lett., vol. 24, no. 11, pp. 2555–2558, Nov. 2020
2020
-
[59]
Multiagent Reinforcement Learning: Methods, Trustworthiness, Applications in Intelligent Vehicles, and Challenges,
Z. Zhou, G. Liu, and Y . Tang, “Multiagent Reinforcement Learning: Methods, Trustworthiness, Applications in Intelligent Vehicles, and Challenges,” IEEE Trans. Intell. Veh., pp. 1–23, 2024
2024
-
[60]
A survey of admm variants for distributed optimization: problems, algorithms and features,
Y . Yang, X. Guan, Q.-S. Jia, L. Yu, B. Xu, and C. J. Spanos, “A survey of admm variants for distributed optimization: problems, algorithms and features,” Aug. 2022, arXiv:2208.03700 [cs]
2022 arXiv
-
[61]
Delay-oriented qos-aware user association and resource allocation in heterogeneous cellular networks,
X. Luo, “Delay-oriented qos-aware user association and resource allocation in heterogeneous cellular networks,” IEEE Trans. Wireless Commun., vol. 16, no. 3, pp. 1809–1822, Mar. 2017
2017
-
[62]
A unified algorithmic framework for block-structured optimization involving big data: with applications in machine learning and signal processing,
M. Hong, M. Razaviyayn, Z.-Q. Luo, and J.-S. Pang, “A unified algorithmic framework for block-structured optimization involving big data: with applications in machine learning and signal processing,” IEEE Signal Process Mag. , vol. 33, no. 1, pp. 57–77, Jan. 2016
2016
-
[63]
Strang, Introduction to Linear Algebra, Sixth Edition (2023) , 2011
G. Strang, Introduction to Linear Algebra, Sixth Edition (2023) , 2011
2023
-
[64]
Multiservice loss models in C-RAN supporting compound poisson traffic,
I.-A. Chousainov, I. Moscholios, P. Sarigiannidis, and M. Logothetis, “Multiservice loss models in C-RAN supporting compound poisson traffic,” Electronics, vol. 11, no. 5, p. 773, Mar. 2022
2022
-
[65]
Proximal policy optimization algorithms,
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” Aug. 2017, arXiv:1707.06347 [cs]. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 18 Supplementary Materials We demonstrate this supplementary results in respon...
2017 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.