Pith. sign in

REVIEW 3 major objections 5 minor 55 references

A quantum-enhanced reinforcement learning agent, Q-PPO, controls stacked intelligent metasurfaces to secure wireless transmissions, claiming about 15% higher secrecy rates and 30% faster convergence than classical deep RL under imperfect ea

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 06:43 UTC pith:CB2YECLW

load-bearing objection Solid adaptation of PPO-Q to SIM security, but the 15%/30% quantum advantage claim is unproven due to an architecture confound and single-run curves. the 3 major comments →

arxiv 2602.13238 v2 pith:CB2YECLW submitted 2026-01-29 cs.NI cs.LG

Securing SIM-Assisted Wireless Networks via Quantum Reinforcement Learning

classification cs.NI cs.LG
keywords quantum reinforcement learningstacked intelligent metasurfacesphysical-layer securityproximal policy optimizationparameterized quantum circuitssecrecy rateMISO downlinkimperfect CSI
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that replacing the classical actor network in proximal policy optimization with a small parameterized quantum circuit yields a better controller for a stacked intelligent metasurface (SIM) that shapes wireless signals for secure downlink transmission. The setting is a multi-user MISO system with a passive eavesdropper whose channel is only imperfectly known. The paper formulates joint optimization of transmit power and SIM phase shifts as a stochastic reinforcement learning problem and proposes Q-PPO, a hybrid classical-quantum actor-critic algorithm. Simulations reportedly show about 15% higher average secrecy rate and about 30% faster convergence than classical PPO, DDPG, and TD3. If correct, this would suggest quantum circuits can mitigate the high-dimensional, strongly coupled control burden of programmable metasurfaces in physical-layer security.

Core claim

On the paper's own terms, the central claim is that a parameterized quantum circuit can serve as the policy representation inside PPO for SIM-assisted secure MISO systems, and that this Q-PPO consistently beats classical DRL baselines: an optimal average secrecy rate of 1.67 bps/Hz, roughly 15% above conventional PPO, reached after about 20,000 training steps versus nearly 30,000 for PPO, under imperfect eavesdropper CSI. The paper attributes the gain to quantum superposition and entanglement enabling compact, expressive policy representations and more efficient exploration in the high-dimensional continuous action space of SIM phase shifts and power allocations.

What carries the argument

The central object is a hardware-efficient parameterized quantum circuit (PQC) embedded in the actor network of PPO, flanked by small pre-encoding and post-processing neural networks. The PQC consists of RY/RZ rotation gates and CZ entangling gates across five qubits and four layers, with data re-uploading; quantum measurements, converted through a softmax policy with an inverse-temperature parameter, produce actions. This hybrid architecture is intended to compress the policy representation and improve exploration while keeping the critic classical.

Load-bearing premise

The paper assumes that the 15% secrecy-rate and 30% convergence gains come from the quantum circuit in the actor, rather than from the different network sizes, architectures, hyperparameters, or random seeds used in the comparison, and that a classical simulation of that circuit faithfully represents quantum policy learning.

What would settle it

Retrain classical PPO with an actor whose parameter count and architecture roughly match the hybrid quantum-classical actor, using the same hyperparameters and multiple random seeds; if the secrecy-rate and convergence gaps shrink to near zero, the claim of quantum advantage collapses. Equivalently, run the same PQC on actual quantum hardware and check whether the simulated training curves reproduce.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the reported gains hold, Q-PPO offers a practical way to adapt SIM configurations in real time without requiring perfect knowledge of the eavesdropper's channel.
  • Q-PPO reportedly achieves a given secrecy level with fewer SIM layers than classical baselines, which could lower hardware and configuration complexity.
  • Increasing qubit count and PQC layer depth improves convergence and secrecy performance up to a point, suggesting a scaling path for larger metasurface arrays.
  • The paper reports fairness indices above 0.7 even as user count grows, so the quantum policy does not appear to improve secrecy at the cost of user fairness.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the reported 15% and 30% margins may reflect architecture and hyperparameter differences, such as the 1024-neuron classical PPO versus the small hybrid quantum-classical actor, rather than an inherent quantum advantage; a controlled comparison with matched parameter counts and repeated seeds would clarify this.
  • Editorial inference: if a classically simulated PQC is a faithful stand-in for real quantum hardware, the same hybrid architecture could transfer to near-term quantum processors, but the paper does not demonstrate that transfer.
  • Editorial inference: the approach plausibly generalizes to other high-dimensional continuous-control problems with strongly coupled action spaces, such as broader RIS or SIM beamforming tasks, though this extension is not tested here.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a hybrid quantum-classical proximal policy optimization (Q-PPO) framework for joint transmit power allocation and phase-shift design in a stacked intelligent metasurface (SIM)-assisted secure multi-user MISO downlink with imperfect eavesdropper CSI. The authors model the problem as an MDP, embed a parameterized quantum circuit (PQC) into the actor network with pre/post classical neural networks, and evaluate the scheme via simulation against PPO, DDPG, TD3, and random baselines. The central reported claims are an approximately 15% higher average secrecy rate and roughly 30% faster convergence for Q-PPO relative to conventional PPO (Abstract, Section IV-C, Fig. 3), along with further scalability and fairness comparisons in Figs. 4–6.

Significance. If the reported gains are real, the paper would provide a useful demonstration of a hybrid quantum-classical policy architecture for a high-dimensional, dynamically varying physical-layer security problem. The system model is clearly specified, the SIM channel and CSI-uncertainty formulation are concrete, and the comparison against several DRL baselines is a reasonable starting point. However, the manuscript does not yet provide the evidence needed to support the headline claim of a quantum-enhanced policy representation: the comparison is confounded by unequal actor architectures, no seed variability or error bars are reported, and the theoretical argument for parameter reduction is not connected to the actual hybrid architecture. The paper is therefore a plausible contribution whose central empirical claim requires substantial additional validation.

major comments (3)
  1. [Section IV-A and Fig. 3] The comparison between Q-PPO and classical PPO is confounded: the classical actor-critic uses four hidden layers of 1024 neurons each, while Q-PPO's actor uses a small Pre-NN (two convolutional layers of 128 and one FC of 64), a 5-qubit, 4-layer PQC, and a Post-NN (FC 62, FC 32). The two architectures differ by orders of magnitude in parameter count, layer type, and inductive bias. The 15% secrecy-rate gain and 30% convergence improvement attributed to the PQC could instead arise from the classical PPO being overparameterized and under-trained at the fixed learning rate and 40k-step horizon, or from the compact Q-PPO actor being easier to optimize. The authors must add a capacity-matched classical control, e.g., a classical PPO with a similarly compact MLP actor or an ablation that replaces the PQC with a classical nonlinear layer of matched input/output size, to isolate the effect of th
  2. [Section IV-C, Fig. 3] All convergence and performance comparisons in Fig. 3 (and Figs. 4–5) are based on single learning curves without multiple random seeds, error bars, confidence intervals, or statistical tests. In RL experiments, the variance across seeds is often substantial, and the reported 15%/30% margins may be within seed noise. The authors should report mean ± standard deviation over at least 5–10 seeds and define 'convergence speed' operationally (e.g., number of steps to reach 95% of the final ASR, or a similar measure). Without this, the quantitative claims in the abstract and Section IV-C are not statistically supported.
  3. [Section III-B (Eqs. 27–37)] The theoretical motivation for Q-PPO rests on assertions such as the PQC 'covers all Q classical parameters' and reduces the parameter count to O(poly(q)), but this is not demonstrated for the hybrid Pre-NN/PQC/Post-NN architecture actually used, and the PQC is classically simulated. The representational argument does not establish that the observed gains come from quantum entanglement or superposition. To support the claim of quantum-enhanced policy learning, the authors need either (i) a controlled experiment showing that a PQC with the same parameter budget outperforms a classical layer of matched capacity inside the same Pre/Post-NN architecture, or (ii) a rigorous analysis of the expressivity or trainability of the specific hybrid policy class. Without this, the phrase 'quantum advantage' in the abstract and Section IV-C is premature.
minor comments (5)
  1. [Section II-A, Eq. (5) and text after Eq. (6)] The definition of W^1 is inconsistent: the text says 'W^1 ∈ C^{N×M}, l∈L\{1}' but should refer to l=1. Also, 'matasurface' is a typo. Please correct.
  2. [Section III-B, Eq. (35)–(36)] The quantum policy is first defined as π(a|s)=⟨P_a⟩ (Eq. 35) and then redefined via a softmax with inverse temperature ζ (Eq. 36). The relationship between these two definitions, and how the trainable weights w_{a,i} in Eq. (37) enter the softmax, should be clarified in the text.
  3. [Algorithm 1, line 4] The symbol M is used both for the number of users in Section II and for the number of sampling iterations in Algorithm 1. Please use a different symbol (e.g., N_iter) to avoid ambiguity.
  4. [Figure captions and text] Typographical issues: 'Comparision' in Fig. 5 captions, '20.000' used in Section IV-C should be '20,000', and 'approximately30%' in the abstract lacks a space. Minor, but should be fixed.
  5. [References and novelty] The paper builds on the authors' conference version [1] and on the PPO-Q framework [35]. Please clarify explicitly which elements are novel beyond [1] (e.g., the full system model, additional baselines, and hyperparameter studies) and the extent to which the actor architecture is imported from [35].

Circularity Check

0 steps flagged

No significant circularity: Q-PPO's claimed advantage is an empirical simulation outcome, not a construction-equivalent prediction or a self-citation-forced result.

full rationale

The paper's central claim—that Q-PPO achieves ~15% higher secrecy rates and ~30% faster convergence—is supported by direct simulation curves (Fig. 3) rather than by fitting a parameter and then predicting a closely related quantity. The objective in (P1) and the reward in (20) are defined independently of the learning algorithm; Q-PPO is trained on that reward and its final secrecy rate is measured against baselines. This is an experimental comparison, not a derivation that reduces to its own inputs. The hybrid quantum-classical actor architecture is motivated by external prior work ([35], [50]), not by the authors' own unverified results, and the self-citations ([1], and background references such as [31]) are not load-bearing for the main claim. The architecture/hyperparameter mismatch between classical PPO (4×1024 neurons) and Q-PPO (Pre-NN, 5-qubit PQC, Post-NN) is a genuine experimental confound and a correctness risk, but it does not make the result circular: the observed margin is not forced by definition, nor is any fitted value renamed as a prediction. No equation in the paper is equivalent to its own input, and no uniqueness theorem or prior self-citation is invoked to forbid alternatives. Therefore, under the hard rules of this pass, the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 0 invented entities

The central claim rests entirely on a simulation whose free parameters (learning rate, q, eta, network sizes, unspecified softmax temperature) are hand-chosen, and whose domain assumptions (diffraction model, Rician fading, Gaussian Eve CSI error, zero-reward QoS handling) are standard but unverified. No new physical entity is introduced. The PQC is a known technique, not an invented entity.

free parameters (6)
  • learning_rate = 3e-4
    Selected as best from {3e-3, 3e-4, 3e-6} in Fig. 4a; all central Q-PPO results use this value.
  • num_qubits q = 5
    Selected as best from {3,4,5} in Fig. 4b; central results use q=5.
  • pqc_layers eta = 4
    Selected as best from {2,3,4} in Fig. 4c; central results use eta=4.
  • inverse_temperature zeta = not specified
    Controls exploration-exploitation in the softmax quantum policy (Eq. 36); no numerical value is given, so the reported behavior depends on an unspecified tuning choice.
  • actor/critic network sizes = PPO: 4x1024; Q-PPO Pre-NN conv 128 + FC 64, Post-NN 62 and 32
    Hand-chosen architectures; the comparison does not equalize parameter counts, so the claimed 15% gain may reflect capacity or architecture differences rather than quantum effects.
  • clip_epsilon and GAE lambda = 0.2 and 0.95
    Taken from Stable-Baselines3 defaults [54]; standard but still hand-set and directly affect the convergence-speed claims.
axioms (6)
  • domain assumption Rayleigh-Sommerfeld diffraction governs all SIM inter-layer and antenna-to-SIM propagation (Eqs. 5-6).
    The SIM transfer function G in Eq. 7 is built from these coefficients; if the diffraction model is inaccurate, all simulated channels and secrecy rates change.
  • domain assumption Channels from SIM to users/Eve follow Rician fading with a spatially correlated NLoS component (Eqs. 8-9).
    The secrecy-rate numbers depend directly on this channel model and its parameters.
  • domain assumption Eve's imperfect CSI is Gaussian additive error treated as interference in the Eve SINR (Eqs. 11 and 15).
    The whole 'imperfect CSI' setting is defined by this model; a different error model could alter the relative algorithm ranking.
  • domain assumption The SIM transfer function is the simple cascade G = Phi_L W_L ... W_1 with uniform layer spacing (Eqs. 4 and 7).
    All simulations assume this cascaded wave-domain model; it ignores multiple reflections and hardware imperfections.
  • ad hoc to paper A 5-qubit, 4-layer PQC simulated classically provides a faithful and advantageous quantum policy representation (Section III-B, IV-A).
    The quantum-advantage assertions in Section III-B.1 are cited from other works and are not demonstrated here; the central performance claim depends on this premise.
  • domain assumption The MDP state is fully captured by current legitimate CSI and imperfect Eve CSI, and rewards are zero when the QoS constraint is violated (Eqs. 18-20).
    The RL formulation and all trained policies depend on this Markovian state and reward definition.

pith-pipeline@v1.3.0-alltime-deepseek · 96 in / 12672 out tokens · 202055 ms · 2026-08-03T06:43:59.992639+00:00 · methodology

0 comments
read the original abstract

Stacked intelligent metasurfaces (SIMs) have recently emerged as a powerful wave-domain technology that enables multi-stage manipulation of electromagnetic signals through multilayer programmable architectures. While SIMs offer unprecedented degrees of freedom for enhancing physical-layer security, their extremely large number of meta-atoms leads to a high-dimensional and strongly coupled optimization space, making conventional design approaches inefficient and difficult to scale. Moreover, existing deep reinforcement learning (DRL) techniques suffer from slow convergence and performance degradation in dynamic wireless environments with imperfect knowledge of passive eavesdroppers. To address these challenges, we propose a hybrid quantum proximal policy optimization (QPPO) framework for SIM-assisted secure communications that jointly optimizes transmit power allocation and SIM phase shifts to maximize the average secrecy rate under power and quality-of-service constraints. Specifically, a parameterized quantum circuit is embedded into the actor network, forming a hybrid classical-quantum policy architecture that enhances policy representation capability and exploration efficiency in high-dimensional continuous action spaces. Extensive simulations demonstrate that the proposed Q-PPO scheme consistently outperforms DRL baselines, achieving approximately 15% higher secrecy rates and 30% faster convergence under imperfect eavesdropper channel state information. These results establish Q-PPO as a powerful optimization paradigm for SIM-enabled secure wireless networks.

Figures

Figures reproduced from arXiv: 2602.13238 by Diep N. Nguyen, Dinh Thai Hoang, Le-Hung Hoang, Quang-Trung Luu, Van-Dinh Nguyen.

Figure 1
Figure 1. Figure 1: The SIM-aided secure multi-user communication sys [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The proposed hybrid quantum-classical Q-PPO framework. [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The convergence of the different algorithms. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Comparison of proposed Q-PPO and benchmark schemes (PPO and Random), where we consider varying (a) learning [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Evaluation of the average secrecy rate with varying (a) number of meta atoms per layers ( [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Jain’s index vs. different number of CUs. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references · 4 linked inside Pith

  1. [1]

    Secure Multiuser Communications with Stacked Intelligent Metasurfaces Using Quantum Reinforcement Learning,

    L.-H. Hoang, M. H. Pham, Q.-T. Luu, and V .-D. Nguyen, “Secure Multiuser Communications with Stacked Intelligent Metasurfaces Using Quantum Reinforcement Learning,” inProc. Int. Conf. Adv. Tech. Commun. (ATC), Hanoi, Vietnam, 2025, pp. 1-6, 2025

  2. [2]

    6G Wireless Networks: Vision, Requirements, Architecture, and Key Technologies,

    Z. Zhang, Y . Xiao, Z. Ma, M. Xiao, Z. Ding, X. Lei, G. K. Karagiannidis, and P. Fan, “6G Wireless Networks: Vision, Requirements, Architecture, and Key Technologies,”IEEE V eh. Technol. Mag., vol. 14, no. 3, pp. 28–41, 2019

  3. [3]

    Terahertz communications for massive connec- tivity and security in 6G and beyond era,

    N. Yang and A. Shafie, “Terahertz communications for massive connec- tivity and security in 6G and beyond era,”IEEE Commun. Mag., vol. 62, no. 2, pp. 72–78, 2022

  4. [4]

    Introduction to wireless endogenous security and safety: Problems, attributes, structures and functions,

    L. Jin, X. Hu, Y . Lou, Z. Zhong, X. Sun, H. Wang, and J. Wu, “Introduction to wireless endogenous security and safety: Problems, attributes, structures and functions,”China Communications, vol. 18, no. 9, pp. 88–99, 2021

  5. [5]

    Safeguarding 5G wireless communication networks using physical layer security,

    N. Yang, L. Wang, G. Geraci, M. Elkashlan, J. Yuan, and M. Di Renzo, “Safeguarding 5G wireless communication networks using physical layer security,”IEEE Commun. Mag., vol. 53, no. 4, pp. 20–27, 2015

  6. [6]

    Guaranteeing secrecy using artificial noise,

    S. Goel and R. Negi, “Guaranteeing secrecy using artificial noise,”IEEE Trans. Wireless Commun., vol. 7, no. 6, pp. 2180–2189, 2008

  7. [7]

    Robust Transmit Beamform- ing for Secure Integrated Sensing and Communication,

    Z. Ren, L. Qiu, J. Xu, and D. W. K. Ng, “Robust Transmit Beamform- ing for Secure Integrated Sensing and Communication,”IEEE Trans. Commun., vol. 71, no. 9, pp. 5549–5564, May May, 2023

  8. [8]

    Beamforming Design for Physical Security in Movable Antenna-aided ISAC Systems: A Reinforcement Learning Approach,

    H. L. Hung, N. H. Huy, N. C. Luong, Q.-V . Pham, D. Niyato, and N. T. Hoa, “Beamforming Design for Physical Security in Movable Antenna-aided ISAC Systems: A Reinforcement Learning Approach,” IEEE Trans. V eh. Commun., pp. 1–5, 2025, Early Access

  9. [9]

    Joint Beamforming and Reflection Design for Secure RIS-ISAC Systems,

    J. Chu, Z. Lu, R. Liu, M. Li, and Q. Liu, “Joint Beamforming and Reflection Design for Secure RIS-ISAC Systems,”IEEE Trans. V eh. Commun., vol. 73, no. 3, pp. 4471–4475, 2024

  10. [10]

    Active Reconfigurable Intelligent Surface Aided Secure Transmission,

    L. Dong, H.-M. Wang, and J. Bai, “Active Reconfigurable Intelligent Surface Aided Secure Transmission,”IEEE Trans. V eh. Commun., vol. 71, no. 2, pp. 2181–2186, 2022

  11. [11]

    Enhancing the Physical Layer Security of Two-Way Relay Systems With RIS and Beamforming,

    Y . Zhang, S. Zhao, Y . Shen, X. Jiang, and N. Shiratori, “Enhancing the Physical Layer Security of Two-Way Relay Systems With RIS and Beamforming,”IEEE Trans. Inf. F orensics Security, vol. 19, pp. 5696– 5711, 2024

  12. [12]

    RIS- assisted ISAC systems for robust secure transmission with imperfect sense estimation,

    C. Jiang, C. Zhang, C. Huang, J. Ge, D. Niyato, and C. Yuen, “RIS- assisted ISAC systems for robust secure transmission with imperfect sense estimation,”IEEE Trans. Wireless Commun., 2025

  13. [13]

    Weighted Sum-Rate Maximization for Reconfigurable Intelligent Surface Aided Wireless Networks,

    H. Guo, Y .-C. Liang, J. Chen, and E. G. Larsson, “Weighted Sum-Rate Maximization for Reconfigurable Intelligent Surface Aided Wireless Networks,”IEEE Trans. Wireless Commun., vol. 19, no. 5, pp. 3064– 3076, 2020. SUBMITTED FOR POSSIBLE PUBLICATION 13

  14. [14]

    Stacked Intelligent Metasurfaces for Multiuser Downlink Beamforming in the Wave Domain,

    J. An, M. D. Renzo, M. Debbah, H. Vincent Poor, and C. Yuen, “Stacked Intelligent Metasurfaces for Multiuser Downlink Beamforming in the Wave Domain,”IEEE Trans. Wireless Commun., pp. 1–1, 2025, Early Access

  15. [15]

    Achievable rate optimization for stacked intelligent metasurface-assisted holographic MIMO communications,

    A. Papazafeiropoulos, J. An, P. Kourtessis, T. Ratnarajah, and S. Chatzinotas, “Achievable rate optimization for stacked intelligent metasurface-assisted holographic MIMO communications,”IEEE Trans. Wireless Commun., vol. 23, no. 10, pp. 13 173–13 186, 2024

  16. [16]

    Efficient beamforming and radiation pattern control using stacked intelligent metasurfaces,

    N. U. Hassan, J. An, M. Di Renzo, M. Debbah, and C. Yuen, “Efficient beamforming and radiation pattern control using stacked intelligent metasurfaces,”IEEE Open J. Commun. Soc., vol. 5, pp. 599–611, 2024

  17. [17]

    Stacked intelligent metasurfaces for wireless communications: Applications and challenges,

    H. Liu, J. An, X. Jia, L. Gan, G. K. Karagiannidis, B. Clerckx, M. Ben- nis, M. Debbah, and T. J. Cui, “Stacked intelligent metasurfaces for wireless communications: Applications and challenges,”IEEE Wireless Commun., vol. 32, no. 4, pp. 46–53, 2025

  18. [18]

    Emerging technologies in intelligent metasurfaces: Shaping the future of wireless communications,

    J. An, M. Debbah, T. J. Cui, Z. N. Chen, and C. Yuen, “Emerging technologies in intelligent metasurfaces: Shaping the future of wireless communications,”IEEE Trans. Antennas Propag., 2025

  19. [19]

    Enhancing Physical Layer Security for SISO Systems Using Stacked Intelligent Metasurfaces,

    H. Niu, J. An, L. Zhang, X. Lei, and C. Yuen, “Enhancing Physical Layer Security for SISO Systems Using Stacked Intelligent Metasurfaces,” in Proc. APWCS, 2024, pp. 1–5

  20. [20]

    On the Efficient Design of Stacked Intelligent Metasurfaces for Secure SISO Transmission,

    H. Niu, X. Lei, J. An, L. Zhang, and C. Yuen, “On the Efficient Design of Stacked Intelligent Metasurfaces for Secure SISO Transmission,”in IEEE Trans. Inf. F orensics Security, vol. 20, pp. 60–70, Jan. 2025

  21. [21]

    Secrecy Rate Maximization in the Presence of Stacked Intelligent Metasurface,

    M. R. Kavianinia, A. Mohammadi, and V . Meghdadi, “Secrecy Rate Maximization in the Presence of Stacked Intelligent Metasurface,”IEEE Trans. Inf. F orensics Security, 2025

  22. [22]

    Stacked Intelligent Metasurfaces for Multiuser Beamforming in the Wave Domain,

    J. An, M. Di Renzo, M. Debbah, and C. Yuen, “Stacked Intelligent Metasurfaces for Multiuser Beamforming in the Wave Domain,” inProc. ICC, 2023, pp. 2834–2839

  23. [23]

    Stacked Intelligent Metasurfaces for Efficient Holo- graphic MIMO Communications in 6G,

    J. An, C. Xu, D. W. K. Ng, G. C. Alexandropoulos, C. Huang, C. Yuen, and L. Hanzo, “Stacked Intelligent Metasurfaces for Efficient Holo- graphic MIMO Communications in 6G,”IEEE J. Sel. Areas Commun., vol. 41, no. 8, pp. 2380–2396, Aug. 2023

  24. [24]

    Channel estimation for stacked intelligent metasurface-assisted wireless networks,

    X. Yao, J. An, L. Gan, M. Di Renzo, and C. Yuen, “Channel estimation for stacked intelligent metasurface-assisted wireless networks,”IEEE Wireless Commun. Lett., vol. 13, no. 5, pp. 1349–1353, 2024

  25. [25]

    Hybrid digital-wave domain channel estimator for stacked intelligent metasurface enabled multi-user MISO systems,

    J. An, A. Chaabanet al., “Hybrid digital-wave domain channel estimator for stacked intelligent metasurface enabled multi-user MISO systems,” inProc. WCNC. IEEE, 2024, pp. 1–6

  26. [26]

    Stacked intelligent metasurfaces for task-oriented semantic communications,

    G. Huang, J. An, Z. Yang, L. Gan, M. Bennis, and M. Debbah, “Stacked intelligent metasurfaces for task-oriented semantic communications,” IEEE Wireless Commun. Lett., 2024

  27. [27]

    Spectrally encoded single-pixel machine vision using diffractive networks,

    J. Li, D. Mengu, N. T. Yardimci, Y . Luo, X. Li, M. Veli, Y . Rivenson, M. Jarrahi, and A. Ozcan, “Spectrally encoded single-pixel machine vision using diffractive networks,”Science Advances, vol. 7, no. 13, p. eabd7690, 2021

  28. [28]

    Distortion Resilience for Goal-Oriented Semantic Communication,

    M.-D. Nguyen, Q. V . Do, Z. Yang, Q.-V . Pham, and W.-J. Hwang, “Distortion Resilience for Goal-Oriented Semantic Communication,” IEEE Trans. Mobile Comput., vol. 24, no. 5, pp. 3489–3501, May 2025

  29. [29]

    Multi-User MISO with Stacked Intelligent Metasurfaces: A DRL-Based Sum-Rate Optimization Approach,

    H. Liu, J. An, G. C. Alexandropoulos, D. W. K. Ng, C. Yuen, and L. Gan, “Multi-User MISO with Stacked Intelligent Metasurfaces: A DRL-Based Sum-Rate Optimization Approach,”IEEE Trans. Cogn. Commun. Netw., pp. 1–1, 2025, Early Access

  30. [30]

    Joint SIM Configuration and Power Allocation for Stacked Intelligent Metasurface-assisted MU-MISO Systems with TD3,

    X. Yang, J. Zhang, E. Shi, Z. Liu, J. Liu, K. Zheng, and B. Ai, “Joint SIM Configuration and Power Allocation for Stacked Intelligent Metasurface-assisted MU-MISO Systems with TD3,” 2024. [Online]. Available: https://arxiv.org/abs/2408.05756

  31. [31]

    D. T. Hoang, N. Van Huynh, D. N. Nguyen, E. Hossain, and D. Niy- ato,Deep Reinforcement Learning for Wireless Communications and Networking: Theory, Applications and Implementation. John Wiley & Sons, 2023

  32. [32]

    Qtrl: Toward practical quantum reinforcement learning via quantum- train,

    C.-Y . Liu, C.-H. A. Lin, C.-H. H. Yang, K.-C. Chen, and M.-H. Hsieh, “Qtrl: Toward practical quantum reinforcement learning via quantum- train,” inProc. QCE, vol. 2. IEEE, 2024, pp. 317–322

  33. [33]

    Q-policy: Quantum-enhanced policy evaluation for scalable reinforcement learning,

    K. Cherukuri, A. Lala, and Y . Yardi, “Q-policy: Quantum-enhanced policy evaluation for scalable reinforcement learning,” 2025. [Online]. Available: https://arxiv.org/abs/2505.11862

  34. [34]

    An Introduction to Quantum Reinforcement Learning (QRL),

    S. Y .-C. Chen, “An Introduction to Quantum Reinforcement Learning (QRL),” inProc. ICTC. IEEE, 2024, pp. 1139–1144

  35. [35]

    PPO-Q: Proximal Policy Optimization with Parametrized Quantum Policies or Values,

    Y .-X. Jin, Z.-W. Wang, H.-Z. Xu, W.-F. Zhuang, M.-J. Hu, and D. E. Liu, “PPO-Q: Proximal Policy Optimization with Parametrized Quantum Policies or Values,” 2025. [Online]. Available: https://arxiv.org/abs/2501.07085

  36. [36]

    A programmable diffractive deep neural network based on a digital-coding metasurface array,

    C. Liu, Q. Ma, Z. J. Luo, Q. R. Hong, Q. Xiao, H. C. Zhang, L. Miao, W. M. Yu, Q. Cheng, L. Li, and T. J. Cui, “A programmable diffractive deep neural network based on a digital-coding metasurface array,” Nature Electronics, vol. 5, no. 2, pp. 113–122, Feb. 2022

  37. [37]

    Coding metamaterials, digital metamaterials and programmable metamaterials,

    T. J. Cui, M. Q. Qi, X. Wan, J. Zhao, and Q. Cheng, “Coding metamaterials, digital metamaterials and programmable metamaterials,” Light: Science & Applications, vol. 3, no. 10, p. e218, Oct. 2014

  38. [38]

    All-optical machine learning using diffractive deep neural networks,

    X. Lin, Y . Rivenson, N. T. Yardimci, M. Veli, Y . Luo, M. Jarrahi, and A. Ozcan, “All-optical machine learning using diffractive deep neural networks,”Science, vol. 361, no. 6406, pp. 1004–1008, Jul. 2018

  39. [39]

    Channel Estimation for Stacked Intelligent Metasurfaces in Rician Fading Channels,

    A. Papazafeiropoulos, P. Kourtessis, D. I. Kaklamani, and I. S. Venieris, “Channel Estimation for Stacked Intelligent Metasurfaces in Rician Fading Channels,”IEEE Wireless Commun. Lett., vol. 14, no. 5, pp. 1411–1415, Feb. 2025

  40. [40]

    Exploiting Amplitude Control in Intelligent Reflecting Surface Aided Wireless Communication With Imperfect CSI,

    M.-M. Zhao, Q. Wu, M.-J. Zhao, and R. Zhang, “Exploiting Amplitude Control in Intelligent Reflecting Surface Aided Wireless Communication With Imperfect CSI,”IEEE Trans. Commun., vol. 69, no. 6, pp. 4216– 4231, Jun. 2021

  41. [41]

    Robust Transmission Design for Intelligent Reflecting Surface-Aided Secure Communication Systems With Imperfect Cascaded CSI,

    S. Hong, C. Pan, H. Ren, K. Wang, K. K. Chai, and A. Nallanathan, “Robust Transmission Design for Intelligent Reflecting Surface-Aided Secure Communication Systems With Imperfect Cascaded CSI,”IEEE Trans. Wireless Commun., vol. 20, no. 4, pp. 2487–2501, Apr. 2021

  42. [42]

    Multi-hop RIS-empowered terahertz com- munications: A DRL-based hybrid beamforming design,

    C. Huang, Z. Yang, G. C. Alexandropoulos, K. Xiong, L. Wei, C. Yuen, Z. Zhang, and M. Debbah, “Multi-hop RIS-empowered terahertz com- munications: A DRL-based hybrid beamforming design,”IEEE J. Sel. Areas Commun., vol. 39, no. 6, pp. 1663–1677, 2021

  43. [43]

    Continuous Control with Deep Reinforcement Learning,

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Sil- ver, and D. Wierstra, “Continuous Control with Deep Reinforcement Learning,”Proc. ICLR, 2016

  44. [44]

    Addressing function approxima- tion error in actor-critic methods,

    S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approxima- tion error in actor-critic methods,” inProc. ICML, 2018, pp. 1587–1596

  45. [45]

    Proximal policy optimization algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017. [Online]. Available: http://arxiv.org/abs/1707.06347

  46. [46]

    Introduction to quantum reinforcement learning: Theory and pennylane-based imple- mentation,

    Y . Kwak, W. J. Yun, S. Jung, J.-K. Kim, and J. Kim, “Introduction to quantum reinforcement learning: Theory and pennylane-based imple- mentation,” inProc. ICTC. IEEE, 2021, pp. 416–420

  47. [47]

    Quantum-train: Rethinking hybrid quantum-classical machine learning in the model compression perspective,

    C.-Y . Liu, E.-J. Kuo, C.-H. Abraham Lin, J. Gemsun Young, Y .- J. Chang, M.-H. Hsieh, and H.-S. Goan, “Quantum-train: Rethinking hybrid quantum-classical machine learning in the model compression perspective,”Quantum Machine Intelligence, vol. 7, no. 2, p. 80, 2025

  48. [48]

    A secure information transmission protocol for healthcare cyber based on quantum image expansion and grover search algorithm,

    Z. Qu and H. Sun, “A secure information transmission protocol for healthcare cyber based on quantum image expansion and grover search algorithm,”IEEE Trans. Netw. Sci. Eng., vol. 10, no. 5, pp. 2551–2563, 2022

  49. [49]

    Quantum-enhanced drl optimization for doa estimation and task offloading in isac systems,

    A. Paul, K. Singh, A. Kaushik, C.-P. Li, O. A. Dobre, M. Di Renzo, and T. Q. Duong, “Quantum-enhanced drl optimization for doa estimation and task offloading in isac systems,”IEEE J. Sel. Areas Commun., 2024

  50. [50]

    Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets,

    A. Kandala, A. Mezzacapo, K. Temme, M. Takita, M. Brink, J. M. Chow, and J. M. Gambetta, “Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets,”Nature, vol. 549, no. 7671, pp. 242–246, 2017

  51. [51]

    Parametrized quantum policies for reinforcement learning,

    S. Jerbi, C. Gyurik, S. Marshall, H. Briegel, and V . Dunjko, “Parametrized quantum policies for reinforcement learning,”Advances in Neural Information Processing Systems, vol. 34, pp. 28 362–28 375, 2021

  52. [52]

    Quantum agents in the gym: a variational quantum algorithm for deep q-learning,

    A. Skolik, S. Jerbi, and V . Dunjko, “Quantum agents in the gym: a variational quantum algorithm for deep q-learning,”Quantum, vol. 6, p. 720, 2022

  53. [53]

    Playing atari with deep reinforcement learning,

    V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforcement learning,” 2013. [Online]. Available: https://arxiv.org/abs/1312.5602

  54. [54]

    Stable-Baselines3: Reliable Reinforcement Learning Implemen- tations,

    A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dor- mann, “Stable-Baselines3: Reliable Reinforcement Learning Implemen- tations,”Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021

  55. [55]

    A quantitative measure of fairness and discrimination,

    R. K. Jain, D.-M. W. Chiu, W. R. Haweet al., “A quantitative measure of fairness and discrimination,”Eastern Research Laboratory, Digital Equipment Corporation, Hudson, MA, vol. 21, no. 1, pp. 2022–2023, 1984