Pith. sign in

REVIEW 4 major objections 7 minor 3 cited by

Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review

T0 review · 4 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This survey claims that DRL-based cloud job scheduling and resource management is best organized by an algorithm-level taxonomy of four method families, and that this taxonomy fills a gap left by earlier surveys.

desk verdict Useful annotated bibliography, not a systematic review; fix the missing methodology and the placeholder before trusting its coverage. read the letter →

arxiv 2501.01007 v1 pith:2K7LO6HE submitted 2025-01-02 cs.DC cs.AI

classification cs.DCcs.AI
keywords deepreinforcementlearningcloudcomputingjobschedulingresourcemanagementworkflowprovisioningmulti-agentsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey tries to establish that the literature on deep reinforcement learning for cloud job scheduling and resource management is best understood at the level of the learning algorithm, and that it can be organized into value-based, policy-based, multi-agent, and advanced DRL families. It argues that existing surveys either focus on heuristic or meta-heuristic methods, or review DRL applications without analyzing the algorithms, leaving a gap that this algorithm-level review fills. The paper models task scheduling, workflow scheduling, resource provisioning, and resource scheduling as Markov decision processes, then classifies the works it reviews according to the DRL method used and the objectives they optimize, such as makespan, cost, energy, and quality of service. If the survey's picture is correct, a reader can rely on its four-way taxonomy and reference map to locate methods, compare design choices, and see where the field is heading, including privacy-aware scheduling, multi-tier resource management, large-scale hierarchical decision-making, and LLM-based scheduling.

What carries the argument

The load-bearing device is the four-family algorithm taxonomy, applied across four problem subdomains: task scheduling, workflow scheduling, resource provisioning, and resource scheduling. The paper defines each family by its learning objective: value-based methods learn an action-value function $Q(s,a)$, policy-based methods learn a policy $\pi(a|s)$ directly, multi-agent methods coordinate several agents under cooperative, competitive, or mixed reward structures, and advanced methods augment DRL with heuristics or quantum circuits. The MDP formulations for each subdomain, specifying the state space, action space, and reward function, are what make the reviewed works comparable within the taxonomy.

What would settle it

Count the DRL job-scheduling and resource-management papers returned by a systematic bibliographic search of the main computing literature databases and check whether the survey's reference list contains them and whether each falls cleanly into one of the four families; a substantial share of absent or unassignable papers would falsify the comprehensive and taxonomically clean picture.

Watch

Extended reading notes

Core claim

The paper's central claim is that prior reviews have not provided an algorithm-level analysis of DRL for cloud job scheduling and resource management, and that such an analysis is needed because DRL methods differ in how they represent states, actions, and rewards. It asserts that DRL overcomes the limitations of heuristic and meta-heuristic approaches, which rely on static models or predefined rules, by learning policies from continuous interaction with the environment. The survey organizes the field into four families: value-based methods such as DQN and its variants, policy-based methods such as actor-critic, PPO, and DDPG, multi-agent DRL with cooperative, competitive, and mixed settings, and advanced techniques that combine DRL with heuristics or quantum elements. It reviews task scheduling and DAG-based workflow scheduling as the two levels of job scheduling, and resource provisioning and resource scheduling as the two functions of resource management. It also explicitly broadens its scope to include edge and edge-cloud computing studies, arguing that these share enough settings with cloud environments that their methods are often applicable.

Load-bearing premise

The survey is only as valid as its implicit assumption that the selected references are a comprehensive and unbiased sample of DRL scheduling and resource-management work, since no search strategy, inclusion criteria, or coverage window is reported.

Editorial extensions

If this is right

  • A reader can use the four-family taxonomy to locate any DRL scheduling or resource-management method and compare how it models state, action, and reward.
  • The survey's argument implies DRL-based schedulers are viable where static heuristics fail, because they learn from continuous environment feedback rather than following predefined rules.
  • Including edge and edge-cloud studies implies those methods can be treated as part of the cloud scheduling toolkit unless stated otherwise.
  • If the survey is right, future work should concentrate on the four directions it names: privacy and security, multi-tier resource management, large-scale and hierarchical decision-making, and integration of large language models.
  • The classification implies that value-based methods dominate discrete-action scheduling while policy-based and multi-agent methods are used for continuous or distributed action spaces.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the taxonomy invites a benchmark study that runs one representative from each family on identical workload traces to test when value-based methods outperform policy-based ones.
  • Beyond the paper, the claim that edge-computing methods transfer to cloud settings could be directly tested by re-running those algorithms on cloud-scale traces, since the survey includes edge work without quantifying transferability.
  • Beyond the paper, the future-directions section implies LLM-based schedulers might reduce retraining cost on changing resource pools, and that is testable by comparing retraining frequency with and without LLM initialization.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This paper surveys deep reinforcement learning (DRL) methods for job scheduling and resource management in cloud computing, organizing the field into a four-way taxonomy (value-based, policy-based, multi-agent, and advanced DRL) and reviewing applications across task scheduling, workflow scheduling, resource provisioning, and resource scheduling. It also outlines future directions such as privacy and security, multi-tier networks, large-scale decision-making, and LLM integration. The paper's central claim is that it provides a comprehensive, algorithm-level review that bridges a gap in existing surveys.

Significance. If the survey's coverage and taxonomy are reliable, it would be a useful entry point for researchers: the taxonomy is clear, the tables aggregate a large number of recent works, and the future-directions section identifies timely topics such as hierarchical RL and LLM integration. The paper's main strength is its organizational framework rather than any new empirical result; the value of the survey depends entirely on whether the reference set is representative and whether the algorithmic descriptions are accurate. Currently, that value is undercut by the absence of a documented selection methodology, unresolved editorial issues in the tables, and a largely descriptive treatment of the individual papers.

major comments (4)
  1. [§I (Our Contributions) and Abstract] The claim of being a 'comprehensive review' is not supported by a reproducible literature search. The manuscript does not report the databases queried, search strings, inclusion/exclusion criteria, or date range, and it broadens the scope to include edge computing ('unless explicitly specified, we broaden the scope of this review to include such works') without defining when an edge paper is relevant. Please add a methodology section and report a systematic screening process; otherwise readers cannot verify that the roughly 219 references are representative or that the four-way taxonomy is a complete partition of the field.
  2. [Table V] Table V contains a row with an unresolved citation placeholder ('[?]') for a PPO-based resource scheduling method, and the same method is described in the text as work [207]. This is an editorial defect that must be fixed. Moreover, many rows across Tables II–V list only 'MADRL' or 'DRL variant' as the method; since the paper's contribution is an 'algorithm-level' review, each row should name the concrete algorithm (e.g., MAPPO, MADDPG, D3QN) rather than a coarse family label.
  3. [§III-D2] The section on Quantum Reinforcement Learning asserts that QNNs 'can achieve comparable or even superior performance with significantly fewer parameters' but provides no comparative evidence or citation to specific studies. As a review, the paper should either attribute this claim to prior work or qualify it, and it should report what QRL applications in cloud scheduling have actually demonstrated rather than presenting a speculative benefit as established.
  4. [§IV and §V] The individual paper discussions are largely descriptive ('this work applies X to optimize Y'), and the promised 'algorithm-level analysis' of methodologies is missing. There is no synthesis of common state/action/reward design choices, no critical comparison of algorithmic variants across the surveyed works, and no discussion of scalability, convergence, or robustness issues that would help a reader choose among approaches. Please add an analytical layer, such as comparative tables of problem formulations, a discussion of algorithmic trade-offs, or a set of design recommendations.
minor comments (7)
  1. [§II-D2 (after Eq. 8)] The text contains a typo: 'TThe type of resources' should be 'The type of resources'.
  2. [§III-A and throughout] The heading 'Valued-Based DRL Methods' uses 'Valued' where 'Value-Based' is the standard term; the same typo appears in several places.
  3. [Table IV, row [173]] The cell reads 'Rraining time cost' and should be 'Training time cost'.
  4. [§IV-B, description of [129]] 'eco-fridenly' should be 'eco-friendly'.
  5. [§V-B, description of [207]] 'scheudling' should be 'scheduling'.
  6. [§VI-A] The recommendation of RC4 for confidentiality and SHA-1 for integrity as 'foundational safeguards' is outdated; both algorithms are cryptographically broken and should not be endorsed. Suggest replacing with modern AEAD ciphers and collision-resistant hash functions.
  7. [Table I] The 'Algorithm-Level Reviewed Method' column is a binary checkbox, but the manuscript never defines what qualifies as an 'algorithm-level' review; the distinction from the existing reviews [19], [20], [21] should be made explicit.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey organizes existing literature and makes no empirical prediction that reduces to its own inputs.

full rationale

This is a literature survey, so the circularity patterns for derivations do not apply. The paper's central product is a taxonomy (value-based, policy-based, multi-agent, and advanced DRL) and a thematic summary of cited works. The taxonomy is a classification judgment, not a fitted parameter or a derived prediction; no equation in Sections II or III is used to generate a result that is then presented as independent evidence. The cited self-authored works (e.g., [11], [31], [41], [58], [106], [115], [116]) appear as examples of DRL-based scheduling and resource-management methods within the surveyed literature, not as load-bearing authorities that define the survey's conclusions. The absence of a search protocol, the unresolved citation placeholder in Table V, and the coarse 'DRL variant' labels are quality and verification concerns; they do not make the survey circular, because the survey does not claim to derive its reference list from a first-principles model. The comprehensiveness claim is an editorial assertion that would need a reproducible methodology to be fully auditable, but failing that standard is a completeness and correctness risk, not circular reasoning. Under the proportionality rule, this is a normal non-finding: score 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper makes no new empirical or theoretical claims; its central claim is descriptive. The main load-bearing assumptions are the accuracy of the secondary summaries and the validity of the chosen categorization.

assumptions (2)
  • domain assumption The reviewed papers are accurately summarized and their reported results are correct.
    The survey's classifications and descriptions rely on the trustworthiness of the cited works; the authors do not independently verify any experiment or result.
  • ad hoc to paper The four-way taxonomy (value-based, policy-based, multi-agent, advanced) is a complete and meaningful partition of DRL methods for this domain.
    The taxonomy is introduced by the authors without justification or comparison to alternative taxonomies; it is the organizing lens of the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review." pith.science (2026). https://pith.science/paper/2K7LO6HE

@misc{pith2026250101007,
  author       = {Pith},
  title        = {Pith review of: Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2K7LO6HE}},
  note         = {Machine review of arXiv:2501.01007}
}
read the original abstract

Cloud computing has revolutionized the provisioning of computing resources, offering scalable, flexible, and on-demand services to meet the diverse requirements of modern applications. At the heart of efficient cloud operations are job scheduling and resource management, which are critical for optimizing system performance and ensuring timely and cost-effective service delivery. However, the dynamic and heterogeneous nature of cloud environments presents significant challenges for these tasks, as workloads and resource availability can fluctuate unpredictably. Traditional approaches, including heuristic and meta-heuristic algorithms, often struggle to adapt to these real-time changes due to their reliance on static models or predefined rules. Deep Reinforcement Learning (DRL) has emerged as a promising solution to these challenges by enabling systems to learn and adapt policies based on continuous observations of the environment, facilitating intelligent and responsive decision-making. This survey provides a comprehensive review of DRL-based algorithms for job scheduling and resource management in cloud computing, analyzing their methodologies, performance metrics, and practical applications. We also highlight emerging trends and future research directions, offering valuable insights into leveraging DRL to advance both job scheduling and resource management in cloud computing.

Figures

Figures reproduced from arXiv: 2501.01007 by the authors.

Figure 1
Figure 1. The general architecture of this review • Comprehensive Survey of DRL-based Approaches: This work provides a detailed review of DRL approaches applied to job scheduling and resource management, including task and workflow scheduling, as well as re￾source provisioning and allocation strategies. We analyze how DRL techniques are customized to optimize these processes, offering insights into their performance, scal￾abi… view at source ↗
Figure 2
Figure 2. Detailed architectures of typical DRL approaches [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. The system architecture of job scheduling in cloud computing [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The system architecture of resource management in cloud computing [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A DRL agent learns bit-interleaving patterns for sparse tensor storage, delivering 1.4 to 1.46x geometric-mean speedups for tensor decomposition over expert formats.

  2. REACH: Reinforcement Learning for Efficient Allocation in Community and Heterogeneous Networks

    cs.NI 2025-08 unverdicted novelty 5.0 of 10

    REACH is claimed to improve task completion by up to 17%, double high-priority success, and cut bandwidth penalties by over 80% in simulations of community GPU scheduling.

  3. GrapheonRL: A Graph Neural Network and Reinforcement Learning Framework for Constraint and Data-Aware Workflow Mapping and Scheduling in Heterogeneous HPC Systems

    cs.DC 2025-05 reject novelty 4.0 of 10

    A GNN-PPO scheduler matches MILP makespans on small workflow benchmarks but is slower than MILP on the largest test, and the evaluation has serious methodological gaps.

Reference graph

Works this paper leans on

220 extracted references · 76 canonical work pages · cited by 3 Pith papers

  1. [207]

    Management of heteroge- neous cloud resources with use of the ppo,

    W. Funika, P. Koperek, and J. Kitowski, “Management of heteroge- neous cloud resources with use of the ppo,” in European Conference on Parallel Processing. Springer, 2020, pp. 148–159

  2. [1]

    D. C. Marinescu, Cloud Computing: Theory and Practice . Morgan Kaufmann, 2022

  3. [2]

    Scalable discovery of hybrid process models in a cloud computing environment,

    L. Cheng, B. F. van Dongen, and W. M. van der Aalst, “Scalable discovery of hybrid process models in a cloud computing environment,” IEEE Transactions on Services Computing, vol. 13, no. 2, pp. 368–380, 2020

  4. [3]

    Differentiate quality of experience scheduling for deep learning in- ferences with docker containers in the cloud,

    Y . Mao, W. Yan, Y . Song, Y . Zeng, M. Chen, L. Cheng, and Q. Liu, “Differentiate quality of experience scheduling for deep learning in- ferences with docker containers in the cloud,” IEEE Transactions on Cloud Computing, vol. 11, no. 2, pp. 1667–1677, 2023

  5. [4]

    Distributed artificial intelligence empowered by end-edge-cloud com- puting: A survey,

    S. Duan, D. Wang, J. Ren, F. Lyu, Y . Zhang, H. Wu, and X. Shen, “Distributed artificial intelligence empowered by end-edge-cloud com- puting: A survey,” IEEE Communications Surveys & Tutorials, vol. 25, no. 1, pp. 591–624, 2022

  6. [5]

    Resource scheduling in edge computing: A survey,

    Q. Luo, S. Hu, C. Li, G. Li, and W. Shi, “Resource scheduling in edge computing: A survey,” IEEE Communications Surveys & Tutorials , vol. 23, no. 4, pp. 2131–2165, 2021

  7. [6]

    A2c-drl: Dynamic scheduling for stochastic edge-cloud environments using a2c and deep reinforcement learning,

    J. Lu, J. Yang, S. Li, Y . Li, W. Jiang, J. Dai, and J. Hu, “A2c-drl: Dynamic scheduling for stochastic edge-cloud environments using a2c and deep reinforcement learning,” IEEE Internet of Things Journal , 2024

  8. [7]

    Cost-aware job scheduling for cloud instances using deep reinforce- ment learning,

    F. Cheng, Y . Huang, B. Tanpure, P. Sawalani, L. Cheng, and C. Liu, “Cost-aware job scheduling for cloud instances using deep reinforce- ment learning,” Cluster Computing, pp. 1–13, 2022

Show all 220 references
  1. [8]

    Dynamic On-Demand Ma- chine Provisioning and Continuous Resource Management,

    A. Chatterjee, P. Dubey, and A. Nigam, “Dynamic On-Demand Ma- chine Provisioning and Continuous Resource Management,” in 2023 International Conference on Computing, Communication, and Intelli- gent Systems, 2023, pp. 1010–1015

  2. [9]

    Real-time dynamic pricing for revenue manage- ment with reusable resources, advance reservation, and deterministic service time requirements,

    Y . Lei and S. Jasin, “Real-time dynamic pricing for revenue manage- ment with reusable resources, advance reservation, and deterministic service time requirements,” Operations Research, vol. 68, no. 3, pp. 676–685, 2020

  3. [10]

    Resource Allocation for Genera- tive AI Workloads: Advanced Cloud Resource Management Strategies for Optimized Model Performance,

    P. Murthy, A. Mehra, and L. Mishra, “Resource Allocation for Genera- tive AI Workloads: Advanced Cloud Resource Management Strategies for Optimized Model Performance,” Iconic Research And Engineering Journals, vol. 6, p. 12, 2023

  4. [11]

    Energy-aware systems for real-time job scheduling in cloud data centers: A deep reinforcement learning approach,

    J. Yan, Y . Huang, A. Gupta, A. Gupta, C. Liu, J. Li, and L. Cheng, “Energy-aware systems for real-time job scheduling in cloud data centers: A deep reinforcement learning approach,” Computers and Electrical Engineering, vol. 99, p. 107688, 2022

  5. [12]

    A low-cost multi-failure resilient replication scheme for high-data availability in cloud storage,

    J. Liu, H. Shen, H. Chi, H. S. Narman, Y . Yang, L. Cheng, and W. Chung, “A low-cost multi-failure resilient replication scheme for high-data availability in cloud storage,” IEEE/ACM Transactions on Networking, vol. 29, no. 4, pp. 1436–1451, 2021

  6. [13]

    Recent advances of resource allocation in network function virtualization,

    S. Yang, F. Li, S. Trajanovski, R. Yahyapour, and X. Fu, “Recent advances of resource allocation in network function virtualization,” IEEE Transactions on Parallel and Distributed Systems, vol. 32, no. 2, pp. 295–314, 2020

  7. [14]

    A survey on resource allocation for 5G heterogeneous networks: Current research, future trends, and challenges,

    Y . Xu, G. Gui, H. Gacanin, and F. Adachi, “A survey on resource allocation for 5G heterogeneous networks: Current research, future trends, and challenges,” IEEE Communications Surveys & Tutorials , vol. 23, no. 2, pp. 668–695, 2021

  8. [15]

    Task scheduling in cloud computing based on meta-heuristics: review, tax- onomy, open challenges, and future trends,

    E. H. Houssein, A. G. Gad, Y . M. Wazery, and P. N. Suganthan, “Task scheduling in cloud computing based on meta-heuristics: review, tax- onomy, open challenges, and future trends,” Swarm and Evolutionary Computation, vol. 62, p. 100841, 2021

  9. [16]

    A survey on PSO based meta- heuristic scheduling mechanism in cloud computing environment,

    A. Pradhan, S. K. Bisoy, and A. Das, “A survey on PSO based meta- heuristic scheduling mechanism in cloud computing environment,” Journal of King Saud University-Computer and Information Sciences , vol. 34, no. 8, pp. 4888–4901, 2022

  10. [17]

    Resource management and scheduling in distributed stream processing systems: a taxonomy, review, and future directions,

    X. Liu and R. Buyya, “Resource management and scheduling in distributed stream processing systems: a taxonomy, review, and future directions,” ACM Computing Surveys , vol. 53, no. 3, pp. 1–41, 2020

  11. [18]

    Resource allocation and task scheduling in fog computing and internet of everything environments: A taxonomy, review, and future directions,

    B. Jamil, H. Ijaz, M. Shojafar, K. Munir, and R. Buyya, “Resource allocation and task scheduling in fog computing and internet of everything environments: A taxonomy, review, and future directions,” ACM Computing Surveys , vol. 54, no. 11s, pp. 1–38, 2022

  12. [19]

    Resource allocation and service provisioning in multi-agent cloud robotics: A comprehensive survey,

    M. Afrin, J. Jin, A. Rahman, A. Rahman, J. Wan, and E. Hossain, “Resource allocation and service provisioning in multi-agent cloud robotics: A comprehensive survey,” IEEE Communications Surveys & Tutorials, vol. 23, no. 2, pp. 842–870, 2021

  13. [20]

    Deep reinforcement learning-based methods for resource scheduling in cloud computing: A review and future directions,

    G. Zhou, W. Tian, R. Buyya, R. Xue, and L. Song, “Deep reinforcement learning-based methods for resource scheduling in cloud computing: A review and future directions,” Artificial Intelligence Review, vol. 57, no. 5, p. 124, 2024

  14. [21]

    Deep rein- forcement learning-based scheduling in distributed systems: a critical review,

    Z. Jalali Khalil Abadi, N. Mansouri, and M. M. Javidi, “Deep rein- forcement learning-based scheduling in distributed systems: a critical review,” Knowledge and Information Systems , pp. 1–74, 2024

  15. [22]

    Comparative study of scheduling al-gorithms in cloud computing environment,

    I. A. Mohialdeen, “Comparative study of scheduling al-gorithms in cloud computing environment,” Journal of Computer Science , vol. 9, no. 2, pp. 252–263, 2013

  16. [23]

    User-priority guided Min-Min scheduling algorithm for load balancing in cloud computing,

    H. Chen, F. Wang, N. Helian, and G. Akanmu, “User-priority guided Min-Min scheduling algorithm for load balancing in cloud computing,” in 2013 National Conference on Parallel Computing Technologies , 2013, pp. 1–8

  17. [24]

    A heuristic- based task scheduling algorithm for scientific workflows in heteroge- neous cloud computing platforms,

    R. NoorianTalouki, M. H. Shirvani, and H. Motameni, “A heuristic- based task scheduling algorithm for scientific workflows in heteroge- neous cloud computing platforms,” Journal of King Saud University- Computer and Information Sciences , vol. 34, no. 8, pp. 4902–4913, 2022

  18. [25]

    PGA: a priority-aware genetic algorithm for task scheduling in heterogeneous fog-cloud computing,

    F. Hoseiny, S. Azizi, M. Shojafar, F. Ahmadiazar, and R. Tafazolli, “PGA: a priority-aware genetic algorithm for task scheduling in heterogeneous fog-cloud computing,” in IEEE INFOCOM 2021-IEEE Conference on Computer Communications Workshops , 2021, pp. 1–6

  19. [26]

    A WOA-based optimization approach for task scheduling in cloud computing systems,

    X. Chen, L. Cheng, C. Liu, Q. Liu, J. Liu, Y . Mao, and J. Murphy, “A WOA-based optimization approach for task scheduling in cloud computing systems,” IEEE Systems Journal , vol. 14, no. 3, pp. 3117– 3128, 2020

  20. [27]

    A Novel Distributed Fog-Based Networked Architecture to Preserve Energy in Fog Data Centers,

    Z. Pooranian, M. Shojafar, P. G. V . Naranjo, L. Chiaraviglio, and M. Conti, “A Novel Distributed Fog-Based Networked Architecture to Preserve Energy in Fog Data Centers,” in 2017 IEEE 14th International Conference on Mobile Ad Hoc and Sensor Systems, 2017, pp. 604–609

  21. [28]

    Resource Allocation Strategy in Fog Computing Based on Priced Timed Petri Nets,

    L. Ni, J. Zhang, C. Jiang, C. Yan, and K. Yu, “Resource Allocation Strategy in Fog Computing Based on Priced Timed Petri Nets,” IEEE Internet of Things Journal , vol. 4, no. 5, pp. 1216–1228, 2017. 25

  22. [29]

    A Multi- Objective Task Scheduling Method for Fog Computing in Cyber- Physical-Social Services,

    M. Yang, H. Ma, S. Wei, Y . Zeng, Y . Chen, and Y . Hu, “A Multi- Objective Task Scheduling Method for Fog Computing in Cyber- Physical-Social Services,” IEEE Access , vol. 8, pp. 65 085–65 095, 2020

  23. [30]

    Optimizing resource scheduling based on extended particle swarm optimization in fog computing envi- ronments,

    N. Potu, C. Jatoth, and P. Parvataneni, “Optimizing resource scheduling based on extended particle swarm optimization in fog computing envi- ronments,” Concurrency and Computation: Practice and Experience , vol. 33, no. 23, p. e6163, 2021

  24. [31]

    A deep reinforcement learning-based preemptive approach for cost-aware cloud job scheduling,

    L. Cheng, Y . Wang, F. Cheng, C. Liu, Z. Zhao, and Y . Wang, “A deep reinforcement learning-based preemptive approach for cost-aware cloud job scheduling,” IEEE Transactions on Sustainable Computing , vol. 9, no. 3, pp. 422–432, 2024

  25. [32]

    Reinforcement learning: An introduction,

    R. S. Sutton, “Reinforcement learning: An introduction,” A Bradford Book, 2018

  26. [33]

    Deep reinforcement learning: An overview,

    Y . Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274, 2017

  27. [34]

    Exploration in deep reinforcement learning: A survey,

    P. Ladosz, L. Weng, M. Kim, and H. Oh, “Exploration in deep reinforcement learning: A survey,” Information Fusion , vol. 85, pp. 1–22, 2022

  28. [35]

    Learning relation in crowd using gated graph convolutional networks for drl-based robot navigation,

    H. Jiang, N. Bhujel, Z. Lin, K.-W. Wan, J. Li, S. Jayavelu, and X. Jiang, “Learning relation in crowd using gated graph convolutional networks for drl-based robot navigation,” IEEE Transactions on Intelligent Transportation Systems, 2023

  29. [36]

    Marp: A cooperative multiagent drl system for connected autonomous vehicle platooning,

    S. Dai, S. Li, H. Tang, X. Ning, F. Fang, Y . Fu, Q. Wang, and L. Cheng, “Marp: A cooperative multiagent drl system for connected autonomous vehicle platooning,” IEEE Internet of Things Journal , vol. 11, no. 20, pp. 32 454–32 463, 2024

  30. [37]

    Deep reinforcement learning for communication flow control in wireless mesh networks,

    Q. Liu, L. Cheng, A. L. Jia, and C. Liu, “Deep reinforcement learning for communication flow control in wireless mesh networks,” IEEE Network, vol. 35, no. 2, pp. 112–119, 2021

  31. [38]

    Deep reinforcement learning for load-balancing aware network control in iot edge systems,

    Q. Liu, T. Xia, L. Cheng, M. Van Eijk, T. Ozcelebi, and Y . Mao, “Deep reinforcement learning for load-balancing aware network control in iot edge systems,” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 6, pp. 1491–1502, 2022

  32. [39]

    Taxonomies of workflow scheduling problem and techniques in the cloud,

    S. Smanchat and K. Viriyapant, “Taxonomies of workflow scheduling problem and techniques in the cloud,” Future Generation Computer Systems, vol. 52, pp. 1–12, 2015

  33. [40]

    An integrated optimization method to task scheduling and vm placement for green datacenters,

    H. Liu, X. Zhou, K. Gao, and Y . Ju, “An integrated optimization method to task scheduling and vm placement for green datacenters,” Simulation Modelling Practice and Theory , vol. 135, p. 102962, 2024

  34. [41]

    Job scheduling in hybrid clouds with privacy constraints: A deep reinforcement learning approach,

    H. He, Y . Gu, Q. Liu, H. Wu, and L. Cheng, “Job scheduling in hybrid clouds with privacy constraints: A deep reinforcement learning approach,” Concurrency and Computation: Practice and Experience , p. e8307, 2024

  35. [42]

    Security-driven heuristics and a fast genetic algorithm for trusted grid job scheduling,

    S. Song, Y .-K. Kwok, and K. Hwang, “Security-driven heuristics and a fast genetic algorithm for trusted grid job scheduling,” in 19th IEEE International Parallel and Distributed Processing Symposium , 2005

  36. [43]

    Scheduling jobs across geo- distributed datacenters with max-min fairness,

    L. Chen, S. Liu, B. Li, and B. Li, “Scheduling jobs across geo- distributed datacenters with max-min fairness,” IEEE Transactions on Network Science and Engineering , vol. 6, no. 3, pp. 488–500, 2018

  37. [44]

    Pricing schemes in cloud computing: an overview,

    A. Mazrekaj, I. Shabani, and B. Sejdiu, “Pricing schemes in cloud computing: an overview,”International Journal of Advanced Computer Science and Applications , vol. 7, no. 2, pp. 80–86, 2016

  38. [45]

    Qos aware job scheduling in a cluster-based web server for multimedia applications,

    J. Guo, L. Bhuyan, R. Kumar, and S. Basu, “Qos aware job scheduling in a cluster-based web server for multimedia applications,” in 19th IEEE International Parallel and Distributed Processing Symposium , 2005

  39. [46]

    Approximate data mapping in refresh-free dram for energy-efficient computing in modern mobile systems,

    S. Li, H. Jin, Y . Gao, Y . Wang, S. Dai, Y . Xu, and L. Cheng, “Approximate data mapping in refresh-free dram for energy-efficient computing in modern mobile systems,” Computer Communications , vol. 216, pp. 151–158, 2024

  40. [47]

    Energy-latency trade-off analysis for scientific workflow in cloud environments: the role of processor utilization ratio and mean grey wolf optimizer,

    M. I. Khaleel, M. Safran, S. Alfarhood, and M. Zhu, “Energy-latency trade-off analysis for scientific workflow in cloud environments: the role of processor utilization ratio and mean grey wolf optimizer,” Engineering Science and Technology, an International Journal, vol. 50, p...

  41. [48]

    A data and task co-scheduling algorithm for scientific cloud workflows,

    K. Deng, K. Ren, M. Zhu, and J. Song, “A data and task co-scheduling algorithm for scientific cloud workflows,” IEEE Transactions on Cloud Computing, vol. 8, no. 2, pp. 349–362, 2015

  42. [49]

    Performance-effective and low-complexity task scheduling for heterogeneous computing,

    H. Topcuoglu, S. Hariri, and M.-Y . Wu, “Performance-effective and low-complexity task scheduling for heterogeneous computing,” IEEE transactions on parallel and distributed systems , vol. 13, no. 3, pp. 260–274, 2002

  43. [50]

    Endpoint communication contention- aware cloud workflow scheduling,

    Q. Wu, M. Zhou, and J. Wen, “Endpoint communication contention- aware cloud workflow scheduling,” IEEE Transactions on Automation Science and Engineering , vol. 19, no. 2, pp. 1137–1150, 2022

  44. [51]

    A deep reinforcement learning based algorithm for time and cost optimized scaling of serverless applications,

    A. Mampage, S. Karunasekera, and R. Buyya, “A deep reinforcement learning based algorithm for time and cost optimized scaling of serverless applications,” arXiv preprint arXiv:2308.11209 , 2023

  45. [52]

    A Hierarchical Framework of Cloud Resource Allocation and Power Management Using Deep Reinforcement Learning,

    N. Liu, Z. Li, J. Xu, Z. Xu, S. Lin, Q. Qiu, J. Tang, and Y . Wang, “A Hierarchical Framework of Cloud Resource Allocation and Power Management Using Deep Reinforcement Learning,” in 2017 IEEE 37th International Conference on Distributed Computing Systems, 2017, pp. 372–382

  46. [53]

    Deep reinforce- ment learning-based computation offloading and resource allocation in security-aware mobile edge computing,

    H. C. Ke, H. Wang, H. W. Zhao, and W. J. Sun, “Deep reinforce- ment learning-based computation offloading and resource allocation in security-aware mobile edge computing,” Wireless Networks, vol. 27, no. 5, pp. 3357–3373, 2021

  47. [54]

    Computing resource allocation scheme of IOV using deep reinforcement learning in edge computing environment,

    Y . Zhang, M. Zhang, C. Fan, F. Li, and B. Li, “Computing resource allocation scheme of IOV using deep reinforcement learning in edge computing environment,” EURASIP Journal on Advances in Signal Processing, vol. 2021, no. 1, p. 33, 2021

  48. [55]

    Multi-Agent Deep Rein- forcement Learning-Based Resource Allocation in HPC/AI Converged Cluster,

    J. Narantuya, J.-S. Shin, S. Park, and J. Kim, “Multi-Agent Deep Rein- forcement Learning-Based Resource Allocation in HPC/AI Converged Cluster,” Computers, Materials & Continua , vol. 72, no. 3, pp. 4375– 4395, 2022

  49. [56]

    Multi-Agent Deep Reinforcement Learning-Based Partial Task Offloading and Resource Allocation in Edge Computing Environment,

    H. Ke, H. Wang, and H. Sun, “Multi-Agent Deep Reinforcement Learning-Based Partial Task Offloading and Resource Allocation in Edge Computing Environment,” Electronics, vol. 11, no. 15, p. 2394, 2022

  50. [57]

    Rcsearcher: Reaction center identification in retrosynthesis via deep Q-learning,

    Z. Lan, Z. Zeng, B. Hong, Z. Liu, and F. Ma, “Rcsearcher: Reaction center identification in retrosynthesis via deep Q-learning,” Pattern Recognition, vol. 150, p. 110318, 2024

  51. [58]

    Deep reinforcement learning for efficient iot data compression in smart railroad management,

    X. Chen, Q. Yu, S. Dai, P. Sun, H. Tang, and L. Cheng, “Deep reinforcement learning for efficient iot data compression in smart railroad management,” IEEE Internet of Things Journal, vol. 11, no. 15, pp. 25 494–25 504, 2024

  52. [59]

    Policy learning with constraints in model-free reinforcement learning: A survey,

    Y . Liu, A. Halev, and X. Liu, “Policy learning with constraints in model-free reinforcement learning: A survey,” inThe 30th International Joint Conference on Artificial Intelligence , 2021

  53. [60]

    Actor-critic algorithms,

    V . Konda and J. Tsitsiklis, “Actor-critic algorithms,” Advances in Neural Information Processing Systems , vol. 12, 1999

  54. [61]

    Performance analysis of deep q networks and advantage actor critic algorithms in designing reinforcement learning-based self-tuning pid controllers,

    R. Mukhopadhyay, S. Bandyopadhyay, A. Sutradhar, and P. Chattopad- hyay, “Performance analysis of deep q networks and advantage actor critic algorithms in designing reinforcement learning-based self-tuning pid controllers,” in 2019 IEEE Bombay Section Signature Conference . IE...

  55. [62]

    Exact numerical simulation of the ornstein-uhlenbeck process and its integral,

    D. T. Gillespie, “Exact numerical simulation of the ornstein-uhlenbeck process and its integral,” Phys. Rev. E, vol. 54, pp. 2084–2091, 1996

  56. [63]

    Pantheonrl: A marl library for dynamic training interactions,

    B. Sarkar, A. Talati, A. Shih, and D. Sadigh, “Pantheonrl: A marl library for dynamic training interactions,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 11, 2022, pp. 13 221– 13 223

  57. [64]

    Imp-marl: a suite of environments for large-scale infrastructure management plan- ning via marl,

    P. Leroy, P. G. Morato, J. Pisane, A. Kolios, and D. Ernst, “Imp-marl: a suite of environments for large-scale infrastructure management plan- ning via marl,” Advances in Neural Information Processing Systems , vol. 36, 2024

  58. [65]

    Multi-agent reinforcement learning as a rehearsal for decentralized planning,

    L. Kraemer and B. Banerjee, “Multi-agent reinforcement learning as a rehearsal for decentralized planning,” Neurocomputing, vol. 190, pp. 82–94, 2016

  59. [66]

    Multi- uav cooperative search based on reinforcement learning with a digital twin driven training framework,

    G. Shen, L. Lei, X. Zhang, Z. Li, S. Cai, and L. Zhang, “Multi- uav cooperative search based on reinforcement learning with a digital twin driven training framework,” IEEE Transactions on Vehicular Technology, vol. 72, pp. 8354–8368, 2023

  60. [67]

    Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient,

    S. Li, Y . Wu, X. Cui, H. Dong, F. Fang, and S. Russell, “Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 4213–4220

  61. [68]

    Independent policy gradient methods for competitive reinforcement learning,

    C. Daskalakis, D. J. Foster, and N. Golowich, “Independent policy gradient methods for competitive reinforcement learning,” Advances in Neural Information Processing Systems, vol. 33, pp. 5527–5540, 2020

  62. [69]

    Distributed nash equilibrium seek- ing in games with partial decision information: A survey,

    M. Ye, Q. Han, L. Ding, and S. Xu, “Distributed nash equilibrium seek- ing in games with partial decision information: A survey,” Proceedings of the IEEE , vol. 111, pp. 140–157, 2023

  63. [70]

    ’other-pla‘ for zero- shot coordination,

    H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster, “’other-pla‘ for zero- shot coordination,” in International Conference on Machine Learning , 2020, pp. 4399–4410

  64. [71]

    Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity,

    K. Zhang, S. Kakade, T. Basar, and L. Yang, “Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity,” Advances in Neural Information Processing Systems, vol. 33, pp. 1166– 1178, 2020. 26

  65. [72]

    A td3-based multi-agent deep reinforce- ment learning method in mixed cooperation-competition environment,

    F. Zhang, J. Li, and Z. Li, “A td3-based multi-agent deep reinforce- ment learning method in mixed cooperation-competition environment,” Neurocomputing, vol. 411, pp. 206–215, 2020

  66. [73]

    Phase transitions in random circuit sampling,

    A. Morvan, B. Villalonga, X. Mi, and et al., “Phase transitions in random circuit sampling,” Nature, vol. 634, no. 8033, pp. 328–333, Oct. 2024

  67. [74]

    Demonstration of logical qubits and repeated error correction with better-than-physical error rates,

    M. Da Silva, C. Ryan-Anderson, J. Bello-Rivas, A. Chernoguzov, J. Dreiling, C. Foltz, J. Gaebler, T. Gatterman, D. Hayes, N. Hewitt et al. , “Demonstration of logical qubits and repeated error correction with better-than-physical error rates,” arXiv preprint, 2024

  68. [75]

    Drlbtsa: Deep reinforcement learning based task-scheduling algorithm in cloud computing,

    S. Mangalampalli, G. R. Karri, M. Kumar, O. I. Khalaf, C. A. T. Romero, and G. A. Sahib, “Drlbtsa: Deep reinforcement learning based task-scheduling algorithm in cloud computing,” Multimedia Tools and Applications, vol. 83, no. 3, pp. 8359–8387, 2024

  69. [76]

    Drl-cloud: Deep reinforcement learning-based resource provisioning and task scheduling for cloud ser- vice providers,

    M. Cheng, J. Li, and S. Nazarian, “Drl-cloud: Deep reinforcement learning-based resource provisioning and task scheduling for cloud ser- vice providers,” in 2018 23rd Asia and South pacific design automation conference. IEEE, 2018, pp. 129–134

  70. [77]

    Adaptive drl-based task scheduling for energy-efficient cloud computing,

    K. Kang, D. Ding, H. Xie, Q. Yin, and J. Zeng, “Adaptive drl-based task scheduling for energy-efficient cloud computing,” IEEE Transactions on Network and Service Management , vol. 19, no. 4, pp. 4948–4961, 2021

  71. [78]

    Optimizing energy efficiency for data center via parameterized deep reinforcement learning,

    Y . Ran, H. Hu, Y . Wen, and X. Zhou, “Optimizing energy efficiency for data center via parameterized deep reinforcement learning,” IEEE Transactions on Services Computing , vol. 16, no. 2, pp. 1310–1323, 2022

  72. [79]

    An agent-based model for resource provisioning and task scheduling in cloud computing using drl,

    T. Oudaa, H. Gharsellaoui, and S. B. Ahmed, “An agent-based model for resource provisioning and task scheduling in cloud computing using drl,” Procedia Computer Science , vol. 192, pp. 3795–3804, 2021

  73. [80]

    Deep reinforcement learning enhanced greedy optimization for online scheduling of batched tasks in cloud hpc systems,

    Y . Yang and H. Shen, “Deep reinforcement learning enhanced greedy optimization for online scheduling of batched tasks in cloud hpc systems,” IEEE Transactions on Parallel and Distributed Systems , vol. 33, no. 11, pp. 3003–3014, 2021

  74. [81]

    Ddqn-ts: A novel bi- objective intelligent scheduling algorithm in the cloud environment,

    Z. Tong, F. Ye, B. Liu, J. Cai, and J. Mei, “Ddqn-ts: A novel bi- objective intelligent scheduling algorithm in the cloud environment,” Neurocomputing, vol. 455, pp. 419–430, 2021

  75. [82]

    Task scheduling for mobile edge computing leveraging deep reinforcement learning,

    F. Han, N. Yu, J. Gong, Y . Ge, and X. Gao, “Task scheduling for mobile edge computing leveraging deep reinforcement learning,” in 2023 42nd Chinese Control Conference. IEEE, 2023, pp. 1921–1926

  76. [83]

    Online dispatching and fair scheduling of edge computing tasks: A learning- based approach,

    H. Yuan, G. Tang, X. Li, D. Guo, L. Luo, and X. Luo, “Online dispatching and fair scheduling of edge computing tasks: A learning- based approach,” IEEE Internet of Things Journal , vol. 8, no. 19, pp. 14 985–14 998, 2021

  77. [84]

    Improved double deep q network-based task scheduling algorithm in edge computing for makespan optimization,

    L. Zeng, Q. Liu, S. Shen, and X. Liu, “Improved double deep q network-based task scheduling algorithm in edge computing for makespan optimization,” Tsinghua Science and Technology , vol. 29, no. 3, pp. 806–817, 2023

  78. [85]

    A double deep q-learning model for energy-efficient edge scheduling,

    Q. Zhang, M. Lin, L. T. Yang, Z. Chen, S. U. Khan, and P. Li, “A double deep q-learning model for energy-efficient edge scheduling,” IEEE Transactions on Services Computing , vol. 12, no. 5, pp. 739– 749, 2018

  79. [86]

    Optimal policy characterization enhanced proximal policy optimization for multitask scheduling in cloud computing,

    J. Jin and Y . Xu, “Optimal policy characterization enhanced proximal policy optimization for multitask scheduling in cloud computing,”IEEE Internet of Things Journal , vol. 9, no. 9, pp. 6418–6433, 2021

  80. [87]

    Cloud task scheduling based on proximal policy optimization algorithm for lowering energy consumption of data center

    Y . Yang, C. He, B. Yin, Z. Wei, and B. Hong, “Cloud task scheduling based on proximal policy optimization algorithm for lowering energy consumption of data center.” KSII Transactions on Internet & Infor- mation Systems, vol. 16, no. 6, 2022

  81. [88]

    A deep reinforcement learning approach to resource management in hybrid clouds harnessing renewable energy and task scheduling,

    J. Zhao, M. A. Rodr ´ıguez, and R. Buyya, “A deep reinforcement learning approach to resource management in hybrid clouds harnessing renewable energy and task scheduling,” in 2021 IEEE 14th Interna- tional Conference on Cloud Computing . IEEE, 2021, pp. 240–249

  82. [89]

    Slas-aware online task scheduling based on deep reinforcement learning method in cloud environment,

    L. Ran, X. Shi, and M. Shang, “Slas-aware online task scheduling based on deep reinforcement learning method in cloud environment,” in 2019 IEEE 21st International Conference on High Performance Computing and Communications, 2019, pp. 1518–1525

  83. [90]

    Performance and cost-aware task scheduling via deep reinforcement learning in cloud environment,

    Z. Zhao, X. Shi, and M. Shang, “Performance and cost-aware task scheduling via deep reinforcement learning in cloud environment,” in International Conference on Service-Oriented Computing . Springer, 2022, pp. 600–615

  84. [91]

    Etpam: An efficient task pre- assignment and migration algorithm in heterogeneous edge-cloud computing environments,

    F. Zhang, L. Jiang, and J. Chen, “Etpam: An efficient task pre- assignment and migration algorithm in heterogeneous edge-cloud computing environments,” in 2024 27th International Conference on Computer Supported Cooperative Work in Design . IEEE, 2024, pp. 2400–2405

  85. [92]

    Dynamic scheduling for stochastic edge-cloud computing environments using a3c learning and residual recurrent neural networks,

    S. Tuli, S. Ilager, K. Ramamohanarao, and R. Buyya, “Dynamic scheduling for stochastic edge-cloud computing environments using a3c learning and residual recurrent neural networks,” IEEE transactions on mobile computing , vol. 21, no. 3, pp. 940–954, 2020

  86. [93]

    Online joint scheduling of delay-sensitive and computation-oriented tasks in edge computing,

    F. Zhang, Z. Tang, J. Lou, and W. Jia, “Online joint scheduling of delay-sensitive and computation-oriented tasks in edge computing,” in 2019 15th International Conference on Mobile Ad-Hoc and Sensor Networks. IEEE, 2019, pp. 303–308

  87. [94]

    Dynamic task allocation and service migration in edge-cloud iot system based on deep reinforcement learning,

    Y . Chen, Y . Sun, C. Wang, and T. Taleb, “Dynamic task allocation and service migration in edge-cloud iot system based on deep reinforcement learning,” IEEE Internet of Things Journal , vol. 9, no. 18, pp. 16 742– 16 757, 2022

  88. [95]

    Reliability-aware: task scheduling in cloud computing using multi-agent reinforcement learn- ing algorithm and neural fitted q

    H. A. Balla, C. G. Sheng, and W. Jing, “Reliability-aware: task scheduling in cloud computing using multi-agent reinforcement learn- ing algorithm and neural fitted q.” Int. Arab J. Inf. Technol. , vol. 18, no. 1, pp. 36–47, 2021

  89. [96]

    Orchestrated scheduling and multi-agent deep reinforcement learning for cloud- assisted multi-UA V charging systems,

    S. Jung, W. J. Yun, M. Shin, J. Kim, and J.-H. Kim, “Orchestrated scheduling and multi-agent deep reinforcement learning for cloud- assisted multi-UA V charging systems,”IEEE Transactions on Vehicular Technology, vol. 70, no. 6, pp. 5362–5377, 2021

  90. [97]

    Multi-agent deep reinforcement learning for collabo- rative task scheduling

    M. I. Gergely, “Multi-agent deep reinforcement learning for collabo- rative task scheduling.” in ICAART (3), 2024, pp. 1076–1083

  91. [98]

    Multi- agent deep reinforcement learning for online request scheduling in edge cooperation networks,

    Y . Zhang, R. Li, Y . Zhao, R. Li, Y . Wang, and Z. Zhou, “Multi- agent deep reinforcement learning for online request scheduling in edge cooperation networks,” Future Generation Computer Systems, vol. 141, pp. 258–268, 2023

  92. [99]

    Distributed task scheduling in serverless edge computing networks for the internet of things: A learning approach,

    Q. Tang, R. Xie, F. R. Yu, T. Chen, R. Zhang, T. Huang, and Y . Liu, “Distributed task scheduling in serverless edge computing networks for the internet of things: A learning approach,” IEEE Internet of Things Journal, vol. 9, no. 20, pp. 19 634–19 648, 2022

  93. [100]

    Multi-agent reinforcement learning based distributed transmission in collaborative cloud-edge systems,

    C. Xu, S. Liu, C. Zhang, Y . Huang, Z. Lu, and L. Yang, “Multi-agent reinforcement learning based distributed transmission in collaborative cloud-edge systems,” IEEE Transactions on Vehicular Technology , vol. 70, no. 2, pp. 1658–1672, 2021

  94. [101]

    Lsia3cs: Deep reinforcement learning-based cloud-edge collaborative task scheduling in large-scale iiot,

    Z. Zhang, F. Zhang, Z. Xiong, K. Zhang, and D. Chen, “Lsia3cs: Deep reinforcement learning-based cloud-edge collaborative task scheduling in large-scale iiot,” IEEE Internet of Things Journal , 2024

  95. [102]

    Load balancing for task scheduling based on multi-agent reinforcement learning in cloud-edge-end collab- orative environments,

    Z. Li, J. Yu, X. Liu, and L. Peng, “Load balancing for task scheduling based on multi-agent reinforcement learning in cloud-edge-end collab- orative environments,” in Proceedings of the 2024 8th International Conference on Machine Learning and Soft Computing , 2024, pp. 94– 100

  96. [103]

    Task placement and resource allocation for edge machine learning: A gnn-based multi-agent reinforcement learning paradigm,

    Y . Li, X. Zhang, T. Zeng, J. Duan, C. Wu, D. Wu, and X. Chen, “Task placement and resource allocation for edge machine learning: A gnn-based multi-agent reinforcement learning paradigm,” IEEE Transactions on Parallel and Distributed Systems , 2023

  97. [104]

    Multiagent meta-reinforcement learning for optimized task scheduling in heterogeneous edge computing systems,

    L. Niu, X. Chen, N. Zhang, Y . Zhu, R. Yin, C. Wu, and Y . Cao, “Multiagent meta-reinforcement learning for optimized task scheduling in heterogeneous edge computing systems,” IEEE Internet of Things Journal, vol. 10, no. 12, pp. 10 519–10 531, 2023

  98. [105]

    Multi-agent driven resource allocation and interference management for deep edge networks,

    Y . Gong, H. Yao, J. Wang, L. Jiang, and F. R. Yu, “Multi-agent driven resource allocation and interference management for deep edge networks,” IEEE Transactions on Vehicular Technology, vol. 71, no. 2, pp. 2018–2030, 2021

  99. [106]

    Deep adversarial imitation reinforcement learning for qos-aware cloud job scheduling,

    Y . Huang, L. Cheng, L. Xue, C. Liu, Y . Li, J. Li, and T. Ward, “Deep adversarial imitation reinforcement learning for qos-aware cloud job scheduling,” IEEE Systems Journal , vol. 16, no. 3, pp. 4232–4242, 2022

  100. [107]

    Deep smart scheduling: A deep learning approach for automated big data schedul- ing over the cloud,

    G. Rjoub, J. Bentahar, O. A. Wahab, and A. Bataineh, “Deep smart scheduling: A deep learning approach for automated big data schedul- ing over the cloud,” in 2019 7th International Conference on Future Internet of Things and Cloud . IEEE, 2019, pp. 189–196

  101. [108]

    Batch jobs load balancing scheduling in cloud computing using distributional reinforcement learn- ing,

    T. Li, S. Ying, Y . Zhao, and J. Shang, “Batch jobs load balancing scheduling in cloud computing using distributional reinforcement learn- ing,” IEEE Transactions on Parallel and Distributed Systems , vol. 35, no. 1, pp. 169–185, 2023

  102. [109]

    DL-DRL: A double-level deep reinforcement learning approach for large-scale task scheduling of multi-UA V,

    X. Mao, G. Wu, M. Fan, Z. Cao, and W. Pedrycz, “DL-DRL: A double-level deep reinforcement learning approach for large-scale task scheduling of multi-UA V,”IEEE Transactions on Automation Science and Engineering, 2024

  103. [110]

    Representation and reinforcement learning for task scheduling in edge computing,

    Z. Tang, W. Jia, X. Zhou, W. Yang, and Y . You, “Representation and reinforcement learning for task scheduling in edge computing,” IEEE Transactions on Big Data , vol. 8, no. 3, pp. 795–808, 2020

  104. [111]

    Scalable parallel task scheduling for autonomous driving using multi- task deep reinforcement learning,

    Q. Qi, L. Zhang, J. Wang, H. Sun, Z. Zhuang, J. Liao, and F. R. Yu, “Scalable parallel task scheduling for autonomous driving using multi- task deep reinforcement learning,” IEEE Transactions on Vehicular Technology, vol. 69, no. 11, pp. 13 861–13 874, 2020. 27

  105. [112]

    A scheduling scheme in the cloud computing environment using deep q-learning,

    Z. Tong, H. Chen, X. Deng, K. Li, and K. Li, “A scheduling scheme in the cloud computing environment using deep q-learning,” Information Sciences, vol. 512, pp. 1170–1191, 2020

  106. [113]

    A deep reinforcement learning-based scheduling framework for real-time workflows in the cloud environment,

    J. Pan and Y . Wei, “A deep reinforcement learning-based scheduling framework for real-time workflows in the cloud environment,” Expert Systems with Applications , vol. 255, p. 124845, 2024

  107. [114]

    Reinforce- ment learning based scheduling in a workflow management system,

    A. M. Kintsakis, F. E. Psomopoulos, and P. A. Mitkas, “Reinforce- ment learning based scheduling in a workflow management system,” Engineering Applications of Artificial Intelligence, vol. 81, pp. 94–106, 2019

  108. [115]

    Cost-aware cloud workflow scheduling using drl and simulated annealing,

    Y . Gu, F. Cheng, L. Yang, J. Xu, X. Chen, and L. Cheng, “Cost-aware cloud workflow scheduling using drl and simulated annealing,” Digital Communications and Networks , 2024

  109. [116]

    Cost-aware scheduling systems for real-time workflows in cloud: An approach based on genetic algorithm and deep reinforcement learning,

    J. Zhang, L. Cheng, C. Liu, Z. Zhao, and Y . Mao, “Cost-aware scheduling systems for real-time workflows in cloud: An approach based on genetic algorithm and deep reinforcement learning,” Expert Systems with Applications , vol. 234, p. 120972, 2023

  110. [117]

    Deep reinforcement learning for fault-tolerant workflow scheduling in cloud environment,

    T. Dong, F. Xue, H. Tang, and C. Xiao, “Deep reinforcement learning for fault-tolerant workflow scheduling in cloud environment,” Applied Intelligence, vol. 53, no. 9, pp. 9916–9932, 2023

  111. [118]

    Weighted double deep q- network based reinforcement learning for bi-objective multi-workflow scheduling in the cloud,

    H. Li, J. Huang, B. Wang, and Y . Fan, “Weighted double deep q- network based reinforcement learning for bi-objective multi-workflow scheduling in the cloud,” Cluster Computing, vol. 25, no. 2, pp. 751– 768, 2022

  112. [119]

    A collaborative scheduling method for cloud computing heterogeneous workflows based on deep reinforcement learning,

    G. Chen, J. Qi, Y . Sun, X. Hu, Z. Dong, and Y . Sun, “A collaborative scheduling method for cloud computing heterogeneous workflows based on deep reinforcement learning,” Future Generation Computer Systems, vol. 141, pp. 284–297, 2023

  113. [120]

    Data-intensive workflow scheduling strategy based on deep reinforcement learning in multi- clouds,

    S. Zhang, Z. Zhao, C. Liu, and S. Qin, “Data-intensive workflow scheduling strategy based on deep reinforcement learning in multi- clouds,” Journal of Cloud Computing , vol. 12, no. 1, p. 125, 2023

  114. [121]

    EdgeTuner: Fast scheduling algorithm tuning for dynamic edge-cloud workloads and resources,

    R. Han, S. Wen, C. H. Liu, Y . Yuan, G. Wang, and L. Y . Chen, “EdgeTuner: Fast scheduling algorithm tuning for dynamic edge-cloud workloads and resources,” in IEEE INFOCOM 2022-IEEE Conference on Computer Communications , 2022, pp. 880–889

  115. [122]

    Deep reinforcement learning-based workload scheduling for edge computing,

    T. Zheng, J. Wan, J. Zhang, and C. Jiang, “Deep reinforcement learning-based workload scheduling for edge computing,” Journal of Cloud Computing, vol. 11, no. 1, p. 3, 2022

  116. [123]

    A workflow-aided Internet of things paradigm with intelligent edge computing,

    Y . Qian, L. Shi, J. Li, Z. Wang, H. Guan, F. Shu, and H. V . Poor, “A workflow-aided Internet of things paradigm with intelligent edge computing,” IEEE Network, vol. 34, no. 6, pp. 92–99, 2020

  117. [124]

    Task Decomposition and Hi- erarchical Scheduling for Collaborative Cloud-Edge-End Computing,

    J. Cai, W. Liu, Z. Huang, and F. R. Yu, “Task Decomposition and Hi- erarchical Scheduling for Collaborative Cloud-Edge-End Computing,” IEEE Transactions on Services Computing , 2024

  118. [125]

    Workflow scheduling in serverless edge computing for the industrial internet of things: A learning approach,

    R. Xie, D. Gu, Q. Tang, T. Huang, and F. R. Yu, “Workflow scheduling in serverless edge computing for the industrial internet of things: A learning approach,” IEEE Transactions on Industrial Informatics , vol. 19, no. 7, pp. 8242–8252, 2022

  119. [126]

    Workflow scheduling based on deep reinforcement learning in the cloud environment,

    T. Dong, F. Xue, C. Xiao, and J. Zhang, “Workflow scheduling based on deep reinforcement learning in the cloud environment,” Journal of Ambient Intelligence and Humanized Computing , vol. 12, no. 12, pp. 10 823–10 835, 2021

  120. [127]

    Deep reinforcement learning for dynamic workflow scheduling in cloud environment,

    ——, “Deep reinforcement learning for dynamic workflow scheduling in cloud environment,” in 2021 IEEE International Conference on Services Computing. IEEE, 2021, pp. 107–115

  121. [128]

    Dag-based workflows scheduling using actor–critic deep reinforcement learning,

    G. P. Koslovski, K. Pereira, and P. R. Albuquerque, “Dag-based workflows scheduling using actor–critic deep reinforcement learning,” Future Generation Computer Systems , vol. 150, pp. 354–363, 2024

  122. [129]

    Reinforcement learning based task scheduling for envi- ronmentally sustainable federated cloud computing,

    Z. Wang, S. Chen, L. Bai, J. Gao, J. Tao, R. R. Bond, and M. D. Mulvenna, “Reinforcement learning based task scheduling for envi- ronmentally sustainable federated cloud computing,” Journal of Cloud Computing, vol. 12, no. 1, p. 174, 2023

  123. [130]

    Towards efficient workflow scheduling over yarn cluster using deep reinforcement learning,

    J. Xue, T. Wang, and P. Cai, “Towards efficient workflow scheduling over yarn cluster using deep reinforcement learning,” in 2023 IEEE Global Communications Conference , 2023, pp. 473–478

  124. [131]

    Lore: a learning-based approach for workflow scheduling in clouds,

    H. Peng, C. Wu, Y . Zhan, and Y . Xia, “Lore: a learning-based approach for workflow scheduling in clouds,” in Proceedings of the conference on research in adaptive and convergent systems , 2022, pp. 47–52

  125. [132]

    SpotDAG: An RL-based algorithm for DAG workflow scheduling in heterogeneous cloud environments,

    L. Lin, L. Pan, and S. Liu, “SpotDAG: An RL-based algorithm for DAG workflow scheduling in heterogeneous cloud environments,” IEEE Transactions on Services Computing , 2024

  126. [133]

    An Effective DDPG-generated Task Scheduling Policy to Minimize Latency in Distributed Computing System,

    X. Yang and B. Hu, “An Effective DDPG-generated Task Scheduling Policy to Minimize Latency in Distributed Computing System,” in 2023 International Conference on Frontiers of Robotics and Software Engineering, 2023, pp. 303–310

  127. [134]

    Deep reinforcement learning for energy and time optimized scheduling of precedence- constrained tasks in edge–cloud computing environments,

    A. Jayanetti, S. Halgamuge, and R. Buyya, “Deep reinforcement learning for energy and time optimized scheduling of precedence- constrained tasks in edge–cloud computing environments,” Future Generation Computer Systems , vol. 137, pp. 14–30, 2022

  128. [135]

    Compo- nentized Task Scheduling in Cloud-Edge Cooperative Scenarios Based on GNN-enhanced DRL,

    J. Li, F. Zhou, W. Li, M. Zhao, X. Yan, Y . Xi, and J. Wu, “Compo- nentized Task Scheduling in Cloud-Edge Cooperative Scenarios Based on GNN-enhanced DRL,” in NOMS 2023-2023 IEEE/IFIP Network Operations and Management Symposium , 2023, pp. 1–4

  129. [136]

    Reinforcement Learning-driven Data-intensive Workflow Scheduling for V olunteer Edge-Cloud,

    M. Mounesan, M. Lemus, H. Yeddulapalli, P. Calyam, and S. Debroy, “Reinforcement Learning-driven Data-intensive Workflow Scheduling for V olunteer Edge-Cloud,” in2024 IEEE 8th International Conference on Fog and Edge Computing , 2024, pp. 79–88

  130. [137]

    Learning to Optimize Workflow Scheduling for an Edge–Cloud Computing Environment,

    K. Zhu, Z. Zhang, S. Zeadally, and F. Sun, “Learning to Optimize Workflow Scheduling for an Edge–Cloud Computing Environment,” IEEE Transactions on Cloud Computing , 2024

  131. [138]

    Deep reinforcement learning-based scheduling for optimizing system load and response time in edge and fog computing environments,

    Z. Wang, M. Goudarzi, M. Gong, and R. Buyya, “Deep reinforcement learning-based scheduling for optimizing system load and response time in edge and fog computing environments,” Future Generation Computer Systems, vol. 152, pp. 55–69, 2024

  132. [139]

    Time-Sensitive and Resource-Aware Concurrent Workflow Scheduling for Edge Computing Platforms Based on Deep Reinforcement Learning,

    J. Zhang, T. Wang, and L. Cheng, “Time-Sensitive and Resource-Aware Concurrent Workflow Scheduling for Edge Computing Platforms Based on Deep Reinforcement Learning,” Applied Sciences, vol. 13, no. 19, p. 10689, 2023

  133. [140]

    Workflow makespan mini- mization for partially connected edge network: A deep reinforcement learning-based approach,

    K. Zhu, Z. Zhang, F. Sun, and B. Shen, “Workflow makespan mini- mization for partially connected edge network: A deep reinforcement learning-based approach,” IEEE Open Journal of the Communications Society, vol. 3, pp. 518–529, 2022

  134. [141]

    A cloud resource management framework for multiple online scientific workflows using cooperative reinforcement learning agents,

    A. Asghari, M. K. Sohrabi, and F. Yaghmaee, “A cloud resource management framework for multiple online scientific workflows using cooperative reinforcement learning agents,” Computer Networks , vol. 179, p. 107340, 2020

  135. [142]

    Multi-Objective Workflow Scheduling With Deep-Q-Network-Based Multi-Agent Reinforcement Learning

    Y . LI, P. CHEN, K. GUO, and H. XIE, “Multi-Objective Workflow Scheduling With Deep-Q-Network-Based Multi-Agent Reinforcement Learning.”

  136. [143]

    Multi-Agent Deep Rein- forcement Learning Framework for Renewable Energy-Aware Work- flow Scheduling on Distributed Cloud Data Centers,

    A. Jayanetti, S. Halgamuge, and R. Buyya, “Multi-Agent Deep Rein- forcement Learning Framework for Renewable Energy-Aware Work- flow Scheduling on Distributed Cloud Data Centers,” IEEE Transac- tions on Parallel and Distributed Systems , 2024

  137. [144]

    A Deep Reinforcement Learning Approach for Cost Optimized Workflow Scheduling in Cloud Computing Environments,

    ——, “A Deep Reinforcement Learning Approach for Cost Optimized Workflow Scheduling in Cloud Computing Environments,” in Proceed- ings of the 2024 Asia Pacific Conference on Computing Technologies, Communications and Networking , 2024, pp. 74–82

  138. [145]

    Telemetry-aided cooperative multi-agent online reinforcement learning for DAG task scheduling in computing power networks,

    Y . Duan, J. Li, H. Sun, F. Zhou, J. Chen, T. Wu, W. Li, and Y . Fan, “Telemetry-aided cooperative multi-agent online reinforcement learning for DAG task scheduling in computing power networks,” Simulation Modelling Practice and Theory , vol. 132, p. 102885, 2024

  139. [146]

    Digital Twin Assisted DAG Task Scheduling Via Evolutionary Selection MARL in Large-Scale Mobile Edge Network,

    J. Huang, F. Zhou, L. Feng, W. Li, M. Zhao, X. Yan, Y . Xi, and J. Wu, “Digital Twin Assisted DAG Task Scheduling Via Evolutionary Selection MARL in Large-Scale Mobile Edge Network,” in 2023 IEEE International Conference on Communications Workshops (ICC Workshops), 2023, pp. 158–163

  140. [147]

    Multi-task scheduling in vehicular edge computing: a multi-agent reinforcement learning approach,

    Y . Zhao, L. Mo, and J. Liu, “Multi-task scheduling in vehicular edge computing: a multi-agent reinforcement learning approach,” CCF Transactions on Pervasive Computing and Interaction, pp. 1–17, 2024

  141. [148]

    Adaptive cloud bundle provisioning and multi-workflow scheduling via coalition reinforcement learning,

    X. Wang, J. Cao, and R. Buyya, “Adaptive cloud bundle provisioning and multi-workflow scheduling via coalition reinforcement learning,” IEEE Transactions on Computers, vol. 72, no. 4, pp. 1041–1054, 2022

  142. [149]

    Transformer- Enhanced DQN Approach for Energy and Cost-Efficient Large-Scale Dynamic Workflow Scheduling in Heterogeneous Environment,

    F. Ding, Y . Yuan, L. Lv, R. Zhang, and W. Zhou, “Transformer- Enhanced DQN Approach for Energy and Cost-Efficient Large-Scale Dynamic Workflow Scheduling in Heterogeneous Environment,” IEEE Internet of Things Journal , 2024

  143. [150]

    GA-DRL: Graph Neural Network-Augmented Deep Reinforcement Learning for DAG Task Scheduling over Dynamic Vehicular Clouds,

    Z. Liu, L. Huang, Z. Gao, M. Luo, S. Hosseinalipour, and H. Dai, “GA-DRL: Graph Neural Network-Augmented Deep Reinforcement Learning for DAG Task Scheduling over Dynamic Vehicular Clouds,” IEEE Transactions on Network and Service Management , 2024

  144. [151]

    A fault-tolerant workflow scheduling method on deep reinforcement learning-based in edge environment,

    T. Long, Y . Xia, Y . Ma, Q. Peng, and J. Zhao, “A fault-tolerant workflow scheduling method on deep reinforcement learning-based in edge environment,” in 2022 IEEE International Conference on Networking, Sensing and Control , 2022, pp. 1–6

  145. [152]

    Quantum ML- Based Cooperative Task Orchestration in Dew-Assisted IoT Frame- work,

    A. Mahapatra, R. Pradhan, S. K. Majhi, and K. Mishra, “Quantum ML- Based Cooperative Task Orchestration in Dew-Assisted IoT Frame- work,” Arabian Journal for Science and Engineering , pp. 1–28, 2024

  146. [153]

    TF-DDRL: A Transformer- enhanced Distributed DRL Technique for Scheduling IoT Applica- tions in Edge and Cloud Computing Environments,

    Z. Wang, M. Goudarzi, and R. Buyya, “TF-DDRL: A Transformer- enhanced Distributed DRL Technique for Scheduling IoT Applica- tions in Edge and Cloud Computing Environments,” arXiv preprint arXiv:2410.14348, 2024. 28

  147. [154]

    Pegasus, a work- flow management system for science automation,

    E. Deelman, K. Vahi, G. Juve, M. Rynge, S. Callaghan, P. J. Maechling, R. Mayani, W. Chen, R. F. Da Silva, M. Livny et al., “Pegasus, a work- flow management system for science automation,” Future Generation Computer Systems, vol. 46, pp. 17–35, 2015

  148. [155]

    A deep reinforcement learning based resource autonomic provisioning approach for cloud services,

    Q. Zong, X. Zheng, Y . Wei, and H. Sun, “A deep reinforcement learning based resource autonomic provisioning approach for cloud services,” in Collaborative Computing: Networking, Applications and Worksharing . Cham: Springer International Publishing, 2021, pp. 132–153

  149. [156]

    CILP: Co-Simulation-Based Imitation Learner for Dynamic Resource Provisioning in Cloud Com- puting Environments,

    S. Tuli, G. Casale, and N. R. Jennings, “CILP: Co-Simulation-Based Imitation Learner for Dynamic Resource Provisioning in Cloud Com- puting Environments,” IEEE Transactions on Network and Service Management, vol. 20, no. 4, pp. 4448–4460, 2023

  150. [157]

    Deep reinforcement learning for provisioning virtualized network function in inter-datacenter elastic optical networks,

    M. Zhu, Q. Chen, J. Gu, and P. Gu, “Deep reinforcement learning for provisioning virtualized network function in inter-datacenter elastic optical networks,” IEEE Transactions on Network and Service Man- agement, vol. 19, no. 3, pp. 3341–3351, 2022

  151. [158]

    AI-based resource provisioning of IoE services in 6G: A deep reinforcement learning approach,

    H. Sami, H. Otrok, J. Bentahar, and A. Mourad, “AI-based resource provisioning of IoE services in 6G: A deep reinforcement learning approach,” IEEE Transactions on Network and Service Management , vol. 18, no. 3, pp. 3527–3540, 2021

  152. [159]

    Drl-based deadline-driven advance reservation allocation in eons for cloud–edge computing,

    R. Zhu, G. Li, P. Wang, M. Xu, and S. Yu, “Drl-based deadline-driven advance reservation allocation in eons for cloud–edge computing,” IEEE Internet of Things Journal , vol. 9, no. 21, pp. 21 444–21 457, 2022

  153. [160]

    A self- learning approach for proactive resource and service provisioning in fog environment,

    M. Faraji-Mehmandar, S. Jabbehdari, and H. H. S. Javadi, “A self- learning approach for proactive resource and service provisioning in fog environment,” The Journal of Supercomputing, vol. 78, no. 15, pp. 16 997–17 026, 2022

  154. [161]

    Resource Pro- visioning in Fog Computing through Deep Reinforcement Learning,

    J. Santos, T. Wauters, B. V olckaert, and F. D. Turck, “Resource Pro- visioning in Fog Computing through Deep Reinforcement Learning,” 2021

  155. [162]

    Adaptive and efficient resource allocation in cloud datacenters using actor-critic deep reinforcement learning,

    Z. Chen, J. Hu, G. Min, C. Luo, and T. El-Ghazawi, “Adaptive and efficient resource allocation in cloud datacenters using actor-critic deep reinforcement learning,” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 8, pp. 1911–1923, 2021

  156. [163]

    Automated cloud resources provisioning with the use of the proximal policy optimization,

    W. Funika, P. Koperek, and J. Kitowski, “Automated cloud resources provisioning with the use of the proximal policy optimization,” The Journal of Supercomputing , vol. 79, no. 6, pp. 6674–6704, 2023

  157. [164]

    You calculate and I provision: A DRL-assisted service framework to realize distributed and tenant-driven virtual network slicing,

    X. Zhang, B. Li, J. Peng, X. Pan, and Z. Zhu, “You calculate and I provision: A DRL-assisted service framework to realize distributed and tenant-driven virtual network slicing,” Journal of Lightwave Technol- ogy, vol. 39, no. 1, pp. 4–16, 2021

  158. [165]

    Edge-AI: IoT Request Service Provisioning in Federated Edge Computing Using Actor-Critic Reinforcement Learning,

    H. Baghban, A. Rezapour, C.-H. Hsu, S. Nuannimnoi, and C.-Y . Huang, “Edge-AI: IoT Request Service Provisioning in Federated Edge Computing Using Actor-Critic Reinforcement Learning,” IEEE Transactions on Engineering Management, vol. 71, pp. 12 519–12 528, 2024

  159. [166]

    Trusted cloud-edge network resource management: DRL-driven service function chain orchestration for IoT,

    S. Guo, Y . Dai, S. Xu, X. Qiu, and F. Qi, “Trusted cloud-edge network resource management: DRL-driven service function chain orchestration for IoT,” IEEE Internet of Things Journal, vol. 7, no. 7, pp. 6010–6022, 2019

  160. [167]

    Intelligent Resource Provisioning and Optimization in Fog Computing using Deep Reinforcement Learning,

    A. S, A. Geetha, and R. K, “Intelligent Resource Provisioning and Optimization in Fog Computing using Deep Reinforcement Learning,” International Journal of Electronics and Communication Engineering , vol. 10, no. 8, pp. 85–97, 2023

  161. [168]

    Deep reinforcement learning for computation offloading in mobile edge computing environment,

    M. Chen, T. Wang, S. Zhang, and A. Liu, “Deep reinforcement learning for computation offloading in mobile edge computing environment,” Computer Communications, vol. 175, pp. 1–12, 2021

  162. [169]

    Dynamic reservation of edge servers via deep reinforcement learning for connected vehicles,

    J. Zhang, S. Chen, X. Wang, and Y . Zhu, “Dynamic reservation of edge servers via deep reinforcement learning for connected vehicles,” IEEE Transactions on Mobile Computing , vol. 22, no. 5, pp. 2661–2674, 2021

  163. [170]

    Dynamic provisioning of resources based on load balancing and service broker policy in cloud computing,

    A. Jyoti and M. Shrimali, “Dynamic provisioning of resources based on load balancing and service broker policy in cloud computing,” Cluster Computing, vol. 23, no. 1, pp. 377–395, 2020

  164. [171]

    Combined use of coral reefs opti- mization and multi-agent deep Q-network for energy-aware resource provisioning in cloud data centers using DVFS technique,

    A. Asghari and M. K. Sohrabi, “Combined use of coral reefs opti- mization and multi-agent deep Q-network for energy-aware resource provisioning in cloud data centers using DVFS technique,” Cluster Computing, vol. 25, no. 1, pp. 119–140, 2022

  165. [172]

    Distributed Multi- Cloud Multi-Access Edge Computing by Multi-Agent Reinforcement Learning,

    Y . Zhang, B. Di, Z. Zheng, J. Lin, and L. Song, “Distributed Multi- Cloud Multi-Access Edge Computing by Multi-Agent Reinforcement Learning,” IEEE Transactions on Wireless Communications , vol. 20, no. 4, pp. 2565–2578, 2021

  166. [173]

    Adaptive resource management for edge network slicing using incremental multi-agent deep reinforcement learning,

    H. Li, Y . Liu, X. Zhou, X. Vasilakos, R. Nejabati, S. Yan, and D. Simeonidou, “Adaptive resource management for edge network slicing using incremental multi-agent deep reinforcement learning,” arXiv preprint arXiv:2310.17523 , 2023

  167. [174]

    Dynamic and efficient resource allocation for 5G end-to-end network slicing: A multi-agent deep rein- forcement learning approach,

    M. Asim Ejaz, G. Wu, and T. Iqbal, “Dynamic and efficient resource allocation for 5G end-to-end network slicing: A multi-agent deep rein- forcement learning approach,” International Journal of Communication Systems, p. e5916

  168. [175]

    Task scheduling, resource provisioning, and load balancing on scientific workflows using parallel SARSA reinforcement learning agents and genetic algorithm,

    A. Asghari, M. K. Sohrabi, and F. Yaghmaee, “Task scheduling, resource provisioning, and load balancing on scientific workflows using parallel SARSA reinforcement learning agents and genetic algorithm,” The Journal of Supercomputing , vol. 77, no. 3, pp. 2800–2828, 2021

  169. [176]

    Generative AI-enabled Quantum Computing Networks and Intelligent Resource Allocation,

    M. Xu, D. Niyato, J. Kang, Z. Xiong, Y . Cao, Y . Gao, C. Ren, and H. Yu, “Generative AI-enabled Quantum Computing Networks and Intelligent Resource Allocation,” 2024

  170. [177]

    Joint computation offloading and resource provisioning for edge-cloud computing environment: A machine learning-based approach,

    A. Shahidinejad and M. Ghobaei-Arani, “Joint computation offloading and resource provisioning for edge-cloud computing environment: A machine learning-based approach,” Software: Practice and Experience, vol. 50, no. 12, pp. 2212–2230, 2020

  171. [178]

    Towards Intelligent Provisioning of Virtualized Network Functions in Cloud of Things: A Deep Reinforcement Learning Based Approach,

    B. He, J. Wang, Q. Qi, H. Sun, and J. Liao, “Towards Intelligent Provisioning of Virtualized Network Functions in Cloud of Things: A Deep Reinforcement Learning Based Approach,” IEEE Transactions on Cloud Computing , vol. 10, no. 2, pp. 1262–1274, 2022

  172. [179]

    DRL-Based Green Resource Provisioning for 5G and Beyond Networks,

    M. Dieye, W. Jaafar, H. Elbiaze, and R. H. Glitho, “DRL-Based Green Resource Provisioning for 5G and Beyond Networks,” IEEE Transactions on Green Communications and Networking, vol. 7, no. 4, pp. 2163–2180, 2023

  173. [180]

    iRAF: A deep reinforcement learning approach for collaborative mobile edge computing IoT networks,

    J. Chen, S. Chen, Q. Wang, B. Cao, G. Feng, and J. Hu, “iRAF: A deep reinforcement learning approach for collaborative mobile edge computing IoT networks,” IEEE Internet of Things Journal , vol. 6, no. 4, pp. 7011–7024, 2019

  174. [181]

    ADRL: A Hybrid Anomaly-Aware Deep Reinforcement Learning-Based Re- source Scaling in Clouds,

    S. Kardani-Moghaddam, R. Buyya, and K. Ramamohanarao, “ADRL: A Hybrid Anomaly-Aware Deep Reinforcement Learning-Based Re- source Scaling in Clouds,” IEEE Transactions on Parallel and Dis- tributed Systems, vol. 32, no. 3, pp. 514–526, 2021

  175. [182]

    Resource Allocation With Workload-Time Windows for Cloud-Based Software Services: A Deep Reinforcement Learning Approach,

    X. Chen, L. Yang, Z. Chen, G. Min, X. Zheng, and C. Rong, “Resource Allocation With Workload-Time Windows for Cloud-Based Software Services: A Deep Reinforcement Learning Approach,” IEEE Transactions on Cloud Computing, vol. 11, no. 2, pp. 1871–1885, 2023

  176. [183]

    DERP: A Deep Rein- forcement Learning Cloud System for Elastic Resource Provisioning,

    C. Bitsakos, I. Konstantinou, and N. Koziris, “DERP: A Deep Rein- forcement Learning Cloud System for Elastic Resource Provisioning,” in 2018 IEEE International Conference on Cloud Computing Technol- ogy and Science , 2018, pp. 21–29

  177. [184]

    Deep Reinforcement Learning- Based Resource Allocation for Content Distribution in IoT-Edge-Cloud Computing Environments,

    T. Cui, R. Yang, C. Fang, and S. Yu, “Deep Reinforcement Learning- Based Resource Allocation for Content Distribution in IoT-Edge-Cloud Computing Environments,” Symmetry, vol. 15, no. 1, p. 217, 2023

  178. [185]

    A Computing Offloading Resource Allocation Scheme Using Deep Reinforcement Learning in Mobile Edge Computing Systems,

    X. Li, “A Computing Offloading Resource Allocation Scheme Using Deep Reinforcement Learning in Mobile Edge Computing Systems,” Journal of Grid Computing , vol. 19, no. 3, p. 35, 2021

  179. [186]

    Deep-Reinforcement-Learning-Based Resource Allocation for Content Distribution in Fog Radio Access Networks,

    C. Fang, H. Xu, Y . Yang, Z. Hu, S. Tu, K. Ota, Z. Yang, M. Dong, Z. Han, F. R. Yu, and Y . Liu, “Deep-Reinforcement-Learning-Based Resource Allocation for Content Distribution in Fog Radio Access Networks,” IEEE Internet of Things Journal, vol. 9, no. 18, pp. 16 874– 16 883, 2022

  180. [187]

    Security computing resource allocation based on deep reinforcement learning in serverless multi- cloud edge computing,

    H. Zhang, J. Wang, H. Zhang, and C. Bu, “Security computing resource allocation based on deep reinforcement learning in serverless multi- cloud edge computing,” Future Generation Computer Systems , vol. 151, pp. 152–161, 2024

  181. [188]

    Optimizing task offloading and resource allocation in edge-cloud networks: A DRL approach,

    I. Ullah, H.-K. Lim, Y .-J. Seok, and Y .-H. Han, “Optimizing task offloading and resource allocation in edge-cloud networks: A DRL approach,” Journal of Cloud Computing , vol. 12, no. 1, p. 112, 2023

  182. [189]

    Deep Rein- forcement Learning Based Approach for Online Service Placement and Computation Resource Allocation in Edge Computing,

    T. Liu, S. Ni, X. Li, Y . Zhu, L. Kong, and Y . Yang, “Deep Rein- forcement Learning Based Approach for Online Service Placement and Computation Resource Allocation in Edge Computing,” IEEE Transactions on Mobile Computing , vol. 22, no. 7, pp. 3870–3881, 2023

  183. [190]

    Resource Man- agement with Deep Reinforcement Learning,

    H. Mao, M. Alizadeh, I. Menache, and S. Kandula, “Resource Man- agement with Deep Reinforcement Learning,” in Proceedings of the 15th ACM Workshop on Hot Topics in Networks , 2016, pp. 50–56

  184. [191]

    Adaptive Resource Allocation in Cloud Data Centers using Actor-Critical Deep Reinforcement Learning for Optimized Load Balancing,

    M. Arvindhan and D. R. Kumar, “Adaptive Resource Allocation in Cloud Data Centers using Actor-Critical Deep Reinforcement Learning for Optimized Load Balancing,” International Journal on Recent and Innovation Trends in Computing and Communication , vol. 11, no. 5s, pp. 310–318, 2023

  185. [192]

    Learning-Based Resource Allocation in Cloud Data Center using Advantage Actor-Critic,

    Z. Chen, J. Hu, and G. Min, “Learning-Based Resource Allocation in Cloud Data Center using Advantage Actor-Critic,” in 2019 IEEE International Conference on Communications (ICC) , 2019, pp. 1–6

  186. [193]

    Blockchain-Based Edge Computing Resource Allocation in IoT: A Deep Reinforcement Learning Approach,

    Y . He, Y . Wang, C. Qiu, Q. Lin, J. Li, and Z. Ming, “Blockchain-Based Edge Computing Resource Allocation in IoT: A Deep Reinforcement Learning Approach,” IEEE Internet of Things Journal , vol. 8, no. 4, pp. 2226–2237, 2021. 29

  187. [194]

    A Deep Reinforcement Learning-Based Resource Management Game in Vehicular Edge Computing,

    X. Zhu, Y . Luo, A. Liu, N. N. Xiong, M. Dong, and S. Zhang, “A Deep Reinforcement Learning-Based Resource Management Game in Vehicular Edge Computing,” IEEE Transactions on Intelligent Trans- portation Systems, vol. 23, no. 3, pp. 2422–2433, 2022

  188. [195]

    Joint Computation Offloading and Resource Allocation for Edge-Cloud Collaboration in Internet of Vehicles via Deep Reinforcement Learning,

    J. Huang, J. Wan, B. Lv, Q. Ye, and Y . Chen, “Joint Computation Offloading and Resource Allocation for Edge-Cloud Collaboration in Internet of Vehicles via Deep Reinforcement Learning,” IEEE Systems Journal, vol. 17, no. 2, pp. 2500–2511, 2023

  189. [196]

    Intelligent multi-agent rein- forcement learning model for resources allocation in cloud computing,

    A. Belgacem, S. Mahmoudi, and M. Kihl, “Intelligent multi-agent rein- forcement learning model for resources allocation in cloud computing,” Journal of King Saud University - Computer and Information Sciences , vol. 34, no. 6, pp. 2391–2404, 2022

  190. [197]

    Multi agent deep reinforcement learn- ing for resource allocation in container-based clouds environments,

    S. Nagarajan, P. S. Rani, M. S. Vinmathi, V . Subba Reddy, A. L. M. Saleth, and D. Abdus Subhahan, “Multi agent deep reinforcement learn- ing for resource allocation in container-based clouds environments,” Expert Systems, p. exsy.13362, 2023

  191. [198]

    Multi-agent reinforcement learning for resource allocation in IoT networks with edge computing,

    X. Liu, J. Yu, Z. Feng, and Y . Gao, “Multi-agent reinforcement learning for resource allocation in IoT networks with edge computing,” China Communications, vol. 17, no. 9, pp. 220–236, 2020

  192. [199]

    Deep Reinforcement Learning Multi-Agent System for Resource Allocation in Industrial Internet of Things,

    J. Rosenberger, M. Urlaub, F. Rauterberg, T. Lutz, A. Selig, M. B ¨uhren, and D. Schramm, “Deep Reinforcement Learning Multi-Agent System for Resource Allocation in Industrial Internet of Things,” Sensors, vol. 22, no. 11, p. 4099, 2022

  193. [200]

    Cloud Resource Schedul- ing With Deep Reinforcement Learning and Imitation Learning,

    W. Guo, W. Tian, Y . Ye, L. Xu, and K. Wu, “Cloud Resource Schedul- ing With Deep Reinforcement Learning and Imitation Learning,” IEEE Internet of Things Journal , vol. 8, no. 5, pp. 3576–3586, 2021

  194. [201]

    Intelligent Cloud Resource Manage- ment with Deep Reinforcement Learning,

    Y . Zhang, J. Yao, and H. Guan, “Intelligent Cloud Resource Manage- ment with Deep Reinforcement Learning,” IEEE Cloud Computing , vol. 4, no. 6, pp. 60–69, 2017

  195. [202]

    Multi- Dimensional Resource Allocation in Distributed Data Centers Using Deep Reinforcement Learning,

    W. Wei, H. Gu, K. Wang, J. Li, X. Zhang, and N. Wang, “Multi- Dimensional Resource Allocation in Distributed Data Centers Using Deep Reinforcement Learning,” IEEE Transactions on Network and Service Management, vol. 20, no. 2, pp. 1817–1829, 2023

  196. [203]

    Stacked Autoencoder-Based Deep Reinforcement Learning for Online Resource Scheduling in Large-Scale MEC Networks,

    F. Jiang, K. Wang, L. Dong, C. Pan, and K. Yang, “Stacked Autoencoder-Based Deep Reinforcement Learning for Online Resource Scheduling in Large-Scale MEC Networks,” IEEE Internet of Things Journal, vol. 7, no. 10, pp. 9278–9290, 2020

  197. [204]

    A deep reinforcement learning based hybrid algorithm for efficient resource scheduling in edge computing environment,

    F. Xue, Q. Hai, T. Dong, Z. Cui, and Y . Gong, “A deep reinforcement learning based hybrid algorithm for efficient resource scheduling in edge computing environment,”Information Sciences, vol. 608, pp. 362– 374, 2022

  198. [205]

    Energy-Efficient Optimization for Mobile Edge Computing With Quantum Machine Learning,

    J. Adu Ansere, D. T. Tran, O. A. Dobre, H. Shin, G. K. Karagian- nidis, and T. Q. Duong, “Energy-Efficient Optimization for Mobile Edge Computing With Quantum Machine Learning,” IEEE Wireless Communications Letters, vol. 13, no. 3, pp. 661–665, 2024

  199. [206]

    Layerwise Quantum Deep Reinforcement Learning for Joint Optimization of UA V Trajectory and Resource Allocation,

    Silvirianti, B. Narottama, and S. Y . Shin, “Layerwise Quantum Deep Reinforcement Learning for Joint Optimization of UA V Trajectory and Resource Allocation,” IEEE Internet of Things Journal , vol. 11, no. 1, pp. 430–443, 2024

  200. [208]

    Efficient, economical and energy-saving multi-workflow scheduling in hybrid cloud,

    Z. Sun, H. Huang, Z. Li, C. Gu, R. Xie, and B. Qian, “Efficient, economical and energy-saving multi-workflow scheduling in hybrid cloud,” Expert Systems with Applications , vol. 228, p. 120401, 2023

  201. [209]

    Et2fa: A hy- brid heuristic algorithm for deadline-constrained workflow scheduling in cloud,

    Z. Sun, B. Zhang, C. Gu, R. Xie, B. Qian, and H. Huang, “Et2fa: A hy- brid heuristic algorithm for deadline-constrained workflow scheduling in cloud,” IEEE Transactions on Services Computing , vol. 16, no. 3, pp. 1807–1821, 2022

  202. [210]

    Scheduling workflows with privacy protection constraints for big data applications on cloud,

    Y . Wen, J. Liu, W. Dou, X. Xu, B. Cao, and J. Chen, “Scheduling workflows with privacy protection constraints for big data applications on cloud,” Future Generation Computer Systems , vol. 108, pp. 1084– 1091, 2020

  203. [211]

    Symmetric encryption algorithms: Review and evaluation study,

    M. N. Alenezi, H. Alabdulrazzaq, and N. Q. Mohammad, “Symmetric encryption algorithms: Review and evaluation study,” International Journal of Communication Networks and Information Security, vol. 12, no. 2, pp. 256–272, 2020

  204. [212]

    Performance analysis of md5 and sha-256 algorithms to maintain data integrity,

    S. R. Prasanna and B. Premananda, “Performance analysis of md5 and sha-256 algorithms to maintain data integrity,” in 2021 International Conference on Recent Trends on Electronics, Information, Communi- cation & Technology, 2021, pp. 246–250

  205. [213]

    Digital twin- empowered network planning for multi-tier computing,

    C. Zhou, J. Gao, M. Li, X. S. Shen, and W. Zhuang, “Digital twin- empowered network planning for multi-tier computing,” Journal of Communications and Information Networks , vol. 7, no. 3, pp. 221– 238, 2022

  206. [214]

    Multi-tier multi- domain network slicing: A resource allocation perspective,

    S. O. Oladejo, S. O. Ekwe, and L. A. Akinyemi, “Multi-tier multi- domain network slicing: A resource allocation perspective,” in 2021 IEEE AFRICON, 2021, pp. 1–6

  207. [215]

    Two- Tier Resource Allocation for Multitenant Network Slicing: A Federated Deep Reinforcement Learning Approach,

    R. Ou, G. Sun, D. Ayepah-Mensah, G. O. Boateng, and G. Liu, “Two- Tier Resource Allocation for Multitenant Network Slicing: A Federated Deep Reinforcement Learning Approach,” IEEE Internet of Things Journal, vol. 10, no. 22, pp. 20 174–20 187, 2023

  208. [216]

    HierRL: Hierarchical reinforcement learning for task scheduling in distributed systems,

    Y . Guan, Y . Liu, Y . Li, and X. Xu, “HierRL: Hierarchical reinforcement learning for task scheduling in distributed systems,” in 2022 Interna- tional Joint Conference on Neural Networks , 2022, pp. 1–8

  209. [217]

    An adaptive multi-objective multi-task scheduling method by hierarchical deep reinforcement learning,

    J. Zhang, B. Guo, X. Ding, D. Hu, J. Tang, K. Du, C. Tang, and Y . Jiang, “An adaptive multi-objective multi-task scheduling method by hierarchical deep reinforcement learning,” Applied Soft Computing , vol. 154, p. 111342, 2024

  210. [218]

    Efficient and scalable re- inforcement learning for large-scale network control,

    C. Ma, A. Li, Y . Du, H. Dong, and Y . Yang, “Efficient and scalable re- inforcement learning for large-scale network control,” Nature Machine Intelligence, pp. 1–15, 2024

  211. [219]

    Large language models as traffic signal control agents: Capacity and opportunity,

    S. Lai, Z. Xu, W. Zhang, H. Liu, and H. Xiong, “Large language models as traffic signal control agents: Capacity and opportunity,”arXiv preprint arXiv:2312.16044, 2023. Yan Gu received the B.E. degree from Nanjing Institute of Technology, Nanjing, China, in 2020, and M.S. degr...

  212. [2014]

    He has published more than 110 papers in refereed journals and conferences

    He was an Assistant Professor at Dublin City University, and a Marie Curie Fellow at University College Dublin. He has published more than 110 papers in refereed journals and conferences. His research focuses on distributed computing and deep reinforcement learning. Prof Cheng...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.