Pith. sign in

REVIEW 1 cited by

A Rollout-Based Algorithm and Reward Function for Resource Allocation in Business Processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11250 v2 pith:IKDRNBJ7 submitted 2025-04-15 cs.LG cs.AI

classification cs.LGcs.AI
keywords rewardbusinessprocessesalgorithmfunctionobjectivepolicyallocation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Resource allocation plays a critical role in minimizing cycle time and improving the efficiency of business processes. Recently, Deep Reinforcement Learning (DRL) has emerged as a powerful technique to optimize resource allocation policies in business processes. In the DRL framework, an agent learns a policy through interaction with the environment, guided solely by reward signals that indicate the quality of its decisions. However, existing algorithms are not suitable for dynamic environments such as business processes. Furthermore, existing DRL-based methods rely on engineered reward functions that approximate the desired objective, but a misalignment between reward and objective can lead to undesired decisions or suboptimal policies. To address these issues, we propose a rollout-based DRL algorithm and a reward function to optimize the objective directly. Our algorithm iteratively improves the policy by evaluating execution trajectories following different actions. Our reward function directly decomposes the objective function of minimizing the cycle time, such that trial-and-error reward engineering becomes unnecessary. We evaluated our method in six scenarios, for which the optimal policy can be computed, and on a set of increasingly complex, realistically sized process models. The results show that our algorithm can learn the optimal policy for the scenarios and outperform or match the best heuristics on the realistically sized business processes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GymPN: A Library for Decision-Making in Process Management Systems

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A software library, GymPN, extends the A-E Petri net framework with partial observability and multiple action transitions, and learns optimal task assignment policies on eight workflow patterns.

Pith tools