Pith. sign in

REVIEW 2 cited by

Achieving Fairness in Multi-Agent Markov Decision Processes Using Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00324 v1 pith:T6FYN75S submitted 2023-06-01 cs.LG cs.MA

classification cs.LGcs.MA
keywords fairnessapproachmulti-agentboundproposeagentsconfidencedecision
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Fairness plays a crucial role in various multi-agent systems (e.g., communication networks, financial markets, etc.). Many multi-agent dynamical interactions can be cast as Markov Decision Processes (MDPs). While existing research has focused on studying fairness in known environments, the exploration of fairness in such systems for unknown environments remains open. In this paper, we propose a Reinforcement Learning (RL) approach to achieve fairness in multi-agent finite-horizon episodic MDPs. Instead of maximizing the sum of individual agents' value functions, we introduce a fairness function that ensures equitable rewards across agents. Since the classical Bellman's equation does not hold when the sum of individual value functions is not maximized, we cannot use traditional approaches. Instead, in order to explore, we maintain a confidence bound of the unknown environment and then propose an online convex optimization based approach to obtain a policy constrained to this confidence region. We show that such an approach achieves sub-linear regret in terms of the number of episodes. Additionally, we provide a probably approximately correct (PAC) guarantee based on the obtained regret bound. We also propose an offline RL algorithm and bound the optimality gap with respect to the optimal fair solution. To mitigate computational complexity, we introduce a policy-gradient type method for the fair objective. Simulation experiments also demonstrate the efficacy of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fairness in Reinforcement Learning with Bisimulation Metrics

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Bisimulator uses bisimulation metrics to modify rewards and observations, letting unconstrained RL policies approximately satisfy demographic parity in lending and college admissions benchmarks.

  2. Fair Contracts in Principal-Agent Games with Heterogeneous Types

    cs.GT 2025-06 conditional novelty 5.0 of 10

    In a two-agent coin game, a principal trained to minimize variance in wealth learns linear contracts that equalize wealth across heterogeneous agents without reducing total welfare.

Pith tools