Pith. sign in

REVIEW 1 cited by

Multi-Agent Learning in Contextual Games under Unknown Constraints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14685 v2 pith:D5VL5PJY submitted 2023-10-23 cs.GT cs.SYeess.SY

classification cs.GTcs.SYeess.SY
keywords unknownlearningconstraintconstraintscontextualmulti-agentno-violationadanormalgp
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We consider the problem of learning to play a repeated contextual game with unknown reward and unknown constraints functions. Such games arise in applications where each agent's action needs to belong to a feasible set, but the feasible set is a priori unknown. For example, in constrained multi-agent reinforcement learning, the constraints on the agents' policies are a function of the unknown dynamics and hence, are themselves unknown. Under kernel-based regularity assumptions on the unknown functions, we develop a no-regret, no-violation approach which exploits similarities among different reward and constraint outcomes. The no-violation property ensures that the time-averaged sum of constraint violations converges to zero as the game is repeated. We show that our algorithm, referred to as c.z.AdaNormalGP, obtains kernel-dependent regret bounds and that the cumulative constraint violations have sublinear kernel-dependent upper bounds. In addition we introduce the notion of constrained contextual coarse correlated equilibria (c.z.CCE) and show that $\epsilon$-c.z.CCEs can be approached whenever players' follow a no-regret no-violation strategy. Finally, we experimentally demonstrate the effectiveness of c.z.AdaNormalGP on an instance of multi-agent reinforcement learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prediction-Aware Learning in Multi-Agent Systems

    cs.GT 2025-01 accept novelty 6.0 of 10

    A contextual optimistic multiplicative weights algorithm (POMWU) achieves static-game regret, equilibrium convergence, and social welfare guarantees in time-varying games when players can predict the changing state of...

Pith tools