Pith. sign in

REVIEW 2 cited by

Conjectural Online Learning with First-order Beliefs in Asymmetric Information Stochastic Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18781 v4 pith:KPKZJRHP submitted 2024-02-29 cs.GT cs.LGcs.SYeess.SY

classification cs.GTcs.LGcs.SYeess.SY
keywords informationlearningonlineaisgsmethodsadaptasymmetricbayesian
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Asymmetric information stochastic games (AISGs) arise in many complex socio-technical systems, such as cyber-physical systems and IT infrastructures. Existing computational methods for AISGs are primarily offline and can not adapt to equilibrium deviations. Further, current methods are limited to particular information structures to avoid belief hierarchies. Considering these limitations, we propose conjectural online learning (COL), an online learning method under generic information structures in AISGs. COL uses a forecaster-actor-critic (FAC) architecture, where subjective forecasts are used to conjecture the opponents' strategies within a lookahead horizon, and Bayesian learning is used to calibrate the conjectures. To adapt strategies to nonstationary environments based on information feedback, COL uses online rollout with cost function approximation (actor-critic). We prove that the conjectures produced by COL are asymptotically consistent with the information feedback in the sense of a relaxed Bayesian consistency. We also prove that the empirical strategy profile induced by COL converges to the Berk-Nash equilibrium, a solution concept characterizing rationality under subjectivity. Experimental results from an intrusion response use case demonstrate COL's {faster convergence} over state-of-the-art reinforcement learning methods against nonstationary attacks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Security Response to Network Intrusions in IT Systems

    cs.GT 2025-02 conditional novelty 6.0 of 10

    A combination of digital-twin emulation and simulation-based game-theoretic optimization yields near-optimal automated security response strategies for IT infrastructures, demonstrated in emulation.

  2. The Game-Theoretic Symbiosis of Trust and AI in Networked Systems

    cs.AI 2024-11 unverdicted novelty 2.0 of 10

    A survey chapter that combines trust scores, Bayesian updates, and game-theoretic models to argue that AI and trust should be managed as a strategic symbiosis for cybersecurity.

Pith tools