Pith. sign in

REVIEW 2 cited by

Dyadic Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.07843 v6 pith:3Y6CWWXH submitted 2023-08-15 cs.LG stat.APstat.ML

classification cs.LGstat.APstat.ML
keywords dyadichealthcareinterventionsmobiletargetdevelopindividuals
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mobile health aims to enhance health outcomes by delivering interventions to individuals as they go about their daily life. The involvement of care partners and social support networks often proves crucial in helping individuals managing burdensome medical conditions. This presents opportunities in mobile health to design interventions that target the dyadic relationship -- the relationship between a target person and their care partner -- with the aim of enhancing social support. In this paper, we develop dyadic RL, an online reinforcement learning algorithm designed to personalize intervention delivery based on contextual factors and past responses of a target person and their care partner. Here, multiple sets of interventions impact the dyad across multiple time intervals. The developed dyadic RL is Bayesian and hierarchical. We formally introduce the problem setup, develop dyadic RL and establish a regret bound. We demonstrate dyadic RL's empirical performance through simulation studies on both toy scenarios and on a realistic test bed constructed from data collected in a mobile health study.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations

    stat.ML 2026-07 conditional novelty 6.0 of 10

    Synchronized two-timescale hierarchical PPO-style learning converges in average optimality gap at O(T^{-1/2}) (faster under market sharpness) and raises simulated used-car profits under joint shocks.

  2. Reinforcement Learning on Dyads to Enhance Medication Adherence

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A three-agent reinforcement learning framework with domain-informed surrogate rewards outperforms single-agent and random policies for personalizing dyadic medication-adherence interventions in simulation.

Pith tools