Pith. sign in

REVIEW 1 cited by

Federated Temporal Difference Learning with Linear Function Approximation under Environmental Heterogeneity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.02212 v2 pith:EGFGAG5W submitted 2023-02-04 cs.LG math.OC

classification cs.LGmath.OC
keywords agentsfederatedheterogeneityfunctionlearninglinearalgorithmanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We initiate the study of federated reinforcement learning under environmental heterogeneity by considering a policy evaluation problem. Our setup involves $N$ agents interacting with environments that share the same state and action space but differ in their reward functions and state transition kernels. Assuming agents can communicate via a central server, we ask: Does exchanging information expedite the process of evaluating a common policy? To answer this question, we provide the first comprehensive finite-time analysis of a federated temporal difference (TD) learning algorithm with linear function approximation, while accounting for Markovian sampling, heterogeneity in the agents' environments, and multiple local updates to save communication. Our analysis crucially relies on several novel ingredients: (i) deriving perturbation bounds on TD fixed points as a function of the heterogeneity in the agents' underlying Markov decision processes (MDPs); (ii) introducing a virtual MDP to closely approximate the dynamics of the federated TD algorithm; and (iii) using the virtual MDP to make explicit connections to federated optimization. Putting these pieces together, we rigorously prove that in a low-heterogeneity regime, exchanging model estimates leads to linear convergence speedups in the number of agents.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Variance-Reduced Q-Learning over Static and Time-Varying Networks

    cs.LG 2026-07 conditional novelty 7.0 of 10

    VRDQ achieves the optimal collaborative error rate 1/√(NT) for decentralized tabular Q-learning while requiring only O(log²(NT)) communication per agent, on both static and time-varying networks.

Pith tools