Pith. sign in

REVIEW 2 cited by

The Curse of Diversity in Ensemble-Based Exploration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.04342 v1 pith:CKDCPHTV submitted 2024-05-07 cs.LG

classification cs.LG
keywords ensemblecursedatadiversityexplorationlearningperformancetraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We uncover a surprising phenomenon in deep reinforcement learning: training a diverse ensemble of data-sharing agents -- a well-established exploration strategy -- can significantly impair the performance of the individual ensemble members when compared to standard single-agent training. Through careful analysis, we attribute the degradation in performance to the low proportion of self-generated data in the shared training data for each ensemble member, as well as the inefficiency of the individual ensemble members to learn from such highly off-policy data. We thus name this phenomenon the curse of diversity. We find that several intuitive solutions -- such as a larger replay buffer or a smaller ensemble size -- either fail to consistently mitigate the performance loss or undermine the advantages of ensembling. Finally, we demonstrate the potential of representation learning to counteract the curse of diversity with a novel method named Cross-Ensemble Representation Learning (CERL) in both discrete and continuous control domains. Our work offers valuable insights into an unexpected pitfall in ensemble-based exploration and raises important caveats for future applications of similar approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Arbitration Control for an Ensemble of Diversified DQN variants in Continual Reinforcement Learning

    cs.LG 2025-09 conditional novelty 5.0 of 10

    ACED-DQN combines heterogeneous DQN variants with loss-based reliability weighting and experience assignment, but the paper's own ablation indicates that arbitration control is not the key factor behind the performance gain.

  2. The impact of intrinsic rewards on exploration in Reinforcement Learning

    cs.AI 2025-01 conditional novelty 5.0 of 10

    An empirical MiniGrid study shows state-counting is best for low-dimensional observations, maximum entropy is more robust with images, and DIAYN skill learning does not aid exploration.

Pith tools