CODS iteratively picks high-Bellman-residual transitions into a frozen reusable subset, retaining 96.6% of eligible-pool D4RL performance at a 10% budget and beating one-shot and gradient-matching baselines.
International Conference on Machine Learning , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning
CODS iteratively picks high-Bellman-residual transitions into a frozen reusable subset, retaining 96.6% of eligible-pool D4RL performance at a 10% budget and beating one-shot and gradient-matching baselines.