Pith. sign in

REVIEW 15 cited by

Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.09645 v1 pith:FQLV5H64 submitted 2021-07-20 cs.AI cs.LG

classification cs.AIcs.LG
keywords drq-v2controlcontinuousdirectlylearningmodel-freereinforcementtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present DrQ-v2, a model-free reinforcement learning (RL) algorithm for visual continuous control. DrQ-v2 builds on DrQ, an off-policy actor-critic approach that uses data augmentation to learn directly from pixels. We introduce several improvements that yield state-of-the-art results on the DeepMind Control Suite. Notably, DrQ-v2 is able to solve complex humanoid locomotion tasks directly from pixel observations, previously unattained by model-free RL. DrQ-v2 is conceptually simple, easy to implement, and provides significantly better computational footprint compared to prior work, with the majority of tasks taking just 8 hours to train on a single GPU. Finally, we publicly release DrQ-v2's implementation to provide RL practitioners with a strong and computationally efficient baseline.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Across noisy DeepMind Control tasks, explicit bisimulation-metric losses add little denoising benefit beyond plain self-prediction and feature normalization, which dominate performance.

  2. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  3. Error Amplification Limits ANN-to-SNN Conversion in Continuous Control

    cs.NE 2026-01 conditional novelty 6.0 of 10

    Temporally correlated action errors, amplified by closed-loop dynamics, explain ANN-to-SNN conversion failures in continuous control, and cross-step residual potential initialization mitigates them.

  4. Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A sample-efficient SAC-based method with replay-ratio resets, offline data bootstrapping, and expert-driven fine-tuning produces a goalkeeper that outperforms the built-in AI in EA SPORTS FC 25.

  5. DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Depth-guided masking improves visual RL generalization, sample efficiency, and interpretability on manipulation tasks.

  6. Touch begins where vision ends: Generalizable policies for contact-rich manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A localize-then-execute policy that combines vision-language reaching, semantic background augmentation, and residual reinforcement learning with tactile sensing reaches about 90% success on millimeter-precision manip...

  7. The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Introduces LEAST, an adaptive early-episode-stopping rule for off-policy deep RL that improves learning efficiency on MuJoCo and DeepMind Control benchmarks.

  8. Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    AmpAttention and RVAF raise multi-view robotic manipulation success and cut training time by suppressing attention noise with a differential-amplifier-style mechanism plus a CMRR loss.

  9. What Matters for Simulation to Online Reinforcement Learning on Real Robots

    cs.RO 2026-02 conditional novelty 5.0 of 10

    Sim-to-online RL on three real robots is stabilized by retaining data, warm-starting the replay buffer, and using asymmetric actor-critic updates with a low actor learning rate.

  10. ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning

    cs.RO 2025-12 conditional novelty 5.0 of 10

    ReinforceGen uses imitation learning, RL fine-tuning of skill policies, and real-time pose replanning to reach over 80% success on five long-horizon Robosuite tasks from only 10 human demonstrations.

  11. RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation

    cs.RO 2025-10 unverdicted novelty 5.0 of 10

    Abstract describes RoDyn but full text describes iMoWM; the record is internally inconsistent and the headline claims are absent from the body.

  12. AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation

    cs.RO 2025-07 conditional novelty 5.0 of 10

    AC-DiT adds mobility-to-body conditioning and perception-aware 2D/3D weighting to a diffusion transformer, improving success rates on simulated and real-world mobile manipulation tasks.

  13. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

  14. Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

    cs.RO 2025-06 conditional novelty 5.0 of 10

    FiS-VLA embeds a diffusion-based action module into the final transformer blocks of a vision-language model, achieving 69% mean success on RLBench and a claimed 117.7 Hz control frequency.

  15. A Survey of State Representation Learning for Deep Reinforcement Learning

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A six-class taxonomy of state representation learning methods for model-free online deep reinforcement learning, with selection guidelines, evaluation metrics, and future directions.

Pith tools