Pith. sign in

REVIEW 2 cited by

Offline Reinforcement Learning for Autonomous Driving with Safety and Exploration Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07067 v2 pith:66DH5WQB submitted 2021-10-13 cs.RO cs.SYeess.SY

classification cs.ROcs.SYeess.SY
keywords offlineautonomousdrivingpoliciesalgorithmscontrolconventionalenhancement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) is a powerful data-driven control method that has been largely explored in autonomous driving tasks. However, conventional RL approaches learn control policies through trial-and-error interactions with the environment and therefore may cause disastrous consequences such as collisions when testing in real-world traffic. Offline RL has recently emerged as a promising framework to learn effective policies from previously-collected, static datasets without the requirement of active interactions, making it especially appealing for autonomous driving applications. Despite promising, existing offline RL algorithms such as Batch-Constrained deep Q-learning (BCQ) generally lead to rather conservative policies with limited exploration efficiency. To address such issues, this paper presents an enhanced BCQ algorithm by employing a learnable parameter noise scheme in the perturbation model to increase the diversity of observed actions. In addition, a Lyapunov-based safety enhancement strategy is incorporated to constrain the explorable state space within a safe region. Experimental results in highway and parking traffic scenarios show that our approach outperforms the conventional RL method, as well as state-of-the-art offline RL algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A wavelet-Fourier conditioning scheme for trajectory diffusion improves offline RL returns on most D4RL tasks by modeling low- and high-frequency components separately.

  2. RAD: Retrieval High-quality Demonstrations to Enhance Decision-making

    cs.AI 2025-07 conditional novelty 6.0 of 10

    RAD retrieves high-return states from an offline dataset and uses condition-guided diffusion to plan toward them, reporting competitive D4RL MuJoCo scores.

Pith tools