REVIEW 4 cited by
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Evaluating autonomous vehicle stacks (AVs) in simulation typically involves replaying driving logs from real-world recorded traffic. However, agents replayed from offline data are not reactive and hard to intuitively control. Existing approaches address these challenges by proposing methods that rely on heuristics or generative models of real-world data but these approaches either lack realism or necessitate costly iterative sampling procedures to control the generated behaviours. In this work, we take an alternative approach and propose CtRL-Sim, a method that leverages return-conditioned offline reinforcement learning (RL) to efficiently generate reactive and controllable traffic agents. Specifically, we process real-world driving data through a physics-enhanced Nocturne simulator to generate a diverse offline RL dataset, annotated with various rewards. With this dataset, we train a return-conditioned multi-agent behaviour model that allows for fine-grained manipulation of agent behaviours by modifying the desired returns for the various reward components. This capability enables the generation of a wide range of driving behaviours beyond the scope of the initial dataset, including adversarial behaviours. We show that CtRL-Sim can generate realistic safety-critical scenarios while providing fine-grained control over agent behaviours.
Forward citations
Cited by 4 Pith papers
-
Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
On Waymo Sim Agents, LLM-style tokenization, positional embeddings, pretraining, RL post-training, and test-time search can be adapted to improve motion generation, but not all transfer without domain-specific changes.
-
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
Frame-level meta-actions, predicted and injected at every time step in an autoregressive trajectory model, improve alignment between high-level driving decisions and generated motion.
-
Agent-driven Long-tail Simulation for Autonomous Driving
LLM agents with structured actions can drive interactive long-tail road users in nuPlan, and SemanticPlan shows current planners still fail safety and semantic completion there.
-
End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation
A single diffusion model jointly predicts critical background vehicles' states and low-level controls, with gradient guidance steering speed, drivable-area compliance, and collision avoidance/risk during closed-loop s...
Discussion (0). Continue with ORCID to comment.