Pith. sign in

REVIEW 8 cited by

Gen-Drive: Enhancing Diffusion Generative Driving Policies with Reward Modeling and Reinforcement Learning Fine-tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05582 v1 pith:NOJ5QKRB submitted 2024-10-08 cs.RO

classification cs.RO
keywords planningframeworkmodeldiffusiondrivingenhancingfine-tuningreward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Autonomous driving necessitates the ability to reason about future interactions between traffic agents and to make informed evaluations for planning. This paper introduces the \textit{Gen-Drive} framework, which shifts from the traditional prediction and deterministic planning framework to a generation-then-evaluation planning paradigm. The framework employs a behavior diffusion model as a scene generator to produce diverse possible future scenarios, thereby enhancing the capability for joint interaction reasoning. To facilitate decision-making, we propose a scene evaluator (reward) model, trained with pairwise preference data collected through VLM assistance, thereby reducing human workload and enhancing scalability. Furthermore, we utilize an RL fine-tuning framework to improve the generation quality of the diffusion model, rendering it more effective for planning tasks. We conduct training and closed-loop planning tests on the nuPlan dataset, and the results demonstrate that employing such a generation-then-evaluation strategy outperforms other learning-based approaches. Additionally, the fine-tuned generative driving policy shows significant enhancements in planning performance. We further demonstrate that utilizing our learned reward model for evaluation or RL fine-tuning leads to better planning performance compared to relying on human-designed rewards. Project website: https://mczhi.github.io/GenDrive.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

    cs.CV 2025-06 conditional novelty 8.0 of 10

    RDPO builds preference pairs by reverse-sampling real video latents with a pre-trained generator, then fine-tunes with Flow-DPO, improving physics consistency metrics on two video models.

  2. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Hiding future trajectory information until after a driving model forms its decision reduces rationalization and improves verifiable autonomous-driving reasoning in the proposed AD-MCQ and DEFT-RLVR framework.

  3. Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving

    cs.RO 2025-06 reject novelty 6.0 of 10

    R2SE refines pretrained end-to-end driving policies on hard cases via residual LoRA reinforcement learning and switches between specialist and generalist policies using GPD-based uncertainty.

  4. Autoregressive Meta-Actions for Unified Controllable Trajectory Generation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Frame-level meta-actions, predicted and injected at every time step in an autoregressive trajectory model, improve alignment between high-level driving decisions and generated motion.

  5. Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling

    cs.RO 2025-05 conditional novelty 6.0 of 10

    By replacing group normalization in GRPO with fixed scaling, Plan-R1 keeps safety violations dominant in the learning signal and achieves state-of-the-art reactive planning scores on nuPlan.

  6. Pulse Breathing Dynamics in a Mode-Locked Laser measured via SHG autocorrelation

    physics.optics 2026-03 unverdicted novelty 5.0 of 10

    A statistical SHG-autocorrelation Fano analysis is claimed to expose pulse breathing and measure ~3 fs pulse-width fluctuations on two commercial mode-locked lasers.

  7. DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A synthetic dataset and visual prompting framework improve VLM-based driving risk prediction, with the main evidence coming from an unreleased private test set.

  8. A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach

    cs.RO 2025-12 conditional novelty 3.0 of 10

    A position/review paper argues data-driven model predictive control is the best route to safe, adaptive, human-like autonomous-driving motion planning, but provides no new derivation or experiment.

Pith tools