REVIEW 4 cited by
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Traffic simulation aims to learn a policy for traffic agents that, when unrolled in closed-loop, faithfully recovers the joint distribution of trajectories observed in the real world. Inspired by large language models, tokenized multi-agent policies have recently become the state-of-the-art in traffic simulation. However, they are typically trained through open-loop behavior cloning, and thus suffer from covariate shift when executed in closed-loop during simulation. In this work, we present Closest Among Top-K (CAT-K) rollouts, a simple yet effective closed-loop fine-tuning strategy to mitigate covariate shift. CAT-K fine-tuning only requires existing trajectory data, without reinforcement learning or generative adversarial imitation. Concretely, CAT-K fine-tuning enables a small 7M-parameter tokenized traffic simulation policy to outperform a 102M-parameter model from the same model family, achieving the top spot on the Waymo Sim Agent Challenge leaderboard at the time of submission. The code is available at https://github.com/NVlabs/catk.
Forward citations
Cited by 4 Pith papers
-
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Hiding future trajectory information until after a driving model forms its decision reduces rationalization and improves verifiable autonomous-driving reasoning in the proposed AD-MCQ and DEFT-RLVR framework.
-
Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
On Waymo Sim Agents, LLM-style tokenization, positional embeddings, pretraining, RL post-training, and test-time search can be adapted to improve motion generation, but not all transfer without domain-specific changes.
-
Improving Traffic Signal Data Quality for the Waymo Open Motion Dataset
A trajectory-based, ring-and-barrier-constrained method imputes 71.7% of missing traffic signal states in the Waymo Open Motion Dataset and lowers the estimated red-light running rate from 15.7% to 2.9%.
-
Surprise Potential as a Measure of Interactivity in Driving Scenarios
A counterfactual surprise metric, Hist-prim with query-centric feedforward prediction and Wasserstein distance, identifies interactive driving scenarios with 0.82+ Spearman correlation to a human-trained reward model.
Discussion (0). Continue with ORCID to comment.