Pith. sign in

REVIEW 4 cited by

Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05334 v2 pith:2PW4MWMS submitted 2024-12-05 cs.LG

classification cs.LG
keywords trafficclosed-loopfine-tuningsimulationcat-ktokenizedcovariatem-parameter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traffic simulation aims to learn a policy for traffic agents that, when unrolled in closed-loop, faithfully recovers the joint distribution of trajectories observed in the real world. Inspired by large language models, tokenized multi-agent policies have recently become the state-of-the-art in traffic simulation. However, they are typically trained through open-loop behavior cloning, and thus suffer from covariate shift when executed in closed-loop during simulation. In this work, we present Closest Among Top-K (CAT-K) rollouts, a simple yet effective closed-loop fine-tuning strategy to mitigate covariate shift. CAT-K fine-tuning only requires existing trajectory data, without reinforcement learning or generative adversarial imitation. Concretely, CAT-K fine-tuning enables a small 7M-parameter tokenized traffic simulation policy to outperform a 102M-parameter model from the same model family, achieving the top spot on the Waymo Sim Agent Challenge leaderboard at the time of submission. The code is available at https://github.com/NVlabs/catk.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

    cs.AI 2026-08 conditional novelty 7.0 of 10

    Hiding future trajectory information until after a driving model forms its decision reduces rationalization and improves verifiable autonomous-driving reasoning in the proposed AD-MCQ and DEFT-RLVR framework.

  2. Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving

    cs.AI 2025-09 conditional novelty 6.0 of 10

    On Waymo Sim Agents, LLM-style tokenization, positional embeddings, pretraining, RL post-training, and test-time search can be adapted to improve motion generation, but not all transfer without domain-specific changes.

  3. Improving Traffic Signal Data Quality for the Waymo Open Motion Dataset

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A trajectory-based, ring-and-barrier-constrained method imputes 71.7% of missing traffic signal states in the Waymo Open Motion Dataset and lowers the estimated red-light running rate from 15.7% to 2.9%.

  4. Surprise Potential as a Measure of Interactivity in Driving Scenarios

    cs.RO 2025-02 conditional novelty 5.0 of 10

    A counterfactual surprise metric, Hist-prim with query-centric feedforward prediction and Wasserstein distance, identifies interactive driving scenarios with 0.82+ Spearman correlation to a human-trained reward model.

Pith tools