REVIEW 2 major objections 7 minor 58 references
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
T0 review · 2 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read By decoding a traffic agent's future trajectories one at a time, each conditioned on the ones before it, and by rewarding the earliest matching mode during training, ModeSeq obtains diverse, well-calibrated sparse predictions that match…
desk verdict A genuinely new decoding paradigm for multimodal motion prediction, with solid benchmark results and careful ablations, but the 'mode extrapolation' claim outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Sequential mode decoding (Eq. 2) — the factorization of the joint mode-embedding distribution into a chain m_k = Decoder(Ψ, {m_1,...,m_{k-1}}) — together with the Early-Match-Take-All label assignment (Eqs. 7–8). The Memory Transformer makes each mode attend to all previously decoded modes, the Context Transformer fuses scene, map, and agent embeddings, and mode rearrangement between layers sorts embeddings by predicted confidence so the next layer refines the most probable futures first.
What would settle it
Train ModeSeq on a dataset with true mode labels and compare EMTA against a variant that randomizes the decoding order: if random order matches EMTA's mAP and Miss Rate, then the ordering itself, not the sequential conditioning, is irrelevant. Alternatively, find a scenario where decoding more than the training number of modes monotonically increases minFDE, contradicting the extrapolation trend reported in Fig. 5.
Extended reading notes
Core claim
ModeSeq establishes that multimodal motion prediction can be framed as sequence generation over modes: m_k = Decoder(Ψ, {m_1,...,m_{k-1}}), where an ordered chain of K mode embeddings is produced by a recurrent Transformer that attends to a memory bank of earlier modes plus scene context. The matching Early-Match-Take-All loss picks, among all K trajectories that fall within the benchmark's match thresholds, the one decoded at the smallest index as the single positive sample; later matches are treated as negatives, pushing them away from the ground truth to cover other futures. The authors argue this breaks the symmetry of parallel decoding and WTA and produces better-calibrated confidences, and they demonstrate state-of-the-art or balanced results on two benchmarks, plus the ability to decode more modes at inference than seen in training.
Load-bearing premise
The method assumes that a meaningful fixed order over future trajectory modes exists and that the model can learn to place the most likely mode first; if the modes a driver could take have no natural ordering, forcing the earliest matching one to be the sole positive could suppress valid alternatives.
Editorial extensions
If this is right
- Sparse, anchor-free, post-processing-free multimodal prediction can match or surpass dense mode prediction on coverage and scoring metrics.
- Training with fewer modes (e.g., 3) still yields representative, well-scored trajectories, which is useful for onboard latency.
- The model can extrapolate to more modes at inference on demand, helping when future uncertainty is high.
- EMTA creates an inductive bias that other multimodal learning problems could adopt for better diversity and calibration.
Reading between the lines
- The order learned by ModeSeq may act as a form of curriculum or ranking over modes; if so, EMTA is related to learning-to-rank objectives and could be analyzed or extended with ordering losses.
- Mode extrapolation suggests the decoder learns a generative process of modes not tied to a fixed anchor set; a testable consequence is whether extrapolated modes remain diverse and scene-compliant in unfamiliar road topologies.
- Since EMTA treats later matches as negatives, it may under-represent genuinely equiprobable modes; treating later matches as ignored rather than negative would test that boundary.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes ModeSeq, a sequential mode modeling paradigm for multimodal motion prediction. Instead of decoding K trajectory modes in parallel as in DETR-like decoders, ModeSeq generates modes one at a time, conditioning each mode embedding on the previously decoded modes through a Memory Transformer and a Context Transformer, and stacks multiple layers with a mode-rearrangement step. The paper also introduces an Early-Match-Take-All (EMTA) training loss that selects the earliest matching prediction as the positive sample and treats later matches as negatives. Experiments on the Waymo Open Motion Dataset and Argoverse 2 report improvements over QCNet and MTR-series baselines on coverage and confidence metrics, and the abstract and conclusion claim that ModeSeq 'naturally emerges with the capability of mode extrapolation,' i.e., it can generate more than the K=6 modes used in training.
Significance. If validated, ModeSeq would be a meaningful contribution: it offers an end-to-end sparse alternative to dense mode prediction with post-processing, and the ability to vary the number of predicted modes at inference is practically appealing. The paper is well structured, the architecture is clearly described, and the ablations isolate the contributions of sequential decoding, EMTA, iterative refinement, and mode rearrangement. The results on Argoverse 2, where ModeSeq outperforms QCNet and MTR-series on all reported metrics, are particularly encouraging. However, the central 'mode extrapolation' claim is supported only by minFDE and MR curves that improve by construction as K grows, and no confidence-calibration or diversity evidence is provided for K>6. The strengths of the work are the novel sequential factorization and the EMTA training scheme; the main weakness is that the extrapolation evidence does not yet establish the claim as stated.
major comments (2)
- [Section 4.3, Figure 5; Abstract; Section 5] The claim that ModeSeq 'naturally emerges with the capability of mode extrapolation' is not established by the reported evidence. In Figure 5, minFDE and MR are monotonically non-increasing in the number of inference modes by definition: minFDE is the minimum error over the K decoded trajectories, and MR counts cases with no matching trajectory among K candidates, so adding arbitrary extra candidates cannot worsen either metric. The figure therefore does not show that the extra modes are accurate, diverse, or calibrated. To support the claim, please provide a control baseline (e.g., decoding the same top-6 modes and appending perturbed copies, or decoding additional parallel DETR-style heads), and report confidence-aware metrics for K>6, such as mAP6 or Soft mAP6 computed on the top-6 subset of the extrapolated set, or per-mode precision/recall. Without such evidence, the extrapolation claim should be removed or substantially weakened.
- [Section 3.6, Eqs. (7)-(8); Tables 3-4] EMTA relies on the assumption that a fixed sequential order over modes is learnable and that earlier modes should correspond to higher-likelihood futures. The ablations show that label-assignment choices have a material effect: under EMTA, changing from 'None' to 'Other Matches' for ignored samples changes Soft mAP6 from 0.4231 to 0.4098 (Table 3), and the effect of mode rearrangement depends on the ignored-sample definition (Table 4). However, the paper does not analyze what the learned order encodes or whether the benefit of EMTA persists under alternative orderings, such as random permutations of mode indices during training or a confidence-sorted order without the sequential conditioning. Because the method's novelty depends on the ordering being a meaningful inductive bias, please provide such analysis or explicitly discuss the sensitivity.
minor comments (7)
- [Section 3.4, Eq. (3)] The initial mode embedding e is said to be 'randomly initialized at the beginning of training'; please specify the initialization distribution (e.g., truncated normal with a given standard deviation) for reproducibility.
- [Section 3.4, paragraph after Eq. (5)] There is a typo in 'thek-th mode embedding'; it should be 'the k-th mode embedding'.
- [Section 4.1, Metrics] The definition of b-minFDE says 'summing the minFDEK and the Brier scores of the best modes'; this should be 'the Brier score of the best mode' or should clarify the aggregation over modes.
- [Tables 3 and 4] The 'Ignored Samples' configurations are not defined in the main text. A brief explanation of what 'Other Matches' and 'Early Mismatches' mean, or a pointer to the relevant sentence in Section 3.6, would help the reader interpret the ablations.
- [Figure 5] The axis labels 'minFDE' and 'MR' should include the relevant number of modes (e.g., minFDE_K, MR_K) to avoid confusion with the benchmark's fixed K=6 metrics.
- [Section 4.2 and Supplementary Section 7] The single-model results in Tables 1 and 2 appear to come from a single training run without variance estimates. Given the small margins over QCNet on the WOMD validation set (e.g., Soft mAP6 0.4562 vs. 0.4508), reporting multiple seeds or explicitly stating the single-run convention would improve confidence in the comparison.
- [Supplementary Section 7] The hand-tuned scaling factors (1.5, 1.4, 1.4) used in Weighted Trajectory Fusion are a form of post-processing with tuned hyperparameters; the main text should be clear that the 'no heuristic post-processing' claim applies to the single-model results, not the ensemble results.
Circularity Check
Mode-extrapolation evidence in Fig. 5 is entailed by the definitions of minFDE and MR; benchmark and ablation results are otherwise self-contained.
-
self definitional
[Section 4.3, 'Capability of Mode Extrapolation' (Fig. 5); abstract and conclusion 'mode extrapolation' claims]
"We ask the model trained by generating 6 modes to execute more decoding steps at test time. As depicted in Fig. 5, ModeSeq achieves lower prediction error with the increase of the decoded modes, emerging with the capability of mode extrapolation thanks to sequential modeling."
minFDE_K is a minimum over K trajectories and MR_K is the fraction of cases with no matching trajectory; both are monotone non-increasing as K grows for any model, even one emitting arbitrary extra candidates. The paper provides no control (e.g., perturbed copies of top modes, extra parallel heads, random trajectories) and no confidence-aware metric for K>6, so Fig. 5 cannot distinguish learned extrapolation from the trivial effect of adding candidates. The 'emerging capability' claim is therefore entailed by the metric definitions, not by sequential decoding.
full rationale
The main benchmark comparisons (Tables 1-2) and ablations (Tables 3-4) are evaluated on held-out Waymo and Argoverse 2 splits, with the QCNet encoder used as a fixed off-the-shelf component and also as a baseline; the supplementary Table 6 shows ModeSeq improves other encoders, so the self-citations are not load-bearing. Using benchmark-defined match thresholds inside the EMTA loss is a training-time design choice and does not by itself make the held-out mAP/MR numbers circular. The one genuine reduction is the mode-extrapolation claim: Fig. 5's minFDE/MR curves are definitionally monotone non-increasing in the number of modes, so they cannot by themselves establish that ModeSeq 'naturally emerges' with extrapolation. Because that extrapolation claim is advertised in the abstract and conclusion, the paper receives a partial-circularity score of 6; absent that figure, the remainder of the derivation would score near 0.
Assumptions & free parameters
free parameters (1)
- Ensemble trajectory fusion scaling factors =
1.5 (vehicles), 1.4 (pedestrians), 1.4 (cyclists)
assumptions (3)
- domain assumption Sequential factorization m_k = Decoder(Ψ, {m_1,...,m_{k-1}}) can represent the joint distribution of an unordered set of future modes.
- domain assumption Benchmark match thresholds (Eqs. 9-12) define behavioral equivalence between a prediction and the ground truth for both training and evaluation.
- ad hoc to paper Monotonically decreasing confidence over decoded mode order is a learnable and useful target.
Cite this review
Pith. "Pith review of ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling." pith.science (2026). https://pith.science/paper/GHBPCIY4
@misc{pith2026241111911,
author = {Pith},
title = {Pith review of: ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/GHBPCIY4}},
note = {Machine review of arXiv:2411.11911}
}
read the original abstract
Anticipating the multimodality of future events lays the foundation for safe autonomous driving. However, multimodal motion prediction for traffic agents has been clouded by the lack of multimodal ground truth. Existing works predominantly adopt the winner-take-all training strategy to tackle this challenge, yet still suffer from limited trajectory diversity and uncalibrated mode confidence. While some approaches address these limitations by generating excessive trajectory candidates, they necessitate a post-processing stage to identify the most representative modes, a process lacking universal principles and compromising trajectory accuracy. We are thus motivated to introduce ModeSeq, a new multimodal prediction paradigm that models modes as sequences. Unlike the common practice of decoding multiple plausible trajectories in one shot, ModeSeq requires motion decoders to infer the next mode step by step, thereby more explicitly capturing the correlation between modes and significantly enhancing the ability to reason about multimodality. Leveraging the inductive bias of sequential mode prediction, we also propose the Early-Match-Take-All (EMTA) training strategy to diversify the trajectories further. Without relying on dense mode prediction or heuristic post-processing, ModeSeq considerably improves the diversity of multimodal output while attaining satisfactory trajectory accuracy, resulting in balanced performance on motion prediction benchmarks. Moreover, ModeSeq naturally emerges with the capability of mode extrapolation, which supports forecasting more behavior modes when the future is highly uncertain.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
So- cial lstm: Human trajectory prediction in crowded spaces
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. So- cial lstm: Human trajectory prediction in crowded spaces. In CVPR, 2016. 2, 3
work page 2016
- [2]
-
[3]
End-to- end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to- end object detection with transformers. In ECCV, 2020. 2, 5, 7
work page 2020
-
[4]
Multipath: Multiple probabilistic anchor trajec- tory hypotheses for behavior prediction
Yuning Chai, Benjamin Sapp, Mayank Bansal, and Dragomir Anguelov. Multipath: Multiple probabilistic anchor trajec- tory hypotheses for behavior prediction. In CoRL, 2019. 1, 2
work page 2019
-
[5]
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, C ¸ aglar G ¨ulc ¸ehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using rnn encoder-decoder for statistical machine translation. In EMNLP, 2014. 2, 5
work page 2014
-
[6]
Multimodal trajectory predictions for autonomous driving using deep convolutional networks
Henggang Cui, Vladan Radosavljevic, Fang-Chieh Chou, Tsung-Han Lin, Thi Nguyen, Tzu-Kuo Huang, Jeff Schnei- der, and Nemanja Djuric. Multimodal trajectory predictions for autonomous driving using deep convolutional networks. In ICRA, 2019. 1, 2
work page 2019
-
[7]
Convolutional social pooling for vehicle trajectory prediction
Nachiket Deo and Mohan M Trivedi. Convolutional social pooling for vehicle trajectory prediction. In CVPRW, 2018. 2
work page 2018
-
[8]
Ep- silon: An efficient planning system for automated vehicles in highly interactive environments
Wenchao Ding, Lu Zhang, Jing Chen, and Shaojie Shen. Ep- silon: An efficient planning system for automated vehicles in highly interactive environments. T-RO, 2021. 1
work page 2021
Show all 58 references
-
[9]
Qi, Yin Zhou, Zoey Yang, Aur´elien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov
Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu, Hang Zhao, Sabeek Pradhan, Yuning Chai, Ben Sapp, Charles R. Qi, Yin Zhou, Zoey Yang, Aur´elien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov. Large scale i...
2021
-
[10]
Transformer networks for trajectory forecasting
Francesco Giuliari, Irtiza Hasan, Marco Cristani, and Fabio Galasso. Transformer networks for trajectory forecasting. In ICPR, 2020. 2
2020
-
[11]
Social gan: Socially acceptable trajec- tories with generative adversarial networks
Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social gan: Socially acceptable trajec- tories with generative adversarial networks. In CVPR, 2018. 2
2018
-
[12]
Long short-term memory
Sepp Hochreiter and J ¨urgen Schmidhuber. Long short-term memory. Neural Computation, 1997. 2, 5
1997
-
[13]
Rules of the road: Predicting driving behavior with a convolutional model of semantic interactions
Joey Hong, Benjamin Sapp, and James Philbin. Rules of the road: Predicting driving behavior with a convolutional model of semantic interactions. In CVPR, 2019. 2
2019
-
[14]
Desire: Distant future prediction in dynamic scenes with interacting agents
Namhoon Lee, Wongun Choi, Paul Vernaza, Christopher B Choy, Philip HS Torr, and Manmohan Chandraker. Desire: Distant future prediction in dynamic scenes with interacting agents. In CVPR, 2017. 2
2017
-
[15]
Stochastic multiple choice learning for training diverse deep ensembles
Stefan Lee, Senthil Purushwalkam Shiva Prakash, Michael Cogswell, Viresh Ranjan, David Crandall, and Dhruv Batra. Stochastic multiple choice learning for training diverse deep ensembles. In NIPS, 2016. 1, 2, 5
2016
-
[16]
Marc: Multipolicy and risk-aware contingency planning for au- tonomous driving
Tong Li, Lu Zhang, Sikang Liu, and Shaojie Shen. Marc: Multipolicy and risk-aware contingency planning for au- tonomous driving. RA-L, 2023. 1
2023
-
[17]
Learning lane graph representa- tions for motion forecasting
Ming Liang, Bin Yang, Rui Hu, Yun Chen, Renjie Liao, Song Feng, and Raquel Urtasun. Learning lane graph representa- tions for motion forecasting. In ECCV, 2020. 1, 2
2020
-
[18]
Eda: Evolving and distinct anchors for multimodal motion prediction
Longzhong Lin, Xuewu Lin, Tianwei Lin, Lichao Huang, Rong Xiong, and Yue Wang. Eda: Evolving and distinct anchors for multimodal motion prediction. In AAAI, 2024. 1, 5
2024
-
[19]
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll´ar. Focal loss for dense object detection. In CVPR,
-
[20]
Multimodal motion prediction with stacked transformers
Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang, and Bolei Zhou. Multimodal motion prediction with stacked transformers. In CVPR, 2021. 2
2021
-
[21]
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICLR, 2017. 6
2017
-
[22]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 6
2019
-
[23]
Overcoming limitations of mixture density networks: A sam- pling and fitting framework for multimodal future prediction
Osama Makansi, Eddy Ilg, Ozgun Cicek, and Thomas Brox. Overcoming limitations of mixture density networks: A sam- pling and fitting framework for multimodal future prediction. In CVPR, 2019. 1, 2, 5
2019
-
[24]
Multi-head attention for multi-modal joint vehicle mo- tion forecasting
Jean Mercat, Thomas Gilles, Nicole El Zoghby, Guil- laume Sandou, Dominique Beauvois, and Guillermo Pita Gil. Multi-head attention for multi-modal joint vehicle mo- tion forecasting. In ICRA, 2020. 2
2020
-
[25]
Wayformer: Motion forecasting via simple & efficient attention networks
Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou, Kratarth Goel, Khaled S Refaat, and Benjamin Sapp. Wayformer: Motion forecasting via simple & efficient attention networks. In ICRA, 2023. 1, 2, 5
2023
-
[26]
Scene transformer: A unified architecture for predicting multiple agent trajectories
Jiquan Ngiam, Benjamin Caine, Vijay Vasudevan, Zheng- dong Zhang, Hao-Tien Lewis Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu, Ashish Venugopal, David Weiss, Ben Sapp, Zhifeng Chen, and Jonathon Shlens. Scene transformer: A unified architecture for predicting...
2022
-
[27]
Covernet: Multimodal behavior prediction using trajectory sets
Tung Phan-Minh, Elena Corina Grigore, Freddy A Boulton, Oscar Beijbom, and Eric M Wolff. Covernet: Multimodal behavior prediction using trajectory sets. In CVPR, 2020. 2
2020
-
[28]
Trajeglish: Traffic modeling as next-token prediction
Jonah Philion, Xue Bin Peng, and Sanja Fidler. Trajeglish: Traffic modeling as next-token prediction. In ICLR, 2024. 2, 3
2024
-
[29]
R2p2: A reparameterized pushforward policy for diverse, precise generative path forecasting
Nicholas Rhinehart, Kris M Kitani, and Paul Vernaza. R2p2: A reparameterized pushforward policy for diverse, precise generative path forecasting. In ECCV, 2018. 2
2018
-
[30]
Precog: Prediction conditioned on goals in visual multi-agent settings
Nicholas Rhinehart, Rowan McAllister, Kris Kitani, and Sergey Levine. Precog: Prediction conditioned on goals in visual multi-agent settings. In ICCV, 2019. 2, 3 9
2019
-
[31]
Fjmp: Factorized joint multi-agent motion prediction over learned directed acyclic interaction graphs
Luke Rowe, Martin Ethier, Eli-Henry Dykhne, and Krzysztof Czarnecki. Fjmp: Factorized joint multi-agent motion prediction over learned directed acyclic interaction graphs. In CVPR, 2023. 3
2023
-
[32]
Learning in an uncertain world: Representing am- biguity through multiple hypotheses
Christian Rupprecht, Iro Laina, Robert DiPietro, Maximil- ian Baust, Federico Tombari, Nassir Navab, and Gregory D Hager. Learning in an uncertain world: Representing am- biguity through multiple hypotheses. In ICCV, 2017. 1, 2, 5
2017
-
[33]
Trajectron++: Dynamically-feasible trajec- tory forecasting with heterogeneous data
Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-feasible trajec- tory forecasting with heterogeneous data. In ECCV, 2020. 2
2020
-
[34]
Motionlm: Multi-agent motion forecasting as language modeling
Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S Refaat, Rami Al-Rfou, and Benjamin Sapp. Motionlm: Multi-agent motion forecasting as language modeling. In ICCV, 2023. 2, 3
2023
-
[35]
Self- attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self- attention with relative position representations. In NAACL,
-
[36]
Mtr v3: 1st place so- lution for 2024 waymo open dataset challenge - motion pre- diction
Chen Shi, Shaoshuai Shi, and Li Jiang. Mtr v3: 1st place so- lution for 2024 waymo open dataset challenge - motion pre- diction. In CVPR 2024 Workshop on Autonomous Driving ,
2024
-
[37]
Motion transformer with global intention localization and lo- cal movement refinement
Shaoshuai Shi, Li Jiang, Dengxin Dai, and Bernt Schiele. Motion transformer with global intention localization and lo- cal movement refinement. In NeurIPS, 2022. 1, 2, 5, 6
2022
-
[38]
Mtr++: Multi-agent motion prediction with symmetric scene modeling and guided intention querying
Shaoshuai Shi, Li Jiang, Dengxin Dai, and Bernt Schiele. Mtr++: Multi-agent motion prediction with symmetric scene modeling and guided intention querying. TPAMI, 2024. 6, 1
2024
-
[39]
Weighted boxes fusion: Ensembling boxes from different ob- ject detection models
Roman Solovyev, Weimin Wang, and Tatiana Gabruseva. Weighted boxes fusion: Ensembling boxes from different ob- ject detection models. Image and Vision Computing, 2021. 1
2021
-
[40]
Rmp-yolo: A robust motion predictor for partially observable scenarios even if you only look once
Jiawei Sun, Jiahui Li, Tingchen Liu, Chengran Yuan, Shuo Sun, Zefan Huang, Anthony Wong, Keng Peng Tee, and Marcelo H Ang Jr. Rmp-yolo: A robust motion predictor for partially observable scenarios even if you only look once. arXiv preprint arXiv:2409.11696, 2024. 6
2024 arXiv
-
[41]
M2i: From factored marginal trajectory pre- diction to interactive prediction
Qiao Sun, Xin Huang, Junru Gu, Brian C Williams, and Hang Zhao. M2i: From factored marginal trajectory pre- diction to interactive prediction. In CVPR, 2022. 3
2022
-
[42]
Multiple futures prediction
Yichuan Charlie Tang and Ruslan Salakhutdinov. Multiple futures prediction. In NeurIPS, 2019. 2, 3
2019
-
[43]
Ana- lyzing the variety loss in the context of probabilistic trajec- tory prediction
Luca Anthony Thiede and Pratik Prabhanjan Brahma. Ana- lyzing the variety loss in the context of probabilistic trajec- tory prediction. In ICCV, 2019. 1, 2, 5
2019
-
[44]
Multipath++: Efficient in- formation fusion and trajectory aggregation for behavior pre- diction
Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivas- tava, Khaled S Refaat, Nigamaa Nayakanti, Andre Cornman, Kan Chen, Bertrand Douillard, Chi Pang Lam, Dragomir Anguelov, and Benjamin Sapp. Multipath++: Efficient in- formation fusion and trajectory aggregation for behavior...
2022
-
[45]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017. 2
2017
-
[46]
Argoverse 2: Next generation datasets for self-driving perception and fore- casting
Benjamin Wilson, William Qi, Tanmay Agarwal, John Lam- bert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Rat- nesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, Deva Ramanan, Peter Carr, and James Hays. Argoverse 2: Next generation datasets for self-driving perception an...
2021
-
[47]
Spatio-temporal graph transformer networks for pedestrian trajectory prediction
Cunjun Yu, Xiao Ma, Jiawei Ren, Haiyu Zhao, and Shuai Yi. Spatio-temporal graph transformer networks for pedestrian trajectory prediction. In ECCV, 2020. 2
2020
-
[48]
Tnt: Target-driven trajectory prediction
Hang Zhao, Jiyang Gao, Tian Lan, Chen Sun, Ben Sapp, Balakrishnan Varadarajan, Yue Shen, Yi Shen, Yuning Chai, Cordelia Schmid, Congcong Li, and Dragomir Anguelov. Tnt: Target-driven trajectory prediction. In CoRL, 2020. 2
2020
-
[49]
Smartrefine: A scenario-adaptive refinement framework for efficient motion prediction
Yang Zhou, Hao Shao, Letian Wang, Steven L Waslander, Hongsheng Li, and Yu Liu. Smartrefine: A scenario-adaptive refinement framework for efficient motion prediction. In CVPR, 2024. 2
2024
-
[50]
Hivt: Hierarchical vector transformer for multi-agent motion prediction
Zikang Zhou, Luyao Ye, Jianping Wang, Kui Wu, and Kejie Lu. Hivt: Hierarchical vector transformer for multi-agent motion prediction. In CVPR, 2022. 1, 2, 6
2022
-
[51]
Query-centric trajectory prediction
Zikang Zhou, Jianping Wang, Yung-Hui Li, and Yu-Kai Huang. Query-centric trajectory prediction. In CVPR, 2023. 1, 2, 3, 5, 6, 7, 8
2023
-
[52]
Behaviorgpt: Smart agent simulation for autonomous driving with next-patch prediction
Zikang Zhou, Haibo Hu, Xinhong Chen, Jianping Wang, Nan Guan, Kui Wu, Yung-Hui Li, Yu-Kai Huang, and Chun Jason Xue. Behaviorgpt: Smart agent simulation for autonomous driving with next-patch prediction. In NeurIPS,
-
[54]
By linearly scaling the 2-meter threshold across time steps, we obtain a distance threshold Γ(t) for each time step t: Γ (t) = t 30
Definition of a Match The Argoverse 2 Motion Forecasting Benchmark [46] de- sires the predictions’ displacement error at the 60-th time step to be less than2 meters. By linearly scaling the 2-meter threshold across time steps, we obtain a distance threshold Γ(t) for each time ...
-
[55]
Our ensemble method is almost the same as WBF, except we are fus- ing trajectories according to distance thresholds rather than Table 6
Ensemble Method on the WOMD Inspired by Weighted Boxes Fusion (WBF) [39], we pro- pose Weighted Trajectory Fusion to aggregate multimodal trajectories produced by multiple models. Our ensemble method is almost the same as WBF, except we are fus- ing trajectories according to d...
1911
-
[56]
early match
Versatility Our ModeSeq framework can be seamlessly integrated with other scene encoders. As shown in Tab. 6, our ap- proach significantly enhances Scene Transformer [26] and HiVT [50], two representative methods adopting scene- centric and agent-centric encoders, respectively...
-
[57]
Parameter Efficiency The ModeSeq decoder comprises 7.5M parameters, total- ing 10.9M parameters for the overall model when combined with the QCNet encoder [51], which is far more parameter- efficient than other state-of-the-art on the WOMD ( e.g., MTR [37] with 65.8M parameter...
-
[58]
4 to demonstrate our approach’s ability to produce representative trajectories and extrapolate more modes
More Qualitative Results Figure 6 supplements the results in Fig. 4 to demonstrate our approach’s ability to produce representative trajectories and extrapolate more modes. 1 (a) #Mode@Training=3, #Mode@Inference=3 (b) #Mode@Training=6, #Mode@Inference=6 (c) #Mode@Training=6, ...
-
[2024]
2, 3 10 ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling Supplementary Material
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.