Pith. sign in

REVIEW 8 cited by

Wayformer: Motion Forecasting via Simple & Efficient Attention Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.05844 v1 pith:GNC4AE76 submitted 2022-07-12 cs.CV

classification cs.CV
keywords attentionforecastingfusionmotionwayformercomplexdesigndiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Motion forecasting for autonomous driving is a challenging task because complex driving scenarios result in a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road geometry, lane connectivity, time-varying traffic light state, and history of a dynamic set of agents and their interactions into an effective encoding. To model this diverse set of input features, many approaches proposed to design an equally complex system with a diverse set of modality specific modules. This results in systems that are difficult to scale, extend, or tune in rigorous ways to trade off quality and efficiency. In this paper, we present Wayformer, a family of attention based architectures for motion forecasting that are simple and homogeneous. Wayformer offers a compact model description consisting of an attention based scene encoder and a decoder. In the scene encoder we study the choice of early, late and hierarchical fusion of the input modalities. For each fusion type we explore strategies to tradeoff efficiency and quality via factorized attention or latent query attention. We show that early fusion, despite its simplicity of construction, is not only modality agnostic but also achieves state-of-the-art results on both Waymo Open MotionDataset (WOMD) and Argoverse leaderboards, demonstrating the effectiveness of our design philosophy

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Class-Incremental Motion Forecasting

    cs.CV 2026-03 conditional novelty 7.0 of 10

    OMEN is the first end-to-end class-incremental motion forecaster that retains old-class accuracy via VLM-filtered future-detection pseudo-labels and variance-based sequence replay.

  2. Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Polynomial representations of trajectories and maps yield competitive prediction accuracy while substantially improving cross-dataset generalization and computational efficiency in autonomous driving.

  3. Multi-Agent Path Finding Among Dynamic Uncontrollable Agents with Statistical Safety Guarantees

    cs.MA 2025-07 conditional novelty 6.0 of 10

    CP-Solver integrates conformal prediction intervals into Enhanced Conflict-Based Search to give statistical collision-safety guarantees for multi-agent path finding among dynamic uncontrollable agents.

  4. GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GoIRL couples maximum entropy inverse reinforcement learning with vectorized lane-graph features to predict multiple future trajectories, reporting competitive benchmark numbers but not the stated state-of-the-art on ...

  5. A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles

    cs.AI 2026-07 conditional novelty 5.5 of 10

    A dynamic graph-attention model jointly predicts every nearby vehicle’s lane-change intention and trajectory, cutting trajectory error by up to ~53% and improving scene coherence on NGSIM and highD.

  6. Learning High-Level Decision Making with an Interaction-Aware Attention-Based Network in Autonomous Driving

    cs.RO 2026-06 conditional novelty 5.0 of 10

    An attention architecture that bottlenecks traffic agents into fixed latent queries plus a finer discrete action set yields higher simulated speeds and lower early-termination rates than DeepSet and Ego-attention on t...

  7. Contrast & Compress: Learning Lightweight Embeddings for Short Trajectories

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A small Transformer trained with a cosine-based triplet loss learns 16-dimensional embeddings that retrieve similar short driving trajectories from Argoverse 2 substantially better than FFT-based triplet training.

  8. PhysVarMix: Physics-Informed Variational Mixture Model for Multi-Modal Trajectory Prediction

    cs.RO 2025-07 reject novelty 3.0 of 10

    PhysVarMix is a claimed physics-informed variational mixture model for multimodal trajectory prediction, but its latent variable never affects the predicted trajectories, so it reduces to a standard mixture density ne...

Pith tools