Pith. sign in

REVIEW 1 cited by

UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-View

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.08536 v2 pith:Z7RA3T4I submitted 2022-07-18 cs.CV

classification cs.CV
keywords fusionunifiedmethodtemporalconventionalmethodsproposedrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Bird's eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusion and merges them into a unified mathematical formulation. The unified fusion could not only provide a new perspective on BEV fusion but also brings new capabilities. With the proposed unified spatial-temporal fusion, our method could support long-range fusion, which is hard to achieve in conventional BEV methods. Moreover, the BEV fusion in our work is temporal-adaptive and the weights of temporal fusion are learnable. In contrast, conventional methods mainly use fixed and equal weights for temporal fusion. Besides, the proposed unified fusion could avoid information lost in conventional BEV fusion methods and make full use of features. Extensive experiments and ablation studies on the NuScenes dataset show the effectiveness of the proposed method and our method gains the state-of-the-art performance in the map segmentation task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Anchor3DLane++: 3D Lane Detection via Sample-Adaptive Sparse 3D Anchor Regression

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Anchor3DLane++ predicts 3D lanes from front-view features using sample-adaptive sparse 3D anchors, improving F1 scores on OpenLane, ApolloSim, and ONCE-3DLanes beyond prior methods.

Pith tools