Pith. sign in

REVIEW 1 cited by

ST-P3: End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07601 v2 pith:OJOWYIDS submitted 2022-07-15 cs.CV

classification cs.CV
keywords vision-basedautonomousdrivingend-to-endfeaturelearningspatial-temporalst-p3
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many existing autonomous driving paradigms involve a multi-stage discrete pipeline of tasks. To better predict the control signals and enhance user safety, an end-to-end approach that benefits from joint spatial-temporal feature learning is desirable. While there are some pioneering works on LiDAR-based input or implicit design, in this paper we formulate the problem in an interpretable vision-based setting. In particular, we propose a spatial-temporal feature learning scheme towards a set of more representative features for perception, prediction and planning tasks simultaneously, which is called ST-P3. Specifically, an egocentric-aligned accumulation technique is proposed to preserve geometry information in 3D space before the bird's eye view transformation for perception; a dual pathway modeling is devised to take past motion variations into account for future prediction; a temporal-based refinement unit is introduced to compensate for recognizing vision-based elements for planning. To the best of our knowledge, we are the first to systematically investigate each part of an interpretable end-to-end vision-based autonomous driving system. We benchmark our approach against previous state-of-the-arts on both open-loop nuScenes dataset as well as closed-loop CARLA simulation. The results show the effectiveness of our method. Source code, model and protocol details are made publicly available at https://github.com/OpenPerceptionX/ST-P3.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    ALN-P3 adds three alignment losses between a driving stack and a language model during training, improving both planning safety and language reasoning on nuScenes, Nu-X, TOD3Cap, and nuScenes-QA.

Pith tools