D²-City: A Large-Scale Dashcam Video Dataset of Diverse Traffic Scenarios

Zhengping Che , Guangyu Li , Tracy Li , Bo Jiang , Xuefeng Shi , Xinsheng Zhang , Ying Lu , Guobin Wu

show 2 more authors

Yan Liu Jieping Ye

Authors on Pith no claims yet

classification 💻 cs.LG cs.CVstat.ML

keywords cityannotationsdatasetsdrivingtrafficvideosdashcamdataset

0 comments

read the original abstract

Driving datasets accelerate the development of intelligent driving and related computer vision technologies, while substantial and detailed annotations serve as fuels and powers to boost the efficacy of such datasets to improve learning-based models. We propose D$^2$-City, a large-scale comprehensive collection of dashcam videos collected by vehicles on DiDi's platform. D$^2$-City contains more than 10000 video clips which deeply reflect the diversity and complexity of real-world traffic scenarios in China. We also provide bounding boxes and tracking annotations of 12 classes of objects in all frames of 1000 videos and detection annotations on keyframes for the remainder of the videos. Compared with existing datasets, D$^2$-City features data in varying weather, road, and traffic conditions and a huge amount of elaborate detection and tracking annotations. By bringing a diverse set of challenging cases to the community, we expect the D$^2$-City dataset will advance the perception and related areas of intelligent driving.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

A global dataset of continuous urban dashcam driving
cs.CV 2026-04 accept novelty 7.0

CROWD is a new global dataset of 51,753 continuous urban dashcam segments spanning over 20,000 hours from 238 countries, with manual labels and automated object detections for routine driving analysis.
Learning from the Unseen: Generative Data Augmentation for Geometric-Semantic Accident Anticipation
cs.CV 2026-04 unverdicted novelty 5.0

A generative video synthesis pipeline paired with a semantic graph neural network yields gains in accident anticipation accuracy and lead time on driving datasets, accompanied by a new benchmark release.