Pith. sign in

REVIEW 15 cited by

BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.04687 v2 pith:BBAXQOYL submitted 2018-05-12 cs.CV

classification cs.CV
keywords datasetdrivingtasksbdd100kheterogeneouslearningmultitaskautonomous
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving. Researchers are usually constrained to study a small set of problems on one dataset, while real-world computer vision applications require performing tasks of various complexities. We construct BDD100K, the largest driving video dataset with 100K videos and 10 tasks to evaluate the exciting progress of image recognition algorithms on autonomous driving. The dataset possesses geographic, environmental, and weather diversity, which is useful for training models that are less likely to be surprised by new conditions. Based on this diverse dataset, we build a benchmark for heterogeneous multitask learning and study how to solve the tasks together. Our experiments show that special training strategies are needed for existing models to perform such heterogeneous tasks. BDD100K opens the door for future studies in this important venue.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A multi-vehicle naturalistic benchmark finds learned driving embeddings retain driver identity under condition matching, while descriptors collapse and video re-ID is mostly route leakage.

  2. Vision as Unified Multimodal Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.

  3. Steadily moving semi-infinite fracture in plane poroelasticity

    physics.geo-ph 2026-04 unverdicted novelty 7.0 of 10

    XEmbodied achieves SOTA on 18 embodied VQA benchmarks by fusing 3D geometric tokens and distilled physical cues into a 30B VLM with progressive curriculum training.

  4. Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning

    cs.CV 2025-11 reject novelty 6.0 of 10

    INTSD, a 6,004-image Indian nighttime traffic-sign dataset, is introduced with LENS-Net, which reports 92.56 mAP@50 detection and 78.89 macro-precision classification, but the paper's internal statistics are inconsistent.

  5. VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors

    cs.CV 2025-10 conditional novelty 6.0 of 10

    An IoU-weighted entropy objective and image-conditioned prompt selection adapt YOLO-World and Grounding DINO at test time, improving robustness on style, weather, low-light, and corruption shifts without labels.

  6. RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 2,100-video database with human opinions shows that current video quality models underperform on robot-generated content, motivating a new VQA subfield.

  7. GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GaRA-SAM improves SAM's robustness to image corruption by using input-dependent gating to adjust the effective rank of low-rank adapters, beating prior methods on robust segmentation benchmarks.

  8. Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Transfer-learned SpikeYOLO achieves mAP 0.937/0.771 and HOTA 0.701/0.445 on KITTI and BDD100K for two-class automotive detection and tracking, competitive with conventional deep networks.

  9. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  10. CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios

    cs.CV 2025-07 conditional novelty 5.0 of 10

    CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.

  11. EV-LayerSegNet: Self-supervised Motion Segmentation using Event Cameras

    cs.CV 2025-06 reject novelty 5.0 of 10

    A self-supervised CNN segments moving objects in event camera streams by jointly learning affine motion and masks, using contrast maximization as the only training signal.

  12. SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM

    cs.CV 2025-11 conditional novelty 4.0 of 10

    SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.

  13. A Survey on Vision-Language-Action Models for Autonomous Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.

  14. Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A structured survey of 2023-2025 LLM and VLM methods for crash detection in video, with notable internal inconsistencies in reported numbers.

  15. Generative AI for Autonomous Driving: A Review

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A review of generative models (VAEs, GANs, diffusion, transformers, LLMs) applied to map generation, scenario generation, trajectory prediction, and motion planning for autonomous driving.

Pith tools