REVIEW 15 cited by
BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving. Researchers are usually constrained to study a small set of problems on one dataset, while real-world computer vision applications require performing tasks of various complexities. We construct BDD100K, the largest driving video dataset with 100K videos and 10 tasks to evaluate the exciting progress of image recognition algorithms on autonomous driving. The dataset possesses geographic, environmental, and weather diversity, which is useful for training models that are less likely to be surprised by new conditions. Based on this diverse dataset, we build a benchmark for heterogeneous multitask learning and study how to solve the tasks together. Our experiments show that special training strategies are needed for existing models to perform such heterogeneous tasks. BDD100K opens the door for future studies in this important venue.
Forward citations
Cited by 15 Pith papers
-
DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification
A multi-vehicle naturalistic benchmark finds learned driving embeddings retain driver identity under condition matching, while descriptors collapse and video re-ID is mostly route leakage.
-
Vision as Unified Multimodal Generation
A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
-
Steadily moving semi-infinite fracture in plane poroelasticity
XEmbodied achieves SOTA on 18 embodied VQA benchmarks by fusing 3D geometric tokens and distilled physical cues into a 30B VLM with progressive curriculum training.
-
Benchmarking Nighttime Traffic Sign Recognition with Illumination-Adaptive Detection and Semantic Attribute Reasoning
INTSD, a 6,004-image Indian nighttime traffic-sign dataset, is introduced with LENS-Net, which reports 92.56 mAP@50 detection and 78.89 macro-precision classification, but the paper's internal statistics are inconsistent.
-
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
An IoU-weighted entropy objective and image-conditioned prompt selection adapt YOLO-World and Grounding DINO at test time, improving robustness on style, weather, low-light, and corruption shifts without labels.
-
RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment
A 2,100-video database with human opinions shows that current video quality models underperform on robot-generated content, motivating a new VQA subfield.
-
GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
GaRA-SAM improves SAM's robustness to image corruption by using input-dependent gating to adjust the effective rank of low-rank adapters, beating prior methods on robust segmentation benchmarks.
-
Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing
Transfer-learned SpikeYOLO achieves mAP 0.937/0.771 and HOTA 0.701/0.445 on KITTI and BDD100K for two-class automotive detection and tracking, competitive with conventional deep networks.
-
BlueGlass: A Framework for Composite AI Safety
BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...
-
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
CrowdTrack is a dense, first-person-view pedestrian tracking benchmark that exposes large performance drops in existing multi-object trackers.
-
EV-LayerSegNet: Self-supervised Motion Segmentation using Event Cameras
A self-supervised CNN segments moving objects in event camera streams by jointly learning affine motion and masks, using contrast maximization as the only training signal.
-
SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM
SAM-MI improves open-vocabulary segmentation by injecting aggregated SAM masks as low- and high-frequency guidance into CLIP cost maps, with sparse text-guided point prompts for speed.
-
A Survey on Vision-Language-Action Models for Autonomous Driving
A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.
-
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
A structured survey of 2023-2025 LLM and VLM methods for crash detection in video, with notable internal inconsistencies in reported numbers.
-
Generative AI for Autonomous Driving: A Review
A review of generative models (VAEs, GANs, diffusion, transformers, LLMs) applied to map generation, scenario generation, trajectory prediction, and motion planning for autonomous driving.
Discussion (0). Continue with ORCID to comment.