REVIEW 7 cited by
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt the material, and they have been collected from Flickr without a predefined list of class names or tags, leading to natural class statistics and avoiding an initial design bias. Open Images V4 offers large scale across several dimensions: 30.1M image-level labels for 19.8k concepts, 15.4M bounding boxes for 600 object classes, and 375k visual relationship annotations involving 57 classes. For object detection in particular, we provide 15x more bounding boxes than the next largest datasets (15.4M boxes on 1.9M images). The images often show complex scenes with several objects (8 annotated objects per image on average). We annotated visual relationships between them, which support visual relationship detection, an emerging task that requires structured reasoning. We provide in-depth comprehensive statistics about the dataset, we validate the quality of the annotations, we study how the performance of several modern models evolves with increasing amounts of training data, and we demonstrate two applications made possible by having unified annotations of multiple types coexisting in the same images. We hope that the scale, quality, and variety of Open Images V4 will foster further research and innovation even beyond the areas of image classification, object detection, and visual relationship detection.
Forward citations
Cited by 7 Pith papers
-
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
SpatialCLI shows that a VLM can learn to use localization, segmentation, depth, and pose tools and then internalize the tool outputs into direct reasoning, improving both tool-enabled and tool-free spatial task performance.
-
TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving
TaCarla releases 2.85M CARLA Leaderboard 2.0 frames with nuScenes-style sensors, multi-task annotations, planning baselines, and a text-based rarity score.
-
Disentangling 3D Modeling from Spatial Reasoning
DiSR answers spatial questions by feeding a language model a text summary of metric 3D positions, sizes, and orientations from frozen expert models, reaching top scores on two benchmarks with 59 GPU-hours of training.
-
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
An encoder-free vision-language model using separate attention, normalization, and feed-forward weights for image versus text tokens outperforms earlier encoder-free models and narrows the gap to encoder-based VLMs wi...
-
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
ChartSketcher has a multimodal LLM sketch intermediate reasoning steps directly on chart images and feed those sketches back as visual feedback, improving chart QA accuracy over its base model.
-
SHeRL-FL: When Representation Learning Meets Split Learning in Hierarchical Federated Learning
The submitted body is an unrelated survey, not the SHeRL-FL method claimed in the metadata.
-
Object Recognition Datasets and Challenges: A Review
A review paper that compiles statistics and descriptions of over 160 object recognition datasets, their associated challenges, and evaluation metrics.
Discussion (0). Continue with ORCID to comment.