Pith. sign in

REVIEW 1 cited by

Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.11154 v2 pith:NCFUDZZK submitted 2023-02-22 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords modelsvisualentitiesrecognitionwikipediaentityexistingimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale multi-modal pre-training models such as CLIP and PaLI exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct OVEN-Wiki by re-purposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. OVEN challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-of-the-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new benchmark integrates AI2-THOR, Google Street View, and functional websites to test agents that must combine physical actions with online information retrieval.

Pith tools