Pith. sign in

REVIEW 6 cited by

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01907 v1 pith:D3N6FQDN submitted 2023-08-03 cs.CV

classification cs.CV
keywords all-seeingdatasetmodelprojectrecognitionunderstandingworldbillion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world, and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including region-text retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Models and the dataset shall be released at https://github.com/OpenGVLab/All-Seeing, and demo can be seen at https://huggingface.co/spaces/OpenGVLab/all-seeing.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CoPa-SG: Dense Scene Graphs with Parametric and Proto-Relations

    cs.CV 2025-06 conditional novelty 7.0 of 10

    The paper introduces CoPa-SG, a synthetic scene graph dataset with more than 86 million relation annotations, plus parametric and proto-relations for richer scene representation.

  2. Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A CLIP plus MLLM pipeline with additive-bias Low-Rank adaptation retrieves 3D objects of unseen categories from multi-view images, outperforming prior art by about 10% mAP on average.

  3. DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DenseWorld-1M provides one million images with detailed object captions, pixel masks, and spatial relations by chaining SAM, APE, RAM++, and VLMs through a three-stage labeling pipeline.

  4. CoMemo: LVLMs Need Image Context with Image Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CoMemo adds a cross-attention image-memory path and thumbnail-anchored position encoding to reduce visual neglect in long-context and multi-image LVLM tasks.

  5. Automatic Fine-grained Segmentation-assisted Report Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    ASaRG concatenates LVM-Med features and 212-class CXAS segmentation maps into LLaVA's projector, raising CE F1 by 2.77% over the baseline on MIMIC-CXR report generation.

  6. Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

    cs.CV 2025-06 reject novelty 5.0 of 10

    PAM extends SAM 2 with a frozen LLM and a Semantic Perceiver to jointly segment and describe regions in images, videos, and streaming video, and contributes a 0.6M-sample region-level streaming video caption dataset.

Pith tools