REVIEW 6 cited by
The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world, and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including region-text retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Models and the dataset shall be released at https://github.com/OpenGVLab/All-Seeing, and demo can be seen at https://huggingface.co/spaces/OpenGVLab/all-seeing.
Forward citations
Cited by 6 Pith papers
-
CoPa-SG: Dense Scene Graphs with Parametric and Proto-Relations
The paper introduces CoPa-SG, a synthetic scene graph dataset with more than 86 million relation annotations, plus parametric and proto-relations for richer scene representation.
-
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
A CLIP plus MLLM pipeline with additive-bias Low-Rank adaptation retrieves 3D objects of unseen categories from multi-view images, outperforming prior art by about 10% mAP on average.
-
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
DenseWorld-1M provides one million images with detailed object captions, pixel masks, and spatial relations by chaining SAM, APE, RAM++, and VLMs through a three-stage labeling pipeline.
-
CoMemo: LVLMs Need Image Context with Image Memory
CoMemo adds a cross-attention image-memory path and thumbnail-anchored position encoding to reduce visual neglect in long-context and multi-image LVLM tasks.
-
Automatic Fine-grained Segmentation-assisted Report Generation
ASaRG concatenates LVM-Med features and 212-class CXAS segmentation maps into LLaVA's projector, raising CE F1 by 2.77% over the baseline on MIMIC-CXR report generation.
-
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
PAM extends SAM 2 with a frozen LLM and a Semantic Perceiver to jointly segment and describe regions in images, videos, and streaming video, and contributes a 0.6M-sample region-level streaming video caption dataset.
Discussion (0). Sign in to comment.