Pith. sign in

REVIEW 6 cited by

Recognize Anything: A Strong Image Tagging Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03514 v3 pith:PHYZ54FA submitted 2023-06-06 cs.CV

classification cs.CV
keywords taggingmodelimagerecognizeannotationsanythingautomaticcomputer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the Recognize Anything Model (RAM): a strong foundation model for image tagging. RAM makes a substantial step for large models in computer vision, demonstrating the zero-shot ability to recognize any common category with high accuracy. RAM introduces a new paradigm for image tagging, leveraging large-scale image-text pairs for training instead of manual annotations. The development of RAM comprises four key steps. Firstly, annotation-free image tags are obtained at scale through automatic text semantic parsing. Subsequently, a preliminary model is trained for automatic annotation by unifying the caption and tagging tasks, supervised by the original texts and parsed tags, respectively. Thirdly, a data engine is employed to generate additional annotations and clean incorrect ones. Lastly, the model is retrained with the processed data and fine-tuned using a smaller but higher-quality dataset. We evaluate the tagging capabilities of RAM on numerous benchmarks and observe impressive zero-shot performance, significantly outperforming CLIP and BLIP. Remarkably, RAM even surpasses the fully supervised manners and exhibits competitive performance with the Google tagging API. We are releasing the RAM at \url{https://recognize-anything.github.io/} to foster the advancements of large models in computer vision.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Differentiable Parameter Propagation trains VLMs to emit design-parameter patches under hard layout and typography constraints, reaching 89% constraint satisfaction versus 52% for GPT-4V.

  2. ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ADAM combines LLM-generated contextual labels, CLIP embeddings, and nearest-neighbor voting to label novel objects without a predefined class list.

  3. Reinforced Visual Perception with Tools

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.

  4. What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A framework that generates image-level object concepts with a vision-language model before region segmentation improves open-vocabulary segmentation on multiple benchmarks.

  5. Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution

    cs.CV 2026-07 reject novelty 4.0 of 10

    DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...

  6. Open-Vocabulary Indoor Object Grounding with 3D Hierarchical Scene Graph

    cs.CV 2025-07 conditional novelty 4.0 of 10

    OVIGo-3DHSG builds a five-level scene graph (building, floor, room, location, object) and uses LLM reasoning over relevant subgraphs to ground open-vocabulary objects in multi-floor indoor scenes.

Pith tools