Pith. sign in

REVIEW 7 cited by

Tag2Text: Guiding Vision-Language Model via Image Tagging

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.05657 v3 pith:DDBUP7C4 submitted 2023-03-10 cs.CV

classification cs.CV
keywords tag2textimagemodelstaggingvision-languageguidancemodelperformance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with an off-the-shelf detector with limited performance, our approach explicitly learns an image tagger using tags parsed from image-paired text and thus provides a strong semantic guidance to vision-language models. In this way, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. As a result, Tag2Text demonstrates the ability of a foundational image tagging model, with superior zero-shot performance even comparable to fully supervised models. Moreover, by leveraging the tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance. Code, demo and pre-trained models are available at https://github.com/xinyu1205/recognize-anything.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Fine-tuning LVLMs on original/altered image pairs with targeted visual instructions reduces object and attribute hallucinations on POPE, LLaVA-Bench, and MMHal-Bench.

  2. Detailed Object Description with Controllable Dimensions

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A training-free post-processing pipeline improves how well multimodal LLMs stick to user-selected object dimensions such as color, texture, and pose.

  3. Object Style Diffusion for Generalized Object Detection in Urban Scene

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GoDiff improves object detection generalization across weather by generating pseudo-target images with a diffusion model and mixing style statistics during training, but the reported gains are modest and depend on kno...

  4. Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios

    cs.CV 2024-12 reject novelty 5.0 of 10

    HCOENet combines multi-model entity cross-checking with open-set detection to delete hallucinated objects and add descriptions of overlooked traffic objects, with reported F1 gains on a modified POPE evaluation.

  5. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

  6. Seamless and Efficient Interactions within a Mixed-Dimensional Information Space

    cs.HC 2025-06 conditional novelty 4.0 of 10

    A thesis that three design strategies, multimodal AI, context-aware placement, and combined 2D/3D views, make mixed-dimensional information spaces seamless and efficient, demonstrated with three systems.

  7. Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models

    cs.CV 2024-11 reject novelty 4.0 of 10

    A Mahalanobis distance computed in SAM feature space is used as an aleatoric uncertainty score for object instances, and filtering and reweighting by this score yields modest AP gains on COCO and BDD100K.

Pith tools