Pith. sign in

REVIEW 4 cited by

AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.06679 v2 pith:2H5JRVLC submitted 2022-11-12 cs.CL

classification cs.CL
keywords clipencodermultilingualtaskstextcapabilitiesextendedlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by OpenAI, we altered its text encoder with a pre-trained multilingual text encoder XLM-R, and aligned both languages and image representations by a two-stage training schema consisting of teacher learning and contrastive learning. We validate our method through evaluations of a wide range of tasks. We set new state-of-the-art performances on a bunch of tasks including ImageNet-CN, Flicker30k-CN, COCO-CN and XTD. Further, we obtain very close performances with CLIP on almost all tasks, suggesting that one can simply alter the text encoder in CLIP for extended capabilities such as multilingual understanding. Our models and code are available at https://github.com/FlagAI-Open/FlagAI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    SimLabel improves zero-shot OOD detection by scoring images based on their consistency with a set of semantically similar class labels.

  2. MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A 20M-parameter adapter trained only on English image-text pairs adapts frozen diffusion models to generate images from prompts in over 110 languages, using a multilingual image-text encoder as its text encoder.

  3. CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities

    cs.CV 2025-08 reject novelty 5.0 of 10

    The paper claims a Chinese-prompt adapter for the Flux text-to-image model, but the manuscript body is an unrelated LLM self-recognition paper, leaving the claim unsupported.

  4. NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment

    cs.CV 2025-05 conditional novelty 4.0 of 10

    The NTIRE 2025 challenge report compares 20 methods for fine-grained text-to-image quality assessment, introduces the EvalMuse-Structure dataset, and finds every participating team outperformed the baselines.

Pith tools