REVIEW 4 cited by
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by OpenAI, we altered its text encoder with a pre-trained multilingual text encoder XLM-R, and aligned both languages and image representations by a two-stage training schema consisting of teacher learning and contrastive learning. We validate our method through evaluations of a wide range of tasks. We set new state-of-the-art performances on a bunch of tasks including ImageNet-CN, Flicker30k-CN, COCO-CN and XTD. Further, we obtain very close performances with CLIP on almost all tasks, suggesting that one can simply alter the text encoder in CLIP for extended capabilities such as multilingual understanding. Our models and code are available at https://github.com/FlagAI-Open/FlagAI.
Forward citations
Cited by 4 Pith papers
-
SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
SimLabel improves zero-shot OOD detection by scoring images based on their consistency with a set of semantically similar class labels.
-
MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
A 20M-parameter adapter trained only on English image-text pairs adapts frozen diffusion models to generate images from prompts in over 110 languages, using a multilingual image-text encoder as its text encoder.
-
CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities
The paper claims a Chinese-prompt adapter for the Flux text-to-image model, but the manuscript body is an unrelated LLM self-recognition paper, leaving the claim unsupported.
-
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
The NTIRE 2025 challenge report compares 20 methods for fine-grained text-to-image quality assessment, introduces the EvalMuse-Structure dataset, and finds every participating team outperformed the baselines.
Discussion (0). Continue with ORCID to comment.