Pith. sign in

REVIEW 12 cited by

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.01335 v3 pith:V4JKKJPJ submitted 2022-11-02 cs.CV cs.CL

classification cs.CVcs.CL
keywords chineseclipachievemodelsperformancepretrainingcontrastivedataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pairs in Chinese, where most data are retrieved from publicly available datasets, and we pretrain Chinese CLIP models on the new dataset. We develop 5 Chinese CLIP models of multiple sizes, spanning from 77 to 958 million parameters. Furthermore, we propose a two-stage pretraining method, where the model is first trained with the image encoder frozen and then trained with all parameters being optimized, to achieve enhanced model performance. Our comprehensive experiments demonstrate that Chinese CLIP can achieve the state-of-the-art performance on MUGE, Flickr30K-CN, and COCO-CN in the setups of zero-shot learning and finetuning, and it is able to achieve competitive performance in zero-shot image classification based on the evaluation on the ELEVATER benchmark (Li et al., 2022). We have released our codes, models, and demos in https://github.com/OFA-Sys/Chinese-CLIP

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Illuminating Visual Identity in Universal Multimodal Embeddings

    cs.CV 2026-08 conditional novelty 6.0 of 10

    By adding identity-aware sampling and a contrastive loss on a new 28-dataset benchmark, the authors build multimodal embeddings that are far better at visual identity matching without losing general retrieval accuracy.

  2. GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A three-stage multimodal recommender pipeline with GRPO-based behavior alignment and adaptive ID-content fusion claims a 0.55% online order-volume increase and small offline AUC gains at Taobao Shangou.

  3. RecGPT-V3 Technical Report

    cs.IR 2026-07 conditional novelty 6.0 of 10

    A stateful LLM recommender with memory, text-plus-Semantic-ID grounding, and latent reasoning reports higher Taobao engagement and sales at ~52% lower serving compute than its predecessor.

  4. Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data

    cs.LG 2026-03 unverdicted novelty 6.0 of 10

    MIPO constructs contrastive preference pairs from correct versus random prompts and uses DPO to maximize mutual information between prompts and responses, producing 3-40% gains on personalization and 1-18% on math tas...

  5. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

    cs.CV 2025-11 unverdicted novelty 6.0 of 10

    Z-Image is an efficient 6B-parameter foundation model for image generation that rivals larger commercial systems in photorealism and bilingual text rendering through a new single-stream diffusion transformer and strea...

  6. FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A two-stage bilingual CLIP-style model with region-text supervision and a new text-side contrastive loss outperforms prior open models on fine-grained vision-language tasks in English and Chinese.

  7. Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A full-image, mask-guided CLIP model with two-stage multi-granularity alignment training sets a new state of the art for Chinese scene text retrieval and introduces a diverse-layout benchmark.

  8. HarmonyCut: Supporting Creative Chinese Paper-cutting Design with Form and Connotation Harmony

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A generative-AI tool with a paper-cutting-specific design space (four factors and a pattern taxonomy) improves designers' exploration, editing, and creativity support compared with generic generative AI.

  9. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  10. JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    JuZhou 1.0 is a 0.387B-parameter T2I diffusion model with 4-step inference achieving 0.69 GenEval, trained on 9M Chinese pairs using Sugon K100 accelerators and deployable on Android/iOS devices.

  11. CHaystack: Benchmarking Chinese Document Retrieval and VQA

    cs.IR 2026-05 conditional novelty 5.0 of 10

    CHaystack, a Chinese document retrieval-and-VQA benchmark, shows Qwen3-VL reaching 71.91 Recall@1 versus 14.40 for the best non-Qwen model, and its VLM filter improves recall.

  12. FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets

    cs.IR 2025-09 conditional novelty 5.0 of 10

    FORGE shows that balancing codebook usage and adding multimodal side information improves semantic identifiers for generative retrieval, validated offline and on Taobao.

Pith tools