REVIEW 9 cited by
Long-CLIP: Unlocking the Long-Text Capability of CLIP
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contrastive Language-Image Pre-training (CLIP) has been the cornerstone for zero-shot classification, text-image retrieval, and text-image generation by aligning image and text modalities. Despite its widespread adoption, a significant limitation of CLIP lies in the inadequate length of text input. The length of the text token is restricted to 77, and an empirical study shows the actual effective length is even less than 20. This prevents CLIP from handling detailed descriptions, limiting its applications for image retrieval and text-to-image generation with extensive prerequisites. To this end, we propose Long-CLIP as a plug-and-play alternative to CLIP that supports long-text input, retains or even surpasses its zero-shot generalizability, and aligns the CLIP latent space, making it readily replace CLIP without any further adaptation in downstream frameworks. Nevertheless, achieving this goal is far from straightforward, as simplistic fine-tuning can result in a significant degradation of CLIP's performance. Moreover, substituting the text encoder with a language model supporting longer contexts necessitates pretraining with vast amounts of data, incurring significant expenses. Accordingly, Long-CLIP introduces an efficient fine-tuning solution on CLIP with two novel strategies designed to maintain the original capabilities, including (1) a knowledge-preserved stretching of positional embedding and (2) a primary component matching of CLIP features. With leveraging just one million extra long text-image pairs, Long-CLIP has shown the superiority to CLIP for about 20% in long caption text-image retrieval and 6% in traditional text-image retrieval tasks, e.g., COCO and Flickr30k. Furthermore, Long-CLIP offers enhanced capabilities for generating images from detailed text descriptions by replacing CLIP in a plug-and-play manner.
Forward citations
Cited by 9 Pith papers
-
AnyStyle: Single-Pass Multimodal Stylization for 3D Gaussian Splatting
A single-pass 3D Gaussian splatting pipeline that stylizes unposed scenes from either a text prompt or a reference image via a lightweight zero-initialized style-injection branch.
-
PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement
PoemTale Diffusion generates a coherent set of images from a poem by combining emotion-based segmentation, multi-stage LLM prompt refinement, and consistent self-attention, outperforming direct poem-to-image approache...
-
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
A dual-branch CLIP training pipeline with regional prompts and hierarchical feature alignment reaches state-of-the-art on long- and short-text retrieval.
-
GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
A hierarchical multi-agent framework with GPT-4o generates escape room puzzle images that are judged more solvable and less shortcut-prone than vanilla text-to-image outputs.
-
JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
JuZhou 1.0 is a 0.387B-parameter T2I diffusion model with 4-step inference achieving 0.69 GenEval, trained on 9M Chinese pairs using Sugon K100 accelerators and deployable on Android/iOS devices.
-
Domain-Enhanced Dual-Branch Model for Efficient and Interpretable Accident Anticipation
The paper claims 87.7% AP on DAD by aligning VGG-16 video features with Long-CLIP embeddings of California DMV accident reports and GPT-generated safe-driving reports.
-
ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model
ANT makes text embeddings change across denoising steps and schedules classifier-free guidance to decay, improving text-motion alignment in diffusion text-to-motion models.
-
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
MultiFakeVerse provides 845,286 person-centric images edited through VLM-generated instructions; state-of-the-art deepfake detectors and human observers misclassify a large fraction of them.
-
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
DCLIP fine-tunes a CLIP student's image encoder to match a YOLO-region, bidirectional cross-attention teacher, improving retrieval while retaining most zero-shot accuracy.
Discussion (0). Continue with ORCID to comment.