REVIEW 13 cited by
Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).
Forward citations
Cited by 13 Pith papers
-
Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs
MME-SID improves LLM-based sequential recommendation by fusing collaborative, text, and image embeddings with quantized semantic IDs, using MMD reconstruction and code-embedding initialization.
-
L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
A distilled 99M-parameter CLIP gives L-CLIPScore, a lightweight caption metric that matches CLIPScore on human correlation and best improves captioning models when mixed with CIDEr.
-
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
The authors propose REIR (retrieval plus localization of object instances from text queries), construct the REIRCOCO benchmark, and report a dual-stream contrastive baseline, CLARE, that outperforms two-stage TIR+REC ...
-
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
ALTA adapts a frozen masked-pretrained X-ray encoder to language with 8% trainable parameters and temporal-multiview inputs, improving medical retrieval and zero-shot classification.
-
FREE: Fast and Robust Vision Language Models with Early Exits
An adversarial early-exit method for frozen-backbone vision language models that reuses the final classifier and reports 1.5x inference speedup with comparable accuracy.
-
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
A monolithic multimodal LLM that cuts pre-training data by 58% and first-token latency by up to 69% while matching or beating its predecessor on 15 benchmarks.
-
LunarFM: A Shared Multimodal Representation of the Moon's Surface
A self-supervised multimodal model fuses 18 channels from six lunar instruments into a shared 768-dimensional embedding per 0.5° chip, enabling mineral regression, similarity search, and geological-unit classification...
-
Qwen-Audio-VAE Technical Report
A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.
-
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
MoE-TTS adds frozen text-expert MoE modules to a Qwen3-based TTS system and reports better out-of-domain description alignment than ElevenLabs and MiniMax on a small hand-built test set.
-
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
MoMa adapts frozen CLIP to video by injecting Mamba-computed scale and bias into each layer, improving accuracy and efficiency on multiple action recognition benchmarks.
-
Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
Manager aggregates multi-layer unimodal representations and improves both two-tower VLMs (ManagerTower) and MLLMs (LLaVA-OV-Manager) on 24 downstream tasks.
-
Two-flow Feedback Multi-scale Progressive Generative Adversarial Network
A GAN paper that proposes several new modules but reports no actual experimental results, with placeholder dataset names and percentages.
-
Dynamic Double Space Tower
The paper claims a four-layer Gestalt-based tower can replace attention in VQA and lift a 3B model to state-of-the-art spatial reasoning, but provides no reproducible method or consistent results.
Discussion (0). Sign in to comment.