Pith. sign in

REVIEW 13 cited by

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.10442 v2 pith:ZYTAPYSC submitted 2022-08-22 cs.CV cs.CL

classification cs.CVcs.CL
keywords cocoimagelanguagepretrainingvisionarchitecturebackbonebeit-3
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 151 citations worldwide. Full citation record

  1. Empowering Large Language Model for Sequential Recommendation via Multimodal Embeddings and Semantic IDs

    cs.IR 2025-09 conditional novelty 6.0 of 10

    MME-SID improves LLM-based sequential recommendation by fusing collaborative, text, and image embeddings with quantized semantic IDs, using MMD reconstruction and code-embedding initialization.

  2. L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A distilled 99M-parameter CLIP gives L-CLIPScore, a lightweight caption metric that matches CLIPScore on human correlation and best improves captioning models when mixed with CIDEr.

  3. Referring Expression Instance Retrieval and A Strong End-to-End Baseline

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors propose REIR (retrieval plus localization of object instances from text queries), construct the REIRCOCO benchmark, and report a dual-stream contrastive baseline, CLARE, that outperforms two-stage TIR+REC ...

  4. Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ALTA adapts a frozen masked-pretrained X-ray encoder to language with 8% trainable parameters and temporal-multiview inputs, improving medical retrieval and zero-shot classification.

  5. FREE: Fast and Robust Vision Language Models with Early Exits

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An adversarial early-exit method for frozen-backbone vision language models that reuses the final classifier and reports 1.5x inference speedup with comparable accuracy.

  6. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.5 of 10

    A monolithic multimodal LLM that cuts pre-training data by 58% and first-token latency by up to 69% while matching or beating its predecessor on 15 benchmarks.

  7. LunarFM: A Shared Multimodal Representation of the Moon's Surface

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A self-supervised multimodal model fuses 18 channels from six lunar instruments into a shared 768-dimensional embedding per 0.5° chip, enabling mineral regression, similarity search, and geological-unit classification...

  8. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  9. MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

    eess.AS 2025-08 conditional novelty 5.0 of 10

    MoE-TTS adds frozen text-expert MoE modules to a Qwen3-based TTS system and reports better out-of-domain description alignment than ElevenLabs and MiniMax on a small hand-built test set.

  10. MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MoMa adapts frozen CLIP to video by injecting Mamba-computed scale and bias into each layer, improving accuracy and efficiency on multiple action recognition benchmarks.

  11. Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Manager aggregates multi-layer unimodal representations and improves both two-tower VLMs (ManagerTower) and MLLMs (LLaVA-OV-Manager) on 24 downstream tasks.

  12. Two-flow Feedback Multi-scale Progressive Generative Adversarial Network

    cs.CV 2025-08 reject novelty 3.0 of 10

    A GAN paper that proposes several new modules but reports no actual experimental results, with placeholder dataset names and percentages.

  13. Dynamic Double Space Tower

    cs.CV 2025-06 reject novelty 3.0 of 10

    The paper claims a four-layer Gestalt-based tower can replace attention in VQA and lift a 3B model to state-of-the-art spatial reasoning, but provides no reproducible method or consistent results.

Pith tools