Pith. sign in

REVIEW 15 cited by

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.10904 v3 pith:LGKCLSOF submitted 2021-08-24 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords pretrainingsimvlmvisualincludinglanguagemodelaccuracyimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream tasks. However, the requirement for expensive annotations including clean image captions and regional labels limits the scalability of existing approaches, and complicates the pretraining procedure with the introduction of multiple dataset-specific objectives. In this work, we relax these constraints and present a minimalist pretraining framework, named Simple Visual Language Model (SimVLM). Unlike prior work, SimVLM reduces the training complexity by exploiting large-scale weak supervision, and is trained end-to-end with a single prefix language modeling objective. Without utilizing extra data or task-specific customization, the resulting model significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA (+3.74% vqa-score), NLVR2 (+1.17% accuracy), SNLI-VE (+1.37% accuracy) and image captioning tasks (+10.1% average CIDEr score). Furthermore, we demonstrate that SimVLM acquires strong generalization and transfer ability, enabling zero-shot behavior including open-ended visual question answering and cross-modality transfer.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SensorLM: Learning the Language of Wearable Sensors

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SensorLM is a sensor-language foundation model trained on 59.7M hours of wearable data with template-generated captions, reporting strong zero-shot, few-shot, and retrieval performance.

  2. FREE: Fast and Robust Vision Language Models with Early Exits

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An adversarial early-exit method for frozen-backbone vision language models that reuses the final classifier and reports 1.5x inference speedup with comparable accuracy.

  3. Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A progressive curriculum that trains open-vocabulary detectors on low-ambiguity, high-signal cross-modal alignments first improves robustness to visual domain shifts, with modest, test-tuned gains.

  4. Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A unified detector with category-aware mixture-of-experts and attribution chain-of-thought reaches 86.7% accuracy on a new combined human-crafted + AI-generated misinformation benchmark.

  5. Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Randomly masking image patches when generating conditional prompts improves unseen-class accuracy for CoCoOp-style CLIP prompt learning by about 1 to 2 points on average, with negligible extra cost.

  6. Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An image autoencoder compresses pictures into caption-like embeddings via a frozen diffusion decoder, and a fine-tuned LLM reads those embeddings into captions claimed to rival GPT-4o at under $1,000 training cost.

  7. From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.

  8. Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BYB is a self-supervised pretraining strategy that predicts a pooled future behavior embedding from a momentum teacher, removing the need for hand-built behavior vocabularies and improving downstream AUC and throughpu...

  9. CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization

    cs.CV 2025-06 conditional novelty 4.0 of 10

    CoCoA-Mix combines cross-entropy with a confidence penalty and a scalar-weighted mixture of specialized and generalized prompts, beating prior prompt-tuning baselines on base-to-new, cross-dataset, and few-shot increm...

  10. Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval

    cs.CV 2025-05 conditional novelty 4.0 of 10

    RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.

  11. Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A structured survey of 2023-2025 LLM and VLM methods for crash detection in video, with notable internal inconsistencies in reported numbers.

  12. Vision Generalist Model: A Survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.

  13. Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model

    cs.CV 2025-05 conditional novelty 3.0 of 10

    Applying beam search, patch self-attention, and cosine scheduling to the K-Replay captioning framework improves knowledge-keyword recognition on KnowCap, though the full combined model is not reported.

  14. Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.

  15. Foundation Model Driven Robotics: A Comprehensive Review

    cs.RO 2025-07 conditional novelty 2.0 of 10

    A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.

Pith tools