Pith. sign in

REVIEW 5 cited by

Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.00849 v2 pith:CYDYX5C6 submitted 2020-04-02 cs.CV cs.CLcs.LGcs.MM

classification cs.CVcs.CLcs.LGcs.MM
keywords languagevisualimagetasksmodelpixel-bertpixelsrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose Pixel-BERT to align image pixels with text by deep multi-modal transformers that jointly learn visual and language embedding in a unified end-to-end framework. We aim to build a more accurate and thorough connection between image pixels and language semantics directly from image and sentence pairs instead of using region-based image features as the most recent vision and language tasks. Our Pixel-BERT which aligns semantic connection in pixel and text level solves the limitation of task-specific visual representation for vision and language tasks. It also relieves the cost of bounding box annotations and overcomes the unbalance between semantic labels in visual task and language semantic. To provide a better representation for down-stream tasks, we pre-train a universal end-to-end model with image and sentence pairs from Visual Genome dataset and MS-COCO dataset. We propose to use a random pixel sampling mechanism to enhance the robustness of visual representation and to apply the Masked Language Model and Image-Text Matching as pre-training tasks. Extensive experiments on downstream tasks with our pre-trained model show that our approach makes the most state-of-the-arts in downstream tasks, including Visual Question Answering (VQA), image-text retrieval, Natural Language for Visual Reasoning for Real (NLVR). Particularly, we boost the performance of a single model in VQA task by 2.17 points compared with SOTA under fair comparison.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving vision-language alignment with graph spiking hybrid Networks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A vision-language model that encodes panoptic image segments with graph attention and spiking neurons reports competitive or better results on VQA, VE, NLVR2, and retrieval benchmarks.

  2. Kinky vortons in the 2HDM

    hep-ph 2026-03 conditional novelty 5.0 of 10

    Multiple dynamically stable kinky vortons exist in the Z2-symmetric global 2HDM and are accurately described by thin-string and elastic-string approximations.

  3. Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Manager aggregates multi-layer unimodal representations and improves both two-tower VLMs (ManagerTower) and MLLMs (LLaVA-OV-Manager) on 24 downstream tasks.

  4. RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

    cs.LG 2025-06 conditional novelty 5.0 of 10

    RollingQ rotates the classification query in a multimodal Transformer toward a rebalanced direction so attention stops over-favoring a single modality, restoring dynamic fusion and improving accuracy.

  5. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.

Pith tools