Pith. sign in

REVIEW 28 cited by

LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.07490 v3 pith:FPIGQHMJ submitted 2019-08-20 cs.CL cs.CVcs.LG

classification cs.CLcs.CVcs.LG
keywords cross-modalitymodelencoderlanguagelxmertlearningpre-trainedanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-and-language reasoning requires an understanding of visual concepts, language semantics, and, most importantly, the alignment and relationships between these two modalities. We thus propose the LXMERT (Learning Cross-Modality Encoder Representations from Transformers) framework to learn these vision-and-language connections. In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder. Next, to endow our model with the capability of connecting vision and language semantics, we pre-train the model with large amounts of image-and-sentence pairs, via five diverse representative pre-training tasks: masked language modeling, masked object prediction (feature regression and label classification), cross-modality matching, and image question answering. These tasks help in learning both intra-modality and cross-modality relationships. After fine-tuning from our pre-trained parameters, our model achieves the state-of-the-art results on two visual question answering datasets (i.e., VQA and GQA). We also show the generalizability of our pre-trained cross-modality model by adapting it to a challenging visual-reasoning task, NLVR2, and improve the previous best result by 22% absolute (54% to 76%). Lastly, we demonstrate detailed ablation studies to prove that both our novel model components and pre-training strategies significantly contribute to our strong results; and also present several attention visualizations for the different encoders. Code and pre-trained models publicly available at: https://github.com/airsplay/lxmert

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Monitoring attention entropy and image-output correlation during decoding, then applying targeted contrastive corrections, reduces hallucination in multimodal LLMs without retraining.

  2. XRFormer: Multiscale Tokenization for XRF Representation Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A multiscale convolutional tokenizer plus MSM/PPP pretraining yields more accurate, parameter-efficient transformers for XRF pigment identification and unmixing than ViT, SpectralFormer, or 1D-CNN baselines.

  3. Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A dual attention adapter using support-image memory and local-global feature mixing improves CLIP few-shot and domain-shift classification.

  4. SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.

  5. DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    A single diffusion policy trained with DAgger, without a waypoint predictor, reports better performance than two-stage waypoint-based models on VLN-CE benchmarks.

  6. Gait-Based Hand Load Estimation via Deep Latent Variable Models with Auxiliary Information

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A VAE-TCN model with bidirectional cross-attention that uses unloaded baseline gait and marginalizes over carrying style cuts hand-load estimation MAE to 5.67 lb on a 22-person IMU dataset.

  7. NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments

    cs.CV 2025-06 conditional novelty 6.0 of 10

    NavMorph combines an RSSM-based latent world model with an online-updated contextual memory, reporting consistent VLN-CE gains on R2R-CE and RxR-CE.

  8. GenRecal: Generation after Recalibration from Large to Small Vision-Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A learnable Recalibrator bridges different tokenizers so that small VLMs can distill knowledge from any large VLM, improving their benchmark scores.

  9. Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LightD creates natural adversarial relighting images with GPT-selected lighting parameters and gradient optimization, outperforming prior non-suspicious attacks on vision-language models.

  10. R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Share-GRPO creates paraphrased and visually augmented versions of reasoning questions and shares answers and reward signals across versions, improving multimodal reasoning without cold-start SFT.

  11. Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval

    cs.CV 2026-01 conditional novelty 5.0 of 10

    By generating an edited "mental image" of a query and synthetic counterparts of database images, and matching in that synthetic space, Paracosm achieves state-of-the-art training-free zero-shot composed image retrieva...

  12. MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models

    cs.AI 2025-12 conditional novelty 5.0 of 10

    MIND improves multimodal reasoning by training on diverse correct and deliberately wrong rationales with two-stage correction and contrastive alignment, reporting SOTA on ScienceQA, A-OKVQA, and M3CoT.

  13. Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    A systematic review of 55 papers finds explainability for multimodal attention-based models is dominated by attention-weight visualizations, while evaluation remains mostly qualitative and non-standardized.

  14. Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Psychology-guided LLM text embeddings fused with audio and facial cues achieved the lowest MSE in the AVI 2025 personality assessment challenge.

  15. MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    MM-Prompt couples the visual and language prompt paths in continual VQA, and reports higher average accuracy and lower forgetting than existing prompt-based methods.

  16. Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis

    cs.CV 2026-08 reject novelty 4.0 of 10

    An unsupervised 3D root skeleton extractor plus evidence-first GPT-4o fine-tuning is claimed to improve root phenotyping VQA accuracy on a private 12-species dataset.

  17. Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

    cs.CV 2026-07 reject novelty 4.0 of 10

    A CP-decomposed, instruction-conditioned latent bottleneck (CompactNav) improves VLN-CE success rate by about 2% over prior state of the art on two benchmarks.

  18. A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications

    eess.IV 2026-01 unverdicted novelty 4.0 of 10

    A survey that classifies visual semantic communication into preservation, expansion, and refinement categories and reviews their machine-learning components and applications.

  19. EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A Qwen-based four-stage retrieval pipeline with RRF ensembling achieves the top-1 score on the EVENTA 2025 Track 2 private test set.

  20. Analyzing the Sensitivity of Vision Language Models in Visual Question Answering

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Adding answer-preserving visual or relational modifiers to visual questions lowers accuracy of GPT-4o, Gemini-1.5-Flash, and Claude-3.5-Sonnet on VQA v2.0.

  21. Can Mental Imagery Improve the Thinking Capabilities of AI Systems?

    cs.LG 2025-07 reject novelty 4.0 of 10

    A framework for machine thinking that adds a Mental Imagery Unit is described, but its demonstrations do not test whether imagery improves reasoning.

  22. Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A new multimodal fusion architecture reports small state-of-the-art gains on two offensive content benchmarks using co-attention, dimension-wise gating, and expert fusion.

  23. Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A meta-learning dissertation showing that distributed memory and hypernetworks can adapt to new tasks with few samples, applied to image classification, text-to-3D generation, and molecular binding prediction, with th...

  24. Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis

    cs.CV 2025-05 reject novelty 3.0 of 10

    A duration-based policy table selects between thresholding and fixed-interval splitting for scene detection, and a sharpness-plus-brightness score picks one keyframe per scene.

  25. A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A survey that categorizes positive and negative pair curation techniques in visual contrastive learning and discusses their trade-offs and open questions.

  26. BERT-VQA: Visual Question Answering on Plots

    cs.LG 2025-08 reject novelty 2.0 of 10

    A VisualBERT-based VQA model underperformed a simple LSTM+CNN+classifier baseline on a subset of PlotQA yes/no questions, but the comparison does not isolate the fusion mechanism.

  27. Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.

  28. Can Argus Judge Them All? Comparing VLMs Across Domains

    cs.IR 2025-06 reject novelty 2.0 of 10

    A VLM benchmark paper that proposes a cross-dataset consistency metric, but the abstract and body evaluate different model sets and the metric's bounds are incorrect.

Pith tools