REVIEW 30 cited by
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-and-language reasoning requires an understanding of visual concepts, language semantics, and, most importantly, the alignment and relationships between these two modalities. We thus propose the LXMERT (Learning Cross-Modality Encoder Representations from Transformers) framework to learn these vision-and-language connections. In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder. Next, to endow our model with the capability of connecting vision and language semantics, we pre-train the model with large amounts of image-and-sentence pairs, via five diverse representative pre-training tasks: masked language modeling, masked object prediction (feature regression and label classification), cross-modality matching, and image question answering. These tasks help in learning both intra-modality and cross-modality relationships. After fine-tuning from our pre-trained parameters, our model achieves the state-of-the-art results on two visual question answering datasets (i.e., VQA and GQA). We also show the generalizability of our pre-trained cross-modality model by adapting it to a challenging visual-reasoning task, NLVR2, and improve the previous best result by 22% absolute (54% to 76%). Lastly, we demonstrate detailed ablation studies to prove that both our novel model components and pre-training strategies significantly contribute to our strong results; and also present several attention visualizations for the different encoders. Code and pre-trained models publicly available at: https://github.com/airsplay/lxmert
Forward citations
Cited by 30 Pith papers
-
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models
Monitoring attention entropy and image-output correlation during decoding, then applying targeted contrastive corrections, reduces hallucination in multimodal LLMs without retraining.
-
XRFormer: Multiscale Tokenization for XRF Representation Learning
A multiscale convolutional tokenizer plus MSM/PPP pretraining yields more accurate, parameter-efficient transformers for XRF pigment identification and unmixing than ViT, SpectralFormer, or 1D-CNN baselines.
-
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
A dual attention adapter using support-image memory and local-global feature mixing improves CLIP few-shot and domain-shift classification.
-
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.
-
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
A single diffusion policy trained with DAgger, without a waypoint predictor, reports better performance than two-stage waypoint-based models on VLN-CE benchmarks.
-
Gait-Based Hand Load Estimation via Deep Latent Variable Models with Auxiliary Information
A VAE-TCN model with bidirectional cross-attention that uses unloaded baseline gait and marginalizes over carrying style cuts hand-load estimation MAE to 5.67 lb on a 22-person IMU dataset.
-
NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
NavMorph combines an RSSM-based latent world model with an online-updated contextual memory, reporting consistent VLN-CE gains on R2R-CE and RxR-CE.
-
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
A learnable Recalibrator bridges different tokenizers so that small VLMs can distill knowledge from any large VLM, improving their benchmark scores.
-
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models
LightD creates natural adversarial relighting images with GPT-selected lighting parameters and gradient optimization, outperforming prior non-suspicious attacks on vision-language models.
-
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
Share-GRPO creates paraphrased and visually augmented versions of reasoning questions and shares answers and reward signals across versions, improving multimodal reasoning without cold-start SFT.
-
A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions
A multimodal transformer predicts ODE/PDE solutions and generates correct scientific text descriptions from numerical and symbolic inputs, with low error on in-distribution and out-of-distribution tests.
-
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
By generating an edited "mental image" of a query and synthetic counterparts of database images, and matching in that synthetic space, Paracosm achieves state-of-the-art training-free zero-shot composed image retrieva...
-
MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
MIND improves multimodal reasoning by training on diverse correct and deliberately wrong rationales with two-stage correction and contrastive alignment, reporting SOTA on ScienceQA, A-OKVQA, and M3CoT.
-
Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
A systematic review of 55 papers finds explainability for multimodal attention-based models is dominated by attention-weight visualizations, while evaluation remains mostly qualitative and non-standardized.
-
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
Psychology-guided LLM text embeddings fused with audio and facial cues achieved the lowest MSE in the AVI 2025 personality assessment challenge.
-
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
MM-Prompt couples the visual and language prompt paths in continual VQA, and reports higher average accuracy and lower forgetting than existing prompt-based methods.
-
Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis
An unsupervised 3D root skeleton extractor plus evidence-first GPT-4o fine-tuning is claimed to improve root phenotyping VQA accuracy on a private 12-species dataset.
-
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
A CP-decomposed, instruction-conditioned latent bottleneck (CompactNav) improves VLN-CE success rate by about 2% over prior state of the art on two benchmarks.
-
A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications
A survey that classifies visual semantic communication into preservation, expansion, and refinement categories and reviews their machine-learning components and applications.
-
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
A Qwen-based four-stage retrieval pipeline with RRF ensembling achieves the top-1 score on the EVENTA 2025 Track 2 private test set.
-
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
Adding answer-preserving visual or relational modifiers to visual questions lowers accuracy of GPT-4o, Gemini-1.5-Flash, and Claude-3.5-Sonnet on VQA v2.0.
-
Can Mental Imagery Improve the Thinking Capabilities of AI Systems?
A framework for machine thinking that adds a Mental Imagery Unit is described, but its demonstrations do not test whether imagery improves reasoning.
-
Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection
A new multimodal fusion architecture reports small state-of-the-art gains on two offensive content benchmarks using co-attention, dimension-wise gating, and expert fusion.
-
Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures
A meta-learning dissertation showing that distributed memory and hypernetworks can adapt to new tasks with few samples, applied to image classification, text-to-3D generation, and molecular binding prediction, with th...
-
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
A duration-based policy table selects between thresholding and fixed-interval splitting for scene detection, and a sharpness-plus-brightness score picks one keyframe per scene.
-
A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters
A survey that categorizes positive and negative pair curation techniques in visual contrastive learning and discusses their trade-offs and open questions.
-
BERT-VQA: Visual Question Answering on Plots
A VisualBERT-based VQA model underperformed a simple LSTM+CNN+classifier baseline on a subset of PlotQA yes/no questions, but the comparison does not isolate the fusion mechanism.
-
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.
-
Can Argus Judge Them All? Comparing VLMs Across Domains
A VLM benchmark paper that proposes a cross-dataset consistency metric, but the abstract and body evaluate different model sets and the metric's bounds are incorrect.
-
Vision-Language Models for Edge Networks: A Comprehensive Survey
A survey of lightweight vision-language models for edge deployment, marred by citation errors, self-citation, and a lack of selection methodology.
Discussion (0). Continue with ORCID to comment.