Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:1908.07490.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T17:02:12.829977Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T00:37:42.291267Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ba334614-f407-4928-8d40-f8e10ac99731 · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f52d41d-6743-4fc0-a56c-4e356703c2aa · inbound
PaLI: A Jointly-Scaled Multilingual Language-Image Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1981de44-76ca-47b1-b9ca-047c7d791076 · inbound
ViperGPT: Visual Inference via Python Execution for Reasoning LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d1f9f3f-7ca6-4556-82eb-4837bfb1c273 · inbound
LRM: Large Reconstruction Model for Single Image to 3D LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 096c54d1-f988-48fe-9848-86c5be5869b2 · inbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d75cec6-4861-42c7-81e5-16e6a0effe0e · inbound
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91a8e37b-1abe-4736-892c-1d79bad6b841 · inbound
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 584047ac-5975-464c-824c-4257dd046313 · inbound
A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bda0e5e-c36d-4fed-bc00-bec0417f5acd · inbound
Vision-Language Models for Edge Networks: A Comprehensive Survey LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6686c7-ab00-4077-a038-34a53b95c176 · inbound
A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8f1508-10d9-4ba2-a231-eccf35f6dd0d · inbound
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad827976-8831-489f-b4b5-16461dbe007a · inbound
Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3966b13-3e4a-4d4c-850c-bd9290508bb3 · inbound
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 438c66c2-4e98-4e3b-b30a-6a4730ea0a76 · inbound
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c12e9f4-35aa-4e84-acf6-a17b997d9deb · inbound
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68915107-b0f2-4811-a498-52b312ddd67e · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d81a9de-1613-4213-bc71-9f7748760389 · inbound
NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8c1f7a-624f-43b8-8446-1332f593c62e · inbound
Can Argus Judge Them All? Comparing VLMs Across Domains LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3ecff2-ba38-4c41-831a-1485eb41c438 · inbound
Gait-Based Hand Load Estimation via Deep Latent Variable Models with Auxiliary Information LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a45b9195-cc8d-4b56-bb83-b2f68133f193 · inbound
Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 175
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6a6227-8c2a-48e5-b1b4-bd001b1f38ea · inbound
Can Mental Imagery Improve the Thinking Capabilities of AI Systems? LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b684e1d3-5f0b-48d6-9500-46836fb80135 · inbound
Boosting Team Modeling through Tempo-Relational Representation Learning LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91065e8e-bc80-41f5-afc7-13d83357015c · inbound
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f2ffce-1378-4e8c-b7c1-c31840025d92 · inbound
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8aa6e69-9107-4cb1-865e-bed3001706f7 · inbound
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 197
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e07906-c825-4fd2-bdb6-8fb4ea00a8d4 · inbound
Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da686ded-f3b3-477c-a560-e3ce5c04f5e7 · inbound
Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1af62fc-58b4-4558-9452-ee96da6a38eb · inbound
Adversarial Video Promotion Against Text-to-Video Retrieval LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88457843-4aaa-4fe1-9559-195a855c8e51 · inbound
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf8c332-6e4a-4f96-a356-45b0093cb043 · inbound
BERT-VQA: Visual Question Answering on Plots LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ead4494-65b0-47ce-81c4-edce3d41d116 · inbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4d9f69-cad2-444b-ba9f-b2c912d277e4 · inbound
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e494f152-e49b-4c76-85be-fa7ec22e0c33 · inbound
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a0c9706-83a1-4e0d-9b84-49a7bad54e68 · inbound
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 480d728a-9bc5-4934-8f61-6d6448b3abc2 · inbound
MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a31957-7a3d-4c76-9f3b-7a0ed32fe78c · inbound
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 65c744d9-19bf-4d3a-990d-f1403b90b09a · inbound
A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 148
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7508e585-66e0-4edc-8a6b-08c9e4727595 · inbound
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f11e4c-a469-426a-aa0c-ef0264832044 · inbound
Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bcef3c78-200a-4a00-a26c-91a2e7ef3f0e · inbound
Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f648b5d-d5d3-458c-8abe-b05034f1d775 · inbound
Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5bb44a68-9975-492f-8ebe-af6f3acc1e06 · inbound
SpecPL: Disentangling Spectral Granularity for Prompt Learning LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ed6773f2-8616-48be-9449-0a1cdc6b48be · inbound
Multimodal LLMs under Pairwise Modalities LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6dd99fb5-5bb4-45b4-9e2a-353802844ec5 · inbound
Disentanglement-Based Equivariant Learning for Compositional VQA LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1d4918e-fecc-46a0-9a86-32c4c830a764 · inbound
Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 90da3d9a-8be5-4109-bcc1-48b211633456 · inbound
Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ea35df6-caa7-4967-98b8-ce76eec7acbf · inbound
Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cdf35859-6f0c-4d34-b4ed-a6c434325ce7 · inbound
XRFormer: Multiscale Tokenization for XRF Representation Learning LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d18b334d-af2e-4686-b0cb-9dd7153719b8 · inbound
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6185bd-e797-4646-8163-e7b2fd376eef · inbound
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990a5aac-30b0-4a58-b45a-5bc86c803de2 · inbound
Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.