Pith. sign in

Paper Citation Record · LEDGER

LXMERT: Learning Cross-Modality Encoder Representations from Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:1908.07490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.07490 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:02:12.829977Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:37:42.291267Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ba334614-f407-4928-8d40-f8e10ac99731 · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:53:08.415123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:fb20943c9a7f7dbfae0b993cc3a0ad121f26f164083479ef6fddd34ee1a74cb0

Observation 7f52d41d-6743-4fc0-a56c-4e356703c2aa · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:29:06.106598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:c8a11cb6cce112b2abefee173470fe41fb6d075a8ff4f567d89dccc2783a30a9

Observation 1981de44-76ca-47b1-b9ca-047c7d791076 · inbound

ViperGPT: Visual Inference via Python Execution for Reasoning cites this paper.

ViperGPT: Visual Inference via Python Execution for Reasoning LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:15:14.578853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T18:15:14.382011Z digest=sha256:8116561e2e51633c5f68d3e93045f489d1fbc0a4dec0dd7bc7fa593c8e339ae4

Observation 3d1f9f3f-7ca6-4556-82eb-4837bfb1c273 · inbound

LRM: Large Reconstruction Model for Single Image to 3D cites this paper.

LRM: Large Reconstruction Model for Single Image to 3D LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:11:00.828458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T10:11:00.777193Z digest=sha256:e13a95e52f01947f56053f7b187632ae09961e0cabe8c6b268b64ca5478b770c

Observation 096c54d1-f988-48fe-9848-86c5be5869b2 · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.088371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:91dde8c86ec862c4ebddc0ea5bae41d255b95c77f1ce9ed0d65af7e3031a2e1e

Observation 5d75cec6-4861-42c7-81e5-16e6a0effe0e · inbound

Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving cites this paper.

Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:24:23.889004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:24:23.756052Z digest=sha256:9ca4377ec3f3f7d5e8a7110d91129709bcfbabb63c40a134b243d89a839d084a

Observation 91a8e37b-1abe-4736-892c-1d79bad6b841 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.414549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:4869b6474aa8554d2e17a852f3f1521e700c14f2284691d277defca050e20ad2

Observation 584047ac-5975-464c-824c-4257dd046313 · inbound

A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions cites this paper.

A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:12.829977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:12.829977Z digest=sha256:897dfd7f8b6cf71e0d1d06229e1605264b59684141e2262b3eacc190266fd6d9

Observation 8bda0e5e-c36d-4fed-bc00-bec0417f5acd · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.319748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.319748Z digest=sha256:4eec49a57a097c6f40b6163b6e1964da9897b802ef3ae9230600cf1628537cf2

Observation 1c6686c7-ab00-4077-a038-34a53b95c176 · inbound

A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters cites this paper.

A Survey on Data Curation for Visual Contrastive Learning: Why Crafting Effective Positive and Negative Pairs Matters LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:04.560381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:04.560381Z digest=sha256:fa52cf47c49faa6364d0ed4863bd482e674f0c44dd4a3b4ff0ab013207e61211

Observation 2a8f1508-10d9-4ba2-a231-eccf35f6dd0d · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:36.258225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:36.258225Z digest=sha256:a770dd2deb8a164bac784c0138b6bcc32601ce7af5ec4a134d650bf0163e8dc0

Observation ad827976-8831-489f-b4b5-16461dbe007a · inbound

Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection cites this paper.

Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:38.844816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:38.844816Z digest=sha256:2636649c8cffd247fa28a64b43e1c20cdd9edc0fe73439246036273a896edc6e

Observation e3966b13-3e4a-4d4c-850c-bd9290508bb3 · inbound

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering cites this paper.

MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:43.801881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:43.801881Z digest=sha256:bfb297e5949e821403a6f819b3b5175004114a45046571a5079d92f9ed266697

Observation 438c66c2-4e98-4e3b-b30a-6a4730ea0a76 · inbound

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models cites this paper.

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.557475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:36.557475Z digest=sha256:41e26b6bff10787731e72528b24077892ffecbe2236427b522f5eb7998e8c499

Observation 5c12e9f4-35aa-4e84-acf6-a17b997d9deb · inbound

Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis cites this paper.

Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:19.148749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:19.148749Z digest=sha256:dbc7db60d6d70a609eb871112ad58f80a89c38e8356586e727eece7f76a0022a

Observation 68915107-b0f2-4811-a498-52b312ddd67e · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:26.034467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:26.034467Z digest=sha256:9a6a627071726092eee5d2c9f7de69f182d15304573e6bbdcbe6fcd5de060b29

Observation 6d81a9de-1613-4213-bc71-9f7748760389 · inbound

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments cites this paper.

NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:28.107508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:28.107508Z digest=sha256:3e0dfebcf9d049915f04cab84e168fde7090f404d0528660ceb731c1daabefa0

Observation 9e8c1f7a-624f-43b8-8446-1332f593c62e · inbound

Can Argus Judge Them All? Comparing VLMs Across Domains cites this paper.

Can Argus Judge Them All? Comparing VLMs Across Domains LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:52.247632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:52.247632Z digest=sha256:21b7bc7bd82fd0e5401a87c3ae5b7e1b2932425f0e56b74ee5bc3164f52e1099

Observation 3a3ecff2-ba38-4c41-831a-1485eb41c438 · inbound

Gait-Based Hand Load Estimation via Deep Latent Variable Models with Auxiliary Information cites this paper.

Gait-Based Hand Load Estimation via Deep Latent Variable Models with Auxiliary Information LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:37.066056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:31:37.066056Z digest=sha256:9bc189a10e6a9f3fa617ae139f3303334a07f57674e9ededa6b9750c167e3e78

Observation a45b9195-cc8d-4b56-bb83-b2f68133f193 · inbound

Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures cites this paper.

Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-06T19:32:54.529492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:32:54.529492Z digest=sha256:2a1b754615802c398792b5cc6ecf0836f7c56e220c7e03d26f591063c5604585

Observation ba6a6227-8c2a-48e5-b1b4-bd001b1f38ea · inbound

Can Mental Imagery Improve the Thinking Capabilities of AI Systems? cites this paper.

Can Mental Imagery Improve the Thinking Capabilities of AI Systems? LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:48:42.087458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:48:42.087458Z digest=sha256:3b43aafd86dd3d09b438c7dec53ce0455ffbafe048dea9c703c820f76ac4bf4b

Observation b684e1d3-5f0b-48d6-9500-46836fb80135 · inbound

Boosting Team Modeling through Tempo-Relational Representation Learning cites this paper.

Boosting Team Modeling through Tempo-Relational Representation Learning LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:47:02.456878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T03:43:01.418431Z digest=sha256:d641d9962446f546959f2fd2f84e8396e251e5dad87c902b50eae9be72cace0f

Observation 91065e8e-bc80-41f5-afc7-13d83357015c · inbound

Analyzing the Sensitivity of Vision Language Models in Visual Question Answering cites this paper.

Analyzing the Sensitivity of Vision Language Models in Visual Question Answering LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:58:01.370083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:58:01.370083Z digest=sha256:04b0bf34f7f16f4995b0f7083596e9130487c8456411aa6b29c7f0db519a774f

Observation e9f2ffce-1378-4e8c-b7c1-c31840025d92 · inbound

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors cites this paper.

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:50:26.388613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:50:26.388613Z digest=sha256:df4eaa2548a70c9c46a089468ce32bd782be2f37cf2bb4ba986e87d78e2d1669

Observation b8aa6e69-9107-4cb1-865e-bed3001706f7 · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:06.945559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:06.945559Z digest=sha256:69e9874f40c66141f60f51d5544c5d5a419aad23110517638c27bebf9cf044c3

Observation 87e07906-c825-4fd2-bdb6-8fb4ea00a8d4 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:36:56.317789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:34:38.247099Z digest=sha256:ab7f82f216ed03f45d9e947a8008d3f541958e95a06178564d64d367222551cb

Observation da686ded-f3b3-477c-a560-e3ce5c04f5e7 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T00:02:58.904022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:02:58.904022Z digest=sha256:0784d7e41508eb2ccfe1ecc765485284555ea25437ce132a333bba3d4f92b046

Observation f1af62fc-58b4-4558-9452-ee96da6a38eb · inbound

Adversarial Video Promotion Against Text-to-Video Retrieval cites this paper.

Adversarial Video Promotion Against Text-to-Video Retrieval LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.172189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:05:07.182361Z digest=sha256:50a3e778dd551dad44762cc183aeca281c6b894cd4ac7cd5c8a1c8c60d204f1a

Observation 88457843-4aaa-4fe1-9559-195a855c8e51 · inbound

DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation cites this paper.

DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T21:05:33.557085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:05:33.557085Z digest=sha256:13131d621b204832c3bd807624d5720e7b3151740b6350d28c361df379919247

Observation 3bf8c332-6e4a-4f96-a356-45b0093cb043 · inbound

BERT-VQA: Visual Question Answering on Plots cites this paper.

BERT-VQA: Visual Question Answering on Plots LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:36:41.132979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:36:41.132979Z digest=sha256:a513cdaecf62a2a3751d47ffea4be674948895fd94f99766d0971100bf6b816a

Observation 1ead4494-65b0-47ce-81c4-edce3d41d116 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.223845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.223845Z digest=sha256:7a42ed3939eb7ffce5b97ca287d139b71cd7281e85b5f03e99a179ed4c677978

Observation aa4d9f69-cad2-444b-ba9f-b2c912d277e4 · inbound

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions cites this paper.

EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:20:24.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:20:24.838309Z digest=sha256:9dfe65bd32a6a22b6623beebc044de834143c5e1ad691e51000adb831bf6d250

Observation e494f152-e49b-4c76-85be-fa7ec22e0c33 · inbound

Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model cites this paper.

Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:38:34.289924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:38:34.289924Z digest=sha256:df196798fa0e0408f92b4fd20a1d42a3633d4a0f19526e2bac3f436f827b424c

Observation 3a0c9706-83a1-4e0d-9b84-49a7bad54e68 · inbound

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles cites this paper.

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:11:29.840785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:09:31.235504Z digest=sha256:48d26c1dd0353d9d8232c9324efba219cc971e609f6adbcbf1b7c544405a9c83

Observation 480d728a-9bc5-4934-8f61-6d6448b3abc2 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:05.662223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:05.662223Z digest=sha256:f1447f5afcfb0f22d283a907c4a2c0cfea4c650ca0930d2c1a89474b18b992ad

Observation 15a31957-7a3d-4c76-9f3b-7a0ed32fe78c · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.362472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:ca1a5f6e3b4cca4dca1d751281a084d8c308adb79f0b3ec30721f4c9d4b9f99a

Observation 65c744d9-19bf-4d3a-990d-f1403b90b09a · inbound

A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications cites this paper.

A Survey on Semantic Communication for Vision: Categories, Frameworks, Enabling Techniques, and Applications LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:16.484745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:16.484745Z digest=sha256:d4feaf25d44c70275ba75bc408f61bc4c6b26baf6e80d42ef5de833e2b8f076e

Observation 7508e585-66e0-4edc-8a6b-08c9e4727595 · inbound

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval cites this paper.

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T05:57:49.346929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:57:49.346929Z digest=sha256:df620f5e1b141c0502ebd6247178f8f6c95aad24c802bb0d94f8a1e53e6ceadf

Observation 26f11e4c-a469-426a-aa0c-ef0264832044 · inbound

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization cites this paper.

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:25:49.069770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:55:16.003348Z digest=sha256:e56aa1fc81073cd6dd3abcdc2ea7405fd7230f1affa17dc32ef1bb8a3b257d87

Observation bcef3c78-200a-4a00-a26c-91a2e7ef3f0e · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 134

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.261357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:e98168af3b745b267d7595057ab49483b91a8548577fc3145e4c054898649d47

Observation 4f648b5d-d5d3-458c-8abe-b05034f1d775 · inbound

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid cites this paper.

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:56:06.200204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T14:33:11.033906Z digest=sha256:e2470ea4833be4281951a6da64dbbadaeab61ffed2b88cb8f457d5a4f5ca5a1d

Observation 5bb44a68-9975-492f-8ebe-af6f3acc1e06 · inbound

SpecPL: Disentangling Spectral Granularity for Prompt Learning cites this paper.

SpecPL: Disentangling Spectral Granularity for Prompt Learning LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:11:10.378395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:51:29.352204Z digest=sha256:0a8c52224ecb148ecff96758c88c76b2614bbc8202982851c61476c5a69dc6c0

Observation ed6773f2-8616-48be-9449-0a1cdc6b48be · inbound

Multimodal LLMs under Pairwise Modalities cites this paper.

Multimodal LLMs under Pairwise Modalities LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:39:40.641308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:37:56.792564Z digest=sha256:7d3c016919c01d82b8606821669e1ed67b9eb53126a445bf2f195c06b07f587d

Observation 6dd99fb5-5bb4-45b4-9e2a-353802844ec5 · inbound

Disentanglement-Based Equivariant Learning for Compositional VQA cites this paper.

Disentanglement-Based Equivariant Learning for Compositional VQA LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:20.180376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:00:24.297044Z digest=sha256:bccc599721693444ae7c20170042c217e6744e1db3ec1ff2bfeb2fed1027ff8e

Observation e1d4918e-fecc-46a0-9a86-32c4c830a764 · inbound

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets cites this paper.

Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.010267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:44:36.059582Z digest=sha256:85e71fa1acea5c3c151406967d97bc829e5108e71a9c15e7023c66cedc9fce4f

Observation 90da3d9a-8be5-4109-bcc1-48b211633456 · inbound

Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction cites this paper.

Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.274834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:46:06.519188Z digest=sha256:0209d58b2825bf04f9abb7267c201f90b99c5024e8f9f633bdfdd7ce4d664fb3

Observation 5ea35df6-caa7-4967-98b8-ce76eec7acbf · inbound

Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts cites this paper.

Geometric and Stochastic Analysis of Discontinuities in Sparse Mixture-of-Experts LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.247106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:20:03.298578Z digest=sha256:42d7a10ac9c496f49a180d0448098a6c2686dd0a9bb9d9de8db6dd23194af99a

Observation cdf35859-6f0c-4d34-b4ed-a6c434325ce7 · inbound

XRFormer: Multiscale Tokenization for XRF Representation Learning cites this paper.

XRFormer: Multiscale Tokenization for XRF Representation Learning LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T00:37:42.318159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-11T00:34:22.775452Z digest=sha256:17a035f3492a24a5043d982dbd82ed0c61651711796db15213074f2b2061c4ee

Observation d18b334d-af2e-4686-b0cb-9dd7153719b8 · inbound

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation cites this paper.

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T03:26:19.195676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:26:19.195676Z digest=sha256:71f5d21c6baf5bfa830a312d012e19815c556da6d5b7d5f9011c8376f02a1aa3

Observation 7f6185bd-e797-4646-8163-e7b2fd376eef · inbound

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models cites this paper.

DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:08.611396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:32:08.611396Z digest=sha256:fbfd9373a125067725c0f29687c698ba78c791c17da71de64d7a959769f531e6

Observation 990a5aac-30b0-4a58-b45a-5bc86c803de2 · inbound

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis cites this paper.

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T00:56:16.616439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:56:16.616439Z digest=sha256:8859a85f417e4700a22d498947b14c749940d8e209376d65ed64c3499ae9905e