Pith. sign in

Paper Citation Record · LEDGER

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:1908.06066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.06066 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:01:13.283168Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:30:37.351291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T18:56:48.399498Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57f21e8d-bbdc-40d9-b965-da369a6c3c88 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.239231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.239231Z digest=sha256:c2a0d40d9ba66f9e8637a03a5a9c34d84d478d76f49190a7dd1392f2f18628ae

Observation 629aeb46-409d-49bc-9c0a-18157c1721af · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training UNITER: UNiversal Image-TExt Representation Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.242260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.242260Z digest=sha256:d7dba57fa3cabb6db79ae8f92c956d40e9340df3608f63d8b30f41dac58327ad

Observation 6d571d1d-7950-465c-a25c-7c2f6d927fe2 · outbound

This paper cites Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T13:01:13.365671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:01:13.254177Z digest=sha256:1720d769f1bfaeefb35b6148407dd46695212231cadff35c3df80acf38665bc2

Observation 55160a59-8588-412b-9488-78404417ba9c · outbound

This paper cites Cross-lingual Language Model Pretraining.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Cross-lingual Language Model Pretraining

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.257054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.257054Z digest=sha256:2bf09cc0327c6d560f47e9c124564848b84e60057a3d3090e67ce4f8c005269f

Observation 6c52cbe4-2736-4864-993a-4a72b1742700 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.259950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.259950Z digest=sha256:83b0a631c02d572095267c5ac9a591746f6c174a5b47801f0cada531e631a32c

Observation 646cbc41-217c-4cf3-a704-543d6adf4a31 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.263107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.263107Z digest=sha256:824861b895c9e135503f7e9ea4bfa30122f4a0cc2e2df97ffe1892ffdf432543

Observation 0ee6b0e0-0cef-45be-b6ad-9436b121f399 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.266000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.266000Z digest=sha256:222cde8ca3135f60a0597ebe679331adfb2e40dfe6b92306315ef17557395bb7

Observation 74967b05-acba-452a-955a-957c1b8c84e2 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.274718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.274718Z digest=sha256:54ec36a9019a9d628ee703afd14a7aedbd6e4094ac5366e78ea557652be2b2fd

Observation e77d5504-8c88-424c-8260-53f0907d3b74 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.277528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.277528Z digest=sha256:a560f254a1bab8905beebd86932429529c49eb72aad5693ffb33ad5e6ff11ede

Observation 34afb34b-0a1b-4e82-8ee2-daa0cc7cba2d · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.280406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.280406Z digest=sha256:ab68b4d23da84000bdf4746e247b5d5830b524afdc4a16d0bfc0f0f04923db6b

Observation d5f54d87-ddce-49b9-a752-40a05e710b58 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.283168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.283168Z digest=sha256:2eaced6b692c1a80d274b657c49225726292e9e4e7b73e6d54afae3248c8dd74

Observation 5b5d464f-2550-46e4-95e0-c764e5ecf590 · outbound

This paper cites In 2009 IEEE conference on computer vision and pattern recognition, 248–255.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training In 2009 IEEE conference on computer vision and pattern recognition, 248–255

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:01:13.410714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:01:13.245337Z digest=sha256:0036f5a71947e9ada9f13d55dcafeaa1739aa1d6f7743e37ce7f11b429050fca

Observation dc70a464-2a05-470d-aef4-48481e2d1098 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.271786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.271786Z digest=sha256:ce1c18db6f37eb80e810dd37a1a631c7c39d6cc9a76df56c9c7f21e1abc675f3

Observation d270b1b0-f24e-48f2-bff2-0345e1f3ea6b · outbound

This paper cites A large annotated corpus for learning natural language inference.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training A large annotated corpus for learning natural language inference

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.236228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.236228Z digest=sha256:7acfb3daaf12e4d5f587bf31c66861d2db31ce99d01e04586ef16e64a28578d5

Observation 6a986e2f-d17c-4e60-aede-58f4c888cc73 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.268745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.268745Z digest=sha256:4193b68198818bcf6363a9019ed07c7d481c8a108af0bcc97655cbfc654b88e8

Observation 3d69f56a-31b7-4b92-9e98-ec399527844f · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.251075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.251075Z digest=sha256:3896369975d2257b787236a39fa931e8527cf25575d422fa8c5abf2fe62102a9

Observation cd129c4a-1d34-4cde-94a7-c8db5588cd4d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.248145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.248145Z digest=sha256:9e532eaa0e1a9eb2355e0326d12f3082374849003143284978a674a86084da06

Observation 58888512-c5e6-46c3-8956-85fd6c948f8a · outbound

This paper cites Fusion of Detected Objects in Text for Visual Question Answering.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Fusion of Detected Objects in Text for Visual Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.232217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.232217Z digest=sha256:877977866eac847e1691b09baf7a38faaf7ebdbad3d220c83ffd14145b27c3e7

Pith citing papers

Observation dfbe82c1-dc26-4f58-bde9-8af72792eba4 · inbound

Fusion of Detected Objects in Text for Visual Question Answering cites this paper.

Fusion of Detected Objects in Text for Visual Question Answering Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.351291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.351291Z digest=sha256:8c450ed5734d5845b3d02da054a853c33ebc52c154556e9ea0f24a5e787713a1

Observation b95b9913-23a5-4f1d-8966-17d993acb34c · inbound

RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer cites this paper.

RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:56:48.403207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:56:47.983074Z digest=sha256:7080b8d9aaed0c6b5dd523957bf477f8dd6d36209dc02f217f86efde1cc7ff8b