Pith. sign in

Paper Citation Record · LEDGER

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

As of 14 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:1908.06066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.06066 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:01:13.283168Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:30:37.351291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T18:56:48.399498Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57f21e8d-bbdc-40d9-b965-da369a6c3c88 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.239231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.239231Z digest=sha256:57ef83b31f00b5a9efccc8da6507b9e6f8a4069c673bde327f1efba6b70660a6

Observation 629aeb46-409d-49bc-9c0a-18157c1721af · outbound

This paper cites UNITER: UNiversal Image-TExt Representation Learning.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training UNITER: UNiversal Image-TExt Representation Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.242260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.242260Z digest=sha256:4955c10cafd2da52414e8644279f220771e39d3663e3118942bce9873fc97629

Observation 6d571d1d-7950-465c-a25c-7c2f6d927fe2 · outbound

This paper cites Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T13:01:13.365671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:01:13.254177Z digest=sha256:9808088806849a03916fd3caf6b7902d92f2d53d8a49a94c387405302bcdd71c

Observation 55160a59-8588-412b-9488-78404417ba9c · outbound

This paper cites Cross-lingual Language Model Pretraining.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Cross-lingual Language Model Pretraining

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.257054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.257054Z digest=sha256:5082217a4dc459da170453f3d59796760d105e22f3b9d2702ce2d5a09357f3f2

Observation 6c52cbe4-2736-4864-993a-4a72b1742700 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.259950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.259950Z digest=sha256:959f17f5407bac40ec4b62cec30191f2ab6a382c3f4e95666c472b2bef6720ae

Observation 646cbc41-217c-4cf3-a704-543d6adf4a31 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.263107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.263107Z digest=sha256:32800d9bd44d163cfe564e3b66c4b9ec657baedbf9f5e8f3a6f31ca95b057921

Observation 0ee6b0e0-0cef-45be-b6ad-9436b121f399 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.266000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.266000Z digest=sha256:6cab5d32889d7b9975739fa9e6dc472e47b9ca8d69246f3df703b12c36984bea

Observation 74967b05-acba-452a-955a-957c1b8c84e2 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.274718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.274718Z digest=sha256:cb4352928b8e493d8ae7d5cf387ae9b2ef47599d7504be3d0be2cd2e9aee9a08

Observation e77d5504-8c88-424c-8260-53f0907d3b74 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.277528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.277528Z digest=sha256:c9eaac0f9dc6945fab35aa8a511696bf8ef6a64cc42d7583cb459c0b58191de7

Observation 34afb34b-0a1b-4e82-8ee2-daa0cc7cba2d · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.280406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.280406Z digest=sha256:1dc53a9166f642b975a1099b5526f3f5f4ce781a9dda63ea35dd8d7bf9d98db5

Observation d5f54d87-ddce-49b9-a752-40a05e710b58 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.283168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.283168Z digest=sha256:9a96bed30c2565554206af6c587f87e628f3ce886c2d2a8049003e294daf1192

Observation 5b5d464f-2550-46e4-95e0-c764e5ecf590 · outbound

This paper cites In 2009 IEEE conference on computer vision and pattern recognition, 248–255.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training In 2009 IEEE conference on computer vision and pattern recognition, 248–255

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:01:13.410714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:01:13.245337Z digest=sha256:3a987925f8117acaa37f87f73763036b3737d98627670c4e694c1b657248df7b

Observation dc70a464-2a05-470d-aef4-48481e2d1098 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.271786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.271786Z digest=sha256:35f2700fd3bba82a468bdfba9e69023fd3bbdc67085610b8256239b66fba6cf4

Observation d270b1b0-f24e-48f2-bff2-0345e1f3ea6b · outbound

This paper cites A large annotated corpus for learning natural language inference.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training A large annotated corpus for learning natural language inference

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.236228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.236228Z digest=sha256:3b74786e37fc9f2285d0602e235203fe2b84a50cae48928f7ac55e4977c89120

Observation 6a986e2f-d17c-4e60-aede-58f4c888cc73 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.268745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.268745Z digest=sha256:13543b0a13423b9ea7bdf94c4899789479977483c64277224d31ea5cc65cd87e

Observation 3d69f56a-31b7-4b92-9e98-ec399527844f · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.251075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.251075Z digest=sha256:0206348126c9741e0dd536b1214633eff9d499661f5ef8cd9db11b22d11b98de

Observation cd129c4a-1d34-4cde-94a7-c8db5588cd4d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.248145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.248145Z digest=sha256:c8d7a31c4dea645a541d60553d3b00b7545b4ce8be6986389f071784231f3117

Observation 58888512-c5e6-46c3-8956-85fd6c948f8a · outbound

This paper cites Fusion of Detected Objects in Text for Visual Question Answering.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Fusion of Detected Objects in Text for Visual Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.232217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.232217Z digest=sha256:0a11b23afdcb9ea0127ad2f3250ffa604b6787764732bdbabb29db3be455ee6d

Pith citing papers

Observation dfbe82c1-dc26-4f58-bde9-8af72792eba4 · inbound

Fusion of Detected Objects in Text for Visual Question Answering cites this paper.

Fusion of Detected Objects in Text for Visual Question Answering Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.351291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.351291Z digest=sha256:e26ee5aa0e259730629ae599858d699a7dfa8989d1a6d523933313e9bc0728db

Observation b95b9913-23a5-4f1d-8966-17d993acb34c · inbound

RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer cites this paper.

RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:56:48.403207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T18:56:47.983074Z digest=sha256:c045cd41732580a7594de1cb85a9ea5804bc37c0377b2212ea82c13d3d56d4b1