Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:1904.01766.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T14:54:41.840774Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T23:59:37.719324Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0e96ad61-628f-4e10-8e03-2e6cf31489c0 · inbound
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b73720ef-e046-481f-aafc-30c09c51412e · inbound
VisualBERT: A Simple and Performant Baseline for Vision and Language VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ec81e925-860a-474d-a44e-a849cdd4ca24 · inbound
Fusion of Detected Objects in Text for Visual Question Answering VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feeec428-6444-44e7-970c-960a54674546 · inbound
Integrating Multimodal Information in Large Pretrained Transformers VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77d5504-8c88-424c-8260-53f0907d3b74 · inbound
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be1c8a77-d72a-4ec6-85a8-9663f517e785 · inbound
LXMERT: Learning Cross-Modality Encoder Representations from Transformers VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a04d4f-aff1-43aa-83a4-bdae4f83e082 · inbound
VL-BERT: Pre-training of Generic Visual-Linguistic Representations VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd433f35-c5af-4b81-8a91-363358b1087f · inbound
CodeBERT: A Pre-Trained Model for Programming and Natural Languages VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a3b4c4f3-f47a-4e99-aecd-871d1b2de118 · inbound
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c44360-d008-4f0e-a0ae-fa6d3816ddef · inbound
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation befd7d9a-3877-4b66-8cf0-7bd50b3d465f · inbound
CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.