Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2005.00200.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.470416Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:38:56.173895Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f4092bb5-b83a-4515-add3-d7bca8a114e5 · inbound
Flamingo: a Visual Language Model for Few-Shot Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e4841d5-236e-42b0-a3c7-faf4e1082a84 · inbound
LVBench: An Extreme Long Video Understanding Benchmark HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e12795ea-5475-42ed-b2fc-9ccfdb18bf82 · inbound
Do Language Models Understand Time? HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f3045a2-0cff-4fa4-a05e-d11e361ffebb · inbound
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4736d93-9ddd-4276-a815-ba364627e2bb · inbound
Stitch-a-Demo: Video Demonstrations from Multistep Descriptions HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 43398dab-a72d-4441-b1e2-55da1f3df5d2 · inbound
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d47c1bce-e1cd-4041-b537-6e78ce544380 · inbound
HierSum: A Global and Local Attention Mechanism for Video Summarization HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365af228-59ce-46e0-ae1c-7e91497c43c6 · inbound
GeoMM: On Geodesic Perspective for Multi-modal Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da815215-73cc-48dd-b574-c1ff35b1a32b · inbound
Generalizing vision-language models to novel domains: A comprehensive survey HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 241
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e9770b-45bc-4ef1-95c7-74c14caa103e · inbound
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca66b65-6b30-4633-b844-1cc8f35be18f · inbound
From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1186ada7-6db8-4df6-9898-174df4cae153 · inbound
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3254beaf-9463-4b14-8c2b-92733863bfeb · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12cd106b-3b83-45b6-ab5a-bb008a2f4481 · inbound
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6319d5-aa72-41db-8405-77fa52d251b9 · inbound
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 490c03fc-1224-4d82-beb0-35458f2dc982 · inbound
Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.