Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2111.12681.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:31:42.408928Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:57:30.204945Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 066ab485-dc68-4c03-afbd-710371346d83 · inbound
Flamingo: a Visual Language Model for Few-Shot Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1b3dfd1-d7f2-4246-a5ad-9a97e7ede817 · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d54d0cf7-6a98-4114-9471-058a461bb26f · inbound
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ab787156-95d0-45d3-be18-1de421898834 · inbound
VideoChat: Chat-Centric Video Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 62b232a8-918e-44f4-bd7a-45d429e693fb · inbound
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc4d57cd-db08-43ab-ae1b-e401303b6ac2 · inbound
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 067cc1fb-e7f2-489f-a50a-ad1df38a4057 · inbound
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99c967b1-335c-40f3-b66f-b18e8b68e8d1 · inbound
Large Language Models for Multi-Robot Systems: A Survey VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 33cddc44-a6fa-4754-a237-163b82c9746b · inbound
Character-Centered Dialogue Generation from Scene-Level Prompts VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 505417fb-eae6-4fa0-b2de-e1c0b6e9f9d2 · inbound
EgoM2P: Egocentric Multimodal Multitask Pretraining VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5093ead1-14f7-42ba-a2bc-e0a2ff8d02f9 · inbound
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74a9922-3306-40c0-8180-d6a268075e48 · inbound
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ada825-88a6-496d-b8f1-1a9fce6c9643 · inbound
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4ed516-d4be-4c5b-a2a9-b686192812be · inbound
Video Understanding by Design: How Datasets Shape Video Models VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 226
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a5bd31-04a1-4392-ad73-e48d61960717 · inbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16f6576a-d254-457f-bec5-29e5253bfc26 · inbound
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bbea9e94-ae46-4d34-92c7-29194f78bfbc · inbound
Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.