Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2103.15691.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:57:29.979050Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f31e5cbf-f44b-4c67-bf61-20a3706da946 · inbound
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models ViViT: A Video Vision Transformer
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bcfb8bbe-a515-4a07-a03a-4f0d26cfcc6e · inbound
Financial Fine-tuning a Large Time Series Model ViViT: A Video Vision Transformer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dcf01d1-605b-4013-9a10-6db96d4dc130 · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey ViViT: A Video Vision Transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789d1064-082d-4a6b-a9d0-1db1d9598107 · inbound
MATEY: multiscale adaptive foundation models for spatiotemporal physical systems ViViT: A Video Vision Transformer
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5215fa92-e0e8-4f76-a95d-f5674f04eff1 · inbound
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video ViViT: A Video Vision Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0c896f-fa65-4a2a-9680-8e1701779dae · inbound
TSLFormer: A Lightweight Transformer Model for Turkish Sign Language Recognition Using Skeletal Landmarks ViViT: A Video Vision Transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c1466b-8666-4691-b0b8-f6c45601fca1 · inbound
Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions ViViT: A Video Vision Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e657ca-e725-432d-850e-cdb03a59b95c · inbound
Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks ViViT: A Video Vision Transformer
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f79dfc7-bd56-47fb-961d-37b147cd9c89 · inbound
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing ViViT: A Video Vision Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aeb6395-7261-4df5-8193-6f0bfff8ac1e · inbound
One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels ViViT: A Video Vision Transformer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426e24a1-26e0-4c4e-b4f4-45e39639b2f8 · inbound
MVP: Winning Solution to SMP Challenge 2025 Video Track ViViT: A Video Vision Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c97d037-7e18-4489-90a6-291bf823bb0d · inbound
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges ViViT: A Video Vision Transformer
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51506fc2-78d5-4c41-8a8a-ad5da45c5bfa · inbound
A Space-Time Transformer for Precipitation Nowcasting ViViT: A Video Vision Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512e385e-51a5-42bc-a25e-99820ab7032c · inbound
Reasoning-Aware Multimodal Fusion for Hateful Video Detection ViViT: A Video Vision Transformer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 519970c0-5cce-46ec-a385-6a77cadd5237 · inbound
DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection ViViT: A Video Vision Transformer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a4ef4c4-b4fd-4ccd-9bac-a8cae9785b62 · inbound
Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction ViViT: A Video Vision Transformer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4ba356d2-41e2-4291-8183-1fb3897e971f · inbound
Exploring High-Order Self-Similarity for Video Understanding ViViT: A Video Vision Transformer
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 133d960b-ffeb-417f-9f00-30c1ea50b6ce · inbound
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe ViViT: A Video Vision Transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b7430e96-a421-4f08-a9b4-b46c21738109 · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing ViViT: A Video Vision Transformer
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation da044df2-9510-4584-b6ba-143292ad2d08 · inbound
TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living ViViT: A Video Vision Transformer
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dfe56956-42da-43b4-8f77-0068d3d55d1f · inbound
Physics-guided spatiotemporal neural models for fuel density prediction ViViT: A Video Vision Transformer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.