Pith. sign in

Paper Citation Record · LEDGER

ViViT: A Video Vision Transformer

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2103.15691.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.15691 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:57:29.979050Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f31e5cbf-f44b-4c67-bf61-20a3706da946 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models ViViT: A Video Vision Transformer

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.097725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:d52a6d550dd60cd2a276198847c57b601b51fb0ffd18a216c449276f9bfd165b

Observation bcfb8bbe-a515-4a07-a03a-4f0d26cfcc6e · inbound

Financial Fine-tuning a Large Time Series Model cites this paper.

Financial Fine-tuning a Large Time Series Model ViViT: A Video Vision Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:21.382552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:21.382552Z digest=sha256:621c86c4c672607c8b233791f7190b55edacfbee3d2c1485c2ddcc0c8488491c

Observation 9dcf01d1-605b-4013-9a10-6db96d4dc130 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey ViViT: A Video Vision Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.107233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.107233Z digest=sha256:de2ef4265c90fbd5d05af0959616e215fe11215d994f5d6d33d52fddb5507b96

Observation 789d1064-082d-4a6b-a9d0-1db1d9598107 · inbound

MATEY: multiscale adaptive foundation models for spatiotemporal physical systems cites this paper.

MATEY: multiscale adaptive foundation models for spatiotemporal physical systems ViViT: A Video Vision Transformer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T23:20:52.220620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:20:52.220620Z digest=sha256:c98b914ee088c0a4c3cf6dd1d0c5c6494d8e8cd7ebf9ed936c829c65752298e3

Observation 5215fa92-e0e8-4f76-a95d-f5674f04eff1 · inbound

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video cites this paper.

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video ViViT: A Video Vision Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:57:29.979050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:57:29.979050Z digest=sha256:2a171b54320267333e381802e6f8ab36d0d0ab3960ad7144d2a7b7559e452eca

Observation 2d0c896f-fa65-4a2a-9680-8e1701779dae · inbound

TSLFormer: A Lightweight Transformer Model for Turkish Sign Language Recognition Using Skeletal Landmarks cites this paper.

TSLFormer: A Lightweight Transformer Model for Turkish Sign Language Recognition Using Skeletal Landmarks ViViT: A Video Vision Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:30:39.771377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:30:39.771377Z digest=sha256:24be9c26c2c84f74b96b1ebe40bf43d889cb83946c1068dfd3aaaabd4a959faf

Observation 70c1466b-8666-4691-b0b8-f6c45601fca1 · inbound

Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions cites this paper.

Time to Embed: Unlocking Foundation Models for Time Series with Channel Descriptions ViViT: A Video Vision Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:33.510207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:33.510207Z digest=sha256:d5a15afdf8714899dd6ab3e3bbcd7db6e3fa161c768584c348f6cc8de8c8e7a0

Observation 05e657ca-e725-432d-850e-cdb03a59b95c · inbound

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks cites this paper.

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks ViViT: A Video Vision Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:14.155465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:14.155465Z digest=sha256:3f1e0191d815825f6d5f927f81d5ce37cd07f7092450432cb0c96c8f12e29f9b

Observation 7f79dfc7-bd56-47fb-961d-37b147cd9c89 · inbound

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing cites this paper.

ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing ViViT: A Video Vision Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:25.158984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:25.158984Z digest=sha256:99e42f4c0d1950f6b89d08a80a04ab51a4e09944a830f929b925673beeb7797c

Observation 7aeb6395-7261-4df5-8193-6f0bfff8ac1e · inbound

One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels cites this paper.

One Video to Steal Them All: 3D-Printing IP Theft through Optical Side-Channels ViViT: A Video Vision Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:13.917814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:13.917814Z digest=sha256:b2c38545bebbd24c669da1ec0c6f44e84dd377cc85de1199d3aad3742a53c59e

Observation 426e24a1-26e0-4c4e-b4f4-45e39639b2f8 · inbound

MVP: Winning Solution to SMP Challenge 2025 Video Track cites this paper.

MVP: Winning Solution to SMP Challenge 2025 Video Track ViViT: A Video Vision Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:56.096420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:56.096420Z digest=sha256:19a2c0630f39839c777be270246d8180ca2abae438e3858d278e7e2cd9352615

Observation 2c97d037-7e18-4489-90a6-291bf823bb0d · inbound

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges cites this paper.

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges ViViT: A Video Vision Transformer

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:06.121709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:06.121709Z digest=sha256:20938ae529691e0cadd07225a552fc53824c0f1da17fd955019c441c157262f4

Observation 51506fc2-78d5-4c41-8a8a-ad5da45c5bfa · inbound

A Space-Time Transformer for Precipitation Nowcasting cites this paper.

A Space-Time Transformer for Precipitation Nowcasting ViViT: A Video Vision Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T22:19:23.918059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:19:23.918059Z digest=sha256:e0139ccbd18852f4d11130a85bffc94829752040804fbc01cb1b9d385d3b8368

Observation 512e385e-51a5-42bc-a25e-99820ab7032c · inbound

Reasoning-Aware Multimodal Fusion for Hateful Video Detection cites this paper.

Reasoning-Aware Multimodal Fusion for Hateful Video Detection ViViT: A Video Vision Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:00:12.936776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:00:12.936776Z digest=sha256:c4b33e54ce4496ecf64dccb812098fe44b5944ff6c5963f3541e11401f0c57dc

Observation 519970c0-5cce-46ec-a385-6a77cadd5237 · inbound

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection cites this paper.

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection ViViT: A Video Vision Transformer

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:52:13.649938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T07:51:04.580617Z digest=sha256:f936303d3a16e7fe249dda54748b20c833d22b0bbcb65bf22edf157fb4bdaae7

Observation 5a4ef4c4-b4fd-4ccd-9bac-a8cae9785b62 · inbound

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction cites this paper.

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction ViViT: A Video Vision Transformer

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:14:36.513569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T23:13:34.852488Z digest=sha256:4a39dbb2935430a2acc15b81c7606c3ac062843751d0238e36cdbb78ec7f4f3e

Observation 4ba356d2-41e2-4291-8183-1fb3897e971f · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding ViViT: A Video Vision Transformer

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.430327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:5b46cadbe9ee9aed701465c5fc35f808e36954574acfedd4cdf9e322b37063bc

Observation 133d960b-ffeb-417f-9f00-30c1ea50b6ce · inbound

AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe cites this paper.

AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe ViViT: A Video Vision Transformer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:14.277963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T22:42:52.758195Z digest=sha256:a748ab9ab5efc671b7042db051b641581f3d8005087855cf41c80c6e54cb7eb5

Observation b7430e96-a421-4f08-a9b4-b46c21738109 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing ViViT: A Video Vision Transformer

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.710238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:e4ba0c5b0e8bff86e5b7ff2eaf3058d81f4145d8120343b34d8ed9654f409ef1

Observation da044df2-9510-4584-b6ba-143292ad2d08 · inbound

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living cites this paper.

TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living ViViT: A Video Vision Transformer

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:30.225196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T18:22:13.147215Z digest=sha256:b9bdf5c6d0382142117970f1199be1d3fc71a1d0223af9dbd9d7ca6dbb2ea32d

Observation dfe56956-42da-43b4-8f77-0068d3d55d1f · inbound

Physics-guided spatiotemporal neural models for fuel density prediction cites this paper.

Physics-guided spatiotemporal neural models for fuel density prediction ViViT: A Video Vision Transformer

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:06:35.216238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T22:03:33.848601Z digest=sha256:9bb7312b6793d52968e4767182c32e1136bc73002282f06580e78ac3f6664281