Pith. sign in

Paper Citation Record · LEDGER

Is Space-Time Attention All You Need for Video Understanding?

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2102.05095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.05095 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:28.412053Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1359
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 140aac23-27fa-41c3-8ea3-ac9d6efb6e82 · inbound

Video Diffusion Models cites this paper.

Video Diffusion Models Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:38:27.964923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T14:38:27.919104Z digest=sha256:e4a752cd51315b375386ea7045fcdbd3ca696f2d401e80d8a83653a7eedc9d53

Observation 5ab418a0-6bfc-4f97-990e-a8b928f6f186 · inbound

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows cites this paper.

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows Is Space-Time Attention All You Need for Video Understanding?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:28.412053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:28.412053Z digest=sha256:02c9ac2164629c81518352473dfd059a14915800f2dcd5ac4dc98d4fd48f6458

Observation eeb68259-9d9e-44d4-b995-f92cc6f0856e · inbound

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks cites this paper.

Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks Is Space-Time Attention All You Need for Video Understanding?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:14.159369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:46:14.159369Z digest=sha256:f68425fdd7d090aa488c8608193c355ca759a6a52f537fab7d47e7ff51bc52ed

Observation 9e0c1bef-2930-49ac-872f-316afaa219da · inbound

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective cites this paper.

Data-Efficient Challenges in Visual Inductive Priors: A Retrospective Is Space-Time Attention All You Need for Video Understanding?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:59.021704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:10:59.021704Z digest=sha256:7063b73467d8f77a928677ac654f97700ae553faef01c54b6d8abe623e19d0e3

Observation a890a18d-7361-412a-8960-615401f929aa · inbound

Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications cites this paper.

Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications Is Space-Time Attention All You Need for Video Understanding?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:10.626908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:59:15.937439Z digest=sha256:75338d70af1ab7d1cbf7629dc2521d5b82c6041f307a0cc632825e5ca9e77da2

Observation bd7e80c6-f1a6-4542-be12-6e0b683c42a7 · inbound

Comparing Learning Paradigms for Egocentric Video Summarization cites this paper.

Comparing Learning Paradigms for Egocentric Video Summarization Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.597659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.597659Z digest=sha256:506ffda2f9602247383cbb0314f16dc6a8febbf618df9b44e87ef0d191f15937

Observation 1c41e3d9-773f-4c9d-bf1e-460f39b99941 · inbound

MVP: Winning Solution to SMP Challenge 2025 Video Track cites this paper.

MVP: Winning Solution to SMP Challenge 2025 Video Track Is Space-Time Attention All You Need for Video Understanding?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:56.108327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:56.108327Z digest=sha256:4bf0b1c5023e03f5af7253a9bf58322e1fe94681cdb38ca3ceb51d5e3946ce91

Observation d9c232d4-c6b1-404e-8f17-5e6da800d8ae · inbound

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes cites this paper.

Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:01.481907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:01.481907Z digest=sha256:b939b5c0515c2a4597f681cf658db761f1fcf0511181824fe8b631ef5a7275d0

Observation 82bce4bf-2976-4f3a-a159-93b903fc27f8 · inbound

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition cites this paper.

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:13.584772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:13.584772Z digest=sha256:d2acf35ba0f7865c89ac594a181e9f6adf1e3da03d34b6bc202b4b5bd4b57ab1

Observation 2a5d2dfb-a2a1-4b8a-bbeb-84eca88c0abd · inbound

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis cites this paper.

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis Is Space-Time Attention All You Need for Video Understanding?

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:40:55.918865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:37:51.737038Z digest=sha256:afda96c6e59b395cfc44d11adbe24abba2233a0eb66e542e4a487a37e675ac3d

Observation 76e0ead7-bfbe-4fa7-9d3f-13c125b192eb · inbound

A Space-Time Transformer for Precipitation Nowcasting cites this paper.

A Space-Time Transformer for Precipitation Nowcasting Is Space-Time Attention All You Need for Video Understanding?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:19:23.932491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:19:23.932491Z digest=sha256:a5932b7162e26bc045e358c4f05495c6bea62e60d67dd35fbc09e1aaf92d386e

Observation 45670124-a771-4715-b6df-cf104cb965f5 · inbound

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection cites this paper.

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection Is Space-Time Attention All You Need for Video Understanding?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:13:39.376251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:13:02.040618Z digest=sha256:4819b071d7674f0188492eb79da3b2e7eaecec994970b22e4d43e05ee4bd3807

Observation 0c425cb9-266e-4991-aa6e-391b8754c243 · inbound

Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition cites this paper.

Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition Is Space-Time Attention All You Need for Video Understanding?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:55:33.481242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:55:17.256153Z digest=sha256:034e27279a6bfa2034b0a73557cc52bb7f5220cfc217aa1621f0dd605ff39c8d

Observation 8fd48442-8300-4968-afb9-7ebebff1f7dc · inbound

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction cites this paper.

Seeing Further and Wider: Joint Spatio-Temporal Enlargement for Micro-Video Popularity Prediction Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:14:36.516021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T23:13:34.852488Z digest=sha256:e481e8f3265cc0fc47bc20859c0c97cc2bbb1bbc429e35e4d800027d4c8f3c36

Observation 8a4b40fe-1be0-441a-bd1e-de4a4bf32084 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:59:49.399058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:ff5aa28fcda5228b45a7e9221d70f8259ccd802052a1960d2d141188b801ada5

Observation 539c8429-561d-4ffc-b7fa-1f728c33144f · inbound

Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models cites this paper.

Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models Is Space-Time Attention All You Need for Video Understanding?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:56:07.521561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T13:13:03.129195Z digest=sha256:91acea907b9904fcd269ec8531f92e7242504203e94099ad03c6cb5d68dc1d85

Observation 6d788f70-2247-45f5-90fb-baff01f58a14 · inbound

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback cites this paper.

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback Is Space-Time Attention All You Need for Video Understanding?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:24.059130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T17:44:28.109134Z digest=sha256:8f66c502a354769ce709a2af5eec9ac90d4b5a0fd09e89b5293c54100766da38

Observation e67b4d71-56db-49d5-85a0-7d6e499a8a29 · inbound

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute cites this paper.

LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute Is Space-Time Attention All You Need for Video Understanding?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:40:52.109484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:36:50.406353Z digest=sha256:576cc8728678ea56bb51ff9e21737f3606f9284dd64315cd4e10f5ed117e4266

Observation edb86348-ec6e-48fb-9f45-737820589425 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection Is Space-Time Attention All You Need for Video Understanding?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.080824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:390cf7e55399f4e242f605a154682aa28032a87dcc674fb5ca3b76492ab7b2ca

Observation a2303355-69b9-4e82-a690-322acc4b77b5 · inbound

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers cites this paper.

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:13:48.619377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:08:42.533804Z digest=sha256:129564588f682667a6b72aeb417344745fa8f8bc970a53381697696bbd59aeaf

Observation daa30a4d-0e4f-4095-b886-38e9453a7445 · inbound

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting cites this paper.

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:45.540511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T07:00:04.541121Z digest=sha256:3e6e863b1d11a34f53b41f564bef7de16076aea25254f1e58860cd31ff8d1555

Observation 681dc00c-6bde-453d-a614-ff4caba6f2f5 · inbound

A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding cites this paper.

A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding Is Space-Time Attention All You Need for Video Understanding?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:38:50.693618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:37:58.322339Z digest=sha256:5929f58ffcb1c363f57e39dba9916cccbf796afcc67134cd40ed9d7ba94baf4a

Observation 520529b3-782c-46a2-a61c-cdd8a937fe75 · inbound

Incentivizing Vision Language Models to Search for Long Video Question Answering cites this paper.

Incentivizing Vision Language Models to Search for Long Video Question Answering Is Space-Time Attention All You Need for Video Understanding?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T05:50:16.895740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:50:16.895740Z digest=sha256:f2e9e8d617f28bfaee898311b31f4f8212392d4a3cc316610e2205babe8d002f

Observation 93a2de9b-9a1a-4845-874b-f4b23582ace7 · inbound

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy cites this paper.

The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy Is Space-Time Attention All You Need for Video Understanding?

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-14T12:06:28.342674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:06:28.342674Z digest=sha256:d90d2e5eef6e4353a9424cb252685b57cbe29195daf6c50a59c172e8db2bab6a

Observation ba800377-3a00-4419-9e01-e6dda56e0606 · inbound

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs cites this paper.

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs Is Space-Time Attention All You Need for Video Understanding?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T04:39:19.873848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:39:19.873848Z digest=sha256:7e64812fff5fb4dac9d03b64df25439de445fe6c4d228246f7a1e85bf03fa60b