Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:37:22.950052Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2411.15284.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:37:22.950052Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:47:17.062406Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T22:13:03.994905Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dcbb4957-5bf8-4b52-96cc-63499caec748 · outbound
When Spatial meets Temporal in Action Recognition Vivit: A video vision transformer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c21acb-54e3-4686-8d7e-7ca079275a10 · outbound
When Spatial meets Temporal in Action Recognition Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb163a11-b637-46d5-850b-cb8bb40eff11 · outbound
When Spatial meets Temporal in Action Recognition Dynamic image networks for action recognition
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea20f808-bdae-49e3-9918-1197d75b4708 · outbound
When Spatial meets Temporal in Action Recognition Quo vadis, action recognition? a new model and the kinetics dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e195e238-28e8-4da5-b15c-46fcfda19077 · outbound
When Spatial meets Temporal in Action Recognition Motion meets attention: Video motion prompts
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d57f68df-f5dc-43c6-844d-dd78c382b14d · outbound
When Spatial meets Temporal in Action Recognition Temporal context network for activity local- ization in videos
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3f5f4097-f5b1-49dc-9677-fba041cc9664 · outbound
When Spatial meets Temporal in Action Recognition Imagenet: A large-scale hierarchical image database
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d999934-6798-4b1f-ad8d-4c53269d9da5 · outbound
When Spatial meets Temporal in Action Recognition An image is worth 16x16 words: Transformers for image recognition at scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d22f25a6-a32b-4e95-b68f-a696a36e35e5 · outbound
When Spatial meets Temporal in Action Recognition X3d: Expanding architectures for efficient video recognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eaae5b96-8ad1-4edb-a831-8260fef3c50c · outbound
When Spatial meets Temporal in Action Recognition Convolutional two-stream network fusion for video action recognition
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ad0506-cb55-41f3-9a49-014e18642c9d · outbound
When Spatial meets Temporal in Action Recognition Slowfast networks for video recognition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f2c177-152c-478c-9e0c-153d6b4e4535 · outbound
When Spatial meets Temporal in Action Recognition Omnimae: Single model masked pretraining on images and videos
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6319f230-35b2-49b3-8bdc-599f5ef84de3 · outbound
When Spatial meets Temporal in Action Recognition Deep residual learning for image recognition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7bd43af2-3af5-48be-be47-c2d8a0929d2b · outbound
When Spatial meets Temporal in Action Recognition Capturing temporal information in a sin- gle frame: Channel sampling strategies for action recogni- tion
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d11146d3-05cf-4aba-93b2-8b3de71e73d6 · outbound
When Spatial meets Temporal in Action Recognition Tensor rep- resentations for action recognition
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a40d2458-edb2-48e1-9f37-a26db7a3a507 · outbound
When Spatial meets Temporal in Action Recognition HMDB: A large video database for human motion recognition
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7aa6fd7-1c47-4cab-8c3f-c7df7282cc64 · outbound
When Spatial meets Temporal in Action Recognition Action recognition based on a bag of 3D points
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation beff1ebc-d42b-4de5-adc8-14c2cf3d7667 · outbound
When Spatial meets Temporal in Action Recognition Tsm: Temporal shift module for efficient video understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 077af4aa-9942-4204-8dc4-7f875ce62157 · outbound
When Spatial meets Temporal in Action Recognition Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 521e680e-45c3-4c29-af10-3c510d94f40d · outbound
When Spatial meets Temporal in Action Recognition HON4D: Histogram of Ori- ented 4D Normals for Activity Recognition from Depth Se- quences
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a6652d33-ebe9-453b-a712-8682b3ae1d5c · outbound
When Spatial meets Temporal in Action Recognition Keeping your eye on the ball: Tra- jectory attention in video transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3209e556-d7fb-4dd9-a898-2e2936449f79 · outbound
When Spatial meets Temporal in Action Recognition Huynh, and Ajmal Mian
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 72e75139-eda6-4d8c-8753-f4fd73b94af8 · outbound
When Spatial meets Temporal in Action Recognition Huynh, and Aj- mal Mian
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b9c8b198-20d0-4b5f-ac45-73c655a78bf5 · outbound
When Spatial meets Temporal in Action Recognition Ryoo, AJ Piergiovanni, Mingxing Tan, and Anelia Angelova
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12fe783b-2283-48b3-b778-fe4244df28df · outbound
When Spatial meets Temporal in Action Recognition Ntu rgb+ d: A large scale dataset for 3d human activity anal- ysis
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e6ac195-9d8d-42e9-a355-11be3fd158df · outbound
When Spatial meets Temporal in Action Recognition Two-stream con- volutional networks for action recognition in videos
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c5f08e-d4ac-4961-bd7a-dcb88dc3ae27 · outbound
When Spatial meets Temporal in Action Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705caf59-eff9-49d8-bc1d-7a1a8ec428f7 · outbound
When Spatial meets Temporal in Action Recognition Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 062eaa9b-20c1-4291-aa6b-1cb294865143 · outbound
When Spatial meets Temporal in Action Recognition Learning spatiotemporal features with 3d convolutional networks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e4e705-437c-468e-a3c3-7ee08eb6b68a · outbound
When Spatial meets Temporal in Action Recognition Self-supervising action recog- nition by statistical moment and subspace descriptors
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0faad10b-1ea4-44b6-9346-a0189581e8f2 · outbound
When Spatial meets Temporal in Action Recognition Flow dynamics correction for action recognition
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a99fad45-19bb-4c8a-a1ae-a9efa0716ccd · outbound
When Spatial meets Temporal in Action Recognition Temporal segment net- works: Towards good practices for deep action recognition
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9c5033ea-9244-4bc2-a98c-27377ba3ffd1 · outbound
When Spatial meets Temporal in Action Recognition Hallucinating idt descriptors and i3d optical flow features for action recog- nition with cnns
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e99828b2-32ce-40c7-ac3d-435e74cb72ab · outbound
When Spatial meets Temporal in Action Recognition Temporal segment networks for action recognition in videos
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4cc24b11-4211-488e-8f5b-7b0e7d2ef18f · outbound
When Spatial meets Temporal in Action Recognition Videomae v2: Scaling video masked autoencoders with dual masking
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cde2db38-350b-49d9-9e53-9bfa1f4b2709 · outbound
When Spatial meets Temporal in Action Recognition High-order tensor pooling with attention for action recognition
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8ea8f59a-aefc-4026-8c95-05b39ff3e1bd · outbound
When Spatial meets Temporal in Action Recognition Taylor videos for action recognition
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 250c1991-3a7d-43e6-a8a0-d9def199b52f · outbound
When Spatial meets Temporal in Action Recognition Internvideo2: Scaling video foundation models for multimodal video understanding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 79bb667e-1c8a-4c09-b716-6f810066c855 · outbound
When Spatial meets Temporal in Action Recognition Towards good practices for missing modality robust action recognition
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 36891865-dbf9-4bc1-8cbc-103fa35f04f4 · outbound
When Spatial meets Temporal in Action Recognition Temporal pyramid network for action recognition
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 862c278e-9696-404b-a4e6-4c40eacfd801 · outbound
When Spatial meets Temporal in Action Recognition Advancing video anomaly detection: A concise re- view and a new dataset
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 90cbbc31-e234-4a38-9c94-fefe1dcff484 · inbound
Do Language Models Understand Time? When Spatial meets Temporal in Action Recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0037c5-c705-4285-bb38-641065632d4a · inbound
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight When Spatial meets Temporal in Action Recognition
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c045b02e-4aad-4f33-b797-c7fb7cede33d · inbound
Evolving Skeletons: Motion Dynamics in Action Recognition When Spatial meets Temporal in Action Recognition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.