Pith. sign in

Paper Citation Record · LEDGER

When Spatial meets Temporal in Action Recognition

As of 15 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 3 inbound Pith citation observations for arXiv:2411.15284.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15284 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:37:22.950052Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:47:17.062406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T22:13:03.994905Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcbb4957-5bf8-4b52-96cc-63499caec748 · outbound

This paper cites Vivit: A video vision transformer.

When Spatial meets Temporal in Action Recognition Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.704587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.704587Z digest=sha256:2005edaddc81a49efbdc6f30c7855ab074a07ba06f74da263afb623a211ff77a

Observation 13c21acb-54e3-4686-8d7e-7ca079275a10 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

When Spatial meets Temporal in Action Recognition Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:24.013939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.710836Z digest=sha256:1a035cdc3eda3d702592f3fb4b76fca9baa023a7112fa9f4d8147b2cf590282d

Observation cb163a11-b637-46d5-850b-cb8bb40eff11 · outbound

This paper cites Dynamic image networks for action recognition.

When Spatial meets Temporal in Action Recognition Dynamic image networks for action recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.994465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.717137Z digest=sha256:61507d7f988bf28af6ee8297540b69582f0b0bbaf149f4f4de1389e8f38ec011

Observation ea20f808-bdae-49e3-9918-1197d75b4708 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

When Spatial meets Temporal in Action Recognition Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.973354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.723830Z digest=sha256:d451672a048a6e55da1b352bc496d0ca9ff907a35565f3c87f5ab38dd443bf90

Observation e195e238-28e8-4da5-b15c-46fcfda19077 · outbound

This paper cites Motion meets attention: Video motion prompts.

When Spatial meets Temporal in Action Recognition Motion meets attention: Video motion prompts

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.951591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.730291Z digest=sha256:8166f939e57f04d992d329d797577974b5c7b9c42de21880bf199081957d300f

Observation d57f68df-f5dc-43c6-844d-dd78c382b14d · outbound

This paper cites Temporal context network for activity local- ization in videos.

When Spatial meets Temporal in Action Recognition Temporal context network for activity local- ization in videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.928064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.737124Z digest=sha256:dde8ac8bb7d899906f33255c28ec71e4ee35b3496afbcbc87e6bbc2885ed6040

Observation 3f5f4097-f5b1-49dc-9677-fba041cc9664 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

When Spatial meets Temporal in Action Recognition Imagenet: A large-scale hierarchical image database

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.742280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.742280Z digest=sha256:697755bb9932bba1bc3968ae0570a44b796086b36b7d5023eb25a7e7aedbfe22

Observation 2d999934-6798-4b1f-ad8d-4c53269d9da5 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

When Spatial meets Temporal in Action Recognition An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.892335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.752411Z digest=sha256:3b15c5e32c0a135e8faa234294761bf8eb6452bfa7d23c2d4ff0516b7b3c3f66

Observation d22f25a6-a32b-4e95-b68f-a696a36e35e5 · outbound

This paper cites X3d: Expanding architectures for efficient video recognition.

When Spatial meets Temporal in Action Recognition X3d: Expanding architectures for efficient video recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.872163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.759559Z digest=sha256:b8202ee4bc8aeefaf65fcbd519a4c0701216293d55c0d5b7e72f6530d7511fbd

Observation eaae5b96-8ad1-4edb-a831-8260fef3c50c · outbound

This paper cites Convolutional two-stream network fusion for video action recognition.

When Spatial meets Temporal in Action Recognition Convolutional two-stream network fusion for video action recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.764522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.764522Z digest=sha256:b08fb909c8acbf5cf52bc19429e06b86e1f6a940f0ed919cb4a429ff6625a61c

Observation 91ad0506-cb55-41f3-9a49-014e18642c9d · outbound

This paper cites Slowfast networks for video recognition.

When Spatial meets Temporal in Action Recognition Slowfast networks for video recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.771016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.771016Z digest=sha256:04e92b15509586cb5443e1d7d85791396fd0c29ea47ac43b987af8f231c8049b

Observation a8f2c177-152c-478c-9e0c-153d6b4e4535 · outbound

This paper cites Omnimae: Single model masked pretraining on images and videos.

When Spatial meets Temporal in Action Recognition Omnimae: Single model masked pretraining on images and videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.828333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.776770Z digest=sha256:8afa5396e1d32ad6ad62e2e0d284e10ed0f8f3b378d0a789e134bbbccb91e6ea

Observation 6319f230-35b2-49b3-8bdc-599f5ef84de3 · outbound

This paper cites Deep residual learning for image recognition.

When Spatial meets Temporal in Action Recognition Deep residual learning for image recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.808952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.782024Z digest=sha256:50492536d41a9849903f03437adf72871093aaa925aed4bcb8d8c06732f9ef68

Observation 7bd43af2-3af5-48be-be47-c2d8a0929d2b · outbound

This paper cites Capturing temporal information in a sin- gle frame: Channel sampling strategies for action recogni- tion.

When Spatial meets Temporal in Action Recognition Capturing temporal information in a sin- gle frame: Channel sampling strategies for action recogni- tion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.790450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.788921Z digest=sha256:53c4abef6d74c09895ea5a2713383c8c6fd610a2cd6428d84fa9c76f8b94a9e9

Observation d11146d3-05cf-4aba-93b2-8b3de71e73d6 · outbound

This paper cites Tensor rep- resentations for action recognition.

When Spatial meets Temporal in Action Recognition Tensor rep- resentations for action recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.772452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.795857Z digest=sha256:d351c8872fca58422d77880884e7db84deb1c5a2db9ef629e63e7f3a68526f6a

Observation a40d2458-edb2-48e1-9f37-a26db7a3a507 · outbound

This paper cites HMDB: A large video database for human motion recognition.

When Spatial meets Temporal in Action Recognition HMDB: A large video database for human motion recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.752819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.802443Z digest=sha256:e7759c6cf6c7ca966f0084fc03ec95874fb84519a1c01c2f4c6e8cbc99357477

Observation a7aa6fd7-1c47-4cab-8c3f-c7df7282cc64 · outbound

This paper cites Action recognition based on a bag of 3D points.

When Spatial meets Temporal in Action Recognition Action recognition based on a bag of 3D points

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.733290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.808890Z digest=sha256:8d691badddb436153728715a535f90e3238303aad29cdeb606069ffbfede3ff3

Observation beff1ebc-d42b-4de5-adc8-14c2cf3d7667 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

When Spatial meets Temporal in Action Recognition Tsm: Temporal shift module for efficient video understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.716205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.814404Z digest=sha256:724c1d4dbe84369f63e52f723367925af73e9fbe1d4ab69aa151057b33029864

Observation 077af4aa-9942-4204-8dc4-7f875ce62157 · outbound

This paper cites an unresolved cited work.

When Spatial meets Temporal in Action Recognition Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:23.698227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.820261Z digest=sha256:3167a875e67944ae7fd60dda231f3d2abfcc2c31148393004c7ca9fba6f62527

Observation 521e680e-45c3-4c29-af10-3c510d94f40d · outbound

This paper cites HON4D: Histogram of Ori- ented 4D Normals for Activity Recognition from Depth Se- quences.

When Spatial meets Temporal in Action Recognition HON4D: Histogram of Ori- ented 4D Normals for Activity Recognition from Depth Se- quences

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.676614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.826912Z digest=sha256:90e7e550daa84e073b4a8e91d90d3d534c78685462f66a42dfeb1dd1dd199a15

Observation a6652d33-ebe9-453b-a712-8682b3ae1d5c · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers.

When Spatial meets Temporal in Action Recognition Keeping your eye on the ball: Tra- jectory attention in video transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.656906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.834276Z digest=sha256:9c46161b5e8868f9040af9e705cf8d1d6fcc9de38ec9ae65505c5012f0c05070

Observation 3209e556-d7fb-4dd9-a898-2e2936449f79 · outbound

This paper cites Huynh, and Ajmal Mian.

When Spatial meets Temporal in Action Recognition Huynh, and Ajmal Mian

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.636595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.839933Z digest=sha256:14b96bafde8c108fe82bd422eee3295b91435a8b47c41bea68f07ecaf2100db4

Observation 72e75139-eda6-4d8c-8753-f4fd73b94af8 · outbound

This paper cites Huynh, and Aj- mal Mian.

When Spatial meets Temporal in Action Recognition Huynh, and Aj- mal Mian

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.618054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.845346Z digest=sha256:2d573453235a225162af4368017c958943495755b642c3cd298040aa112dadea

Observation b9c8b198-20d0-4b5f-ac45-73c655a78bf5 · outbound

This paper cites Ryoo, AJ Piergiovanni, Mingxing Tan, and Anelia Angelova.

When Spatial meets Temporal in Action Recognition Ryoo, AJ Piergiovanni, Mingxing Tan, and Anelia Angelova

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.596967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.852274Z digest=sha256:3483a15970860aa7d026ceaee4b70feefbe1a2e01e157e96380b34738a65705c

Observation 12fe783b-2283-48b3-b778-fe4244df28df · outbound

This paper cites Ntu rgb+ d: A large scale dataset for 3d human activity anal- ysis.

When Spatial meets Temporal in Action Recognition Ntu rgb+ d: A large scale dataset for 3d human activity anal- ysis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.576225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.858923Z digest=sha256:8b34a0a20656f51e0fb9be2d45ce7d41fb3430feebd84987034d33756da38662

Observation 8e6ac195-9d8d-42e9-a355-11be3fd158df · outbound

This paper cites Two-stream con- volutional networks for action recognition in videos.

When Spatial meets Temporal in Action Recognition Two-stream con- volutional networks for action recognition in videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.864152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.864152Z digest=sha256:1a14182b51b7f5f1101b38cfa24c13f58afb1ef6ff499770d59ef1e250b4db30

Observation 02c5f08e-d4ac-4961-bd7a-dcb88dc3ae27 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

When Spatial meets Temporal in Action Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.869058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.869058Z digest=sha256:540d8f76bf079076fa9a23723fa4701798b027f4e5dd9bbdd5449fe1bd2d44c4

Observation 705caf59-eff9-49d8-bc1d-7a1a8ec428f7 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

When Spatial meets Temporal in Action Recognition Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.535506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.875206Z digest=sha256:94d0bce0ca56a647ac918fb3be369f382c4f7abbf3173d5a177adff9d910e882

Observation 062eaa9b-20c1-4291-aa6b-1cb294865143 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

When Spatial meets Temporal in Action Recognition Learning spatiotemporal features with 3d convolutional networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.880408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.880408Z digest=sha256:74f849470969bee5fe78d41d6a513a4c55b3dcb31ed72b4211e49e830fbc5a99

Observation 40e4e705-437c-468e-a3c3-7ee08eb6b68a · outbound

This paper cites Self-supervising action recog- nition by statistical moment and subspace descriptors.

When Spatial meets Temporal in Action Recognition Self-supervising action recog- nition by statistical moment and subspace descriptors

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.500783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.886875Z digest=sha256:5c1ed347ab59023c4dd57023ea23f138efc4a5f7e9609c3f7e032fb8570357be

Observation 0faad10b-1ea4-44b6-9346-a0189581e8f2 · outbound

This paper cites Flow dynamics correction for action recognition.

When Spatial meets Temporal in Action Recognition Flow dynamics correction for action recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.476392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.893095Z digest=sha256:79b652e15e857b01b3819228cb3a34670119debf7d3566e3619c9526bb5539b8

Observation a99fad45-19bb-4c8a-a1ae-a9efa0716ccd · outbound

This paper cites Temporal segment net- works: Towards good practices for deep action recognition.

When Spatial meets Temporal in Action Recognition Temporal segment net- works: Towards good practices for deep action recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.452893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.899022Z digest=sha256:b4f3b5e39c1710e254faebce7b3322a92592a4d4c7d95e73c5af376e6a672b2d

Observation 9c5033ea-9244-4bc2-a98c-27377ba3ffd1 · outbound

This paper cites Hallucinating idt descriptors and i3d optical flow features for action recog- nition with cnns.

When Spatial meets Temporal in Action Recognition Hallucinating idt descriptors and i3d optical flow features for action recog- nition with cnns

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.227789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.905943Z digest=sha256:8c76d4a234e7c20f0bd67fc1db596bcc3ba01d50959775b9c594e96d0d825344

Observation e99828b2-32ce-40c7-ac3d-435e74cb72ab · outbound

This paper cites Temporal segment networks for action recognition in videos.

When Spatial meets Temporal in Action Recognition Temporal segment networks for action recognition in videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.202047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.911650Z digest=sha256:45f9b09cea4af94f685dd34edc4d0d89c04fe8d6d6714e8b9006e1d09b075f9a

Observation 4cc24b11-4211-488e-8f5b-7b0e7d2ef18f · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

When Spatial meets Temporal in Action Recognition Videomae v2: Scaling video masked autoencoders with dual masking

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.178388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.917822Z digest=sha256:922d4c020f5449140ccd4900da938fd05045ace36a1bb8664cd385ebb7980082

Observation cde2db38-350b-49d9-9e53-9bfa1f4b2709 · outbound

This paper cites High-order tensor pooling with attention for action recognition.

When Spatial meets Temporal in Action Recognition High-order tensor pooling with attention for action recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.159005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.924260Z digest=sha256:ab1f0d9dc1d89d82689efbeec4fce76dd507020df609faad57bd730218f45b6e

Observation 8ea8f59a-aefc-4026-8c95-05b39ff3e1bd · outbound

This paper cites Taylor videos for action recognition.

When Spatial meets Temporal in Action Recognition Taylor videos for action recognition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.137389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.929422Z digest=sha256:a100f865b9579643d7823e42f31477703d84ac830a550287df3ffbce78458801

Observation 250c1991-3a7d-43e6-a8a0-d9def199b52f · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

When Spatial meets Temporal in Action Recognition Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.115681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.934251Z digest=sha256:64c82ddf9bdf78f6da73c82467e4ff7498827caedff450c4b1dba9d922c6a16f

Observation 79bb667e-1c8a-4c09-b716-6f810066c855 · outbound

This paper cites Towards good practices for missing modality robust action recognition.

When Spatial meets Temporal in Action Recognition Towards good practices for missing modality robust action recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:23.089306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.939733Z digest=sha256:b3e818a9a624215d4443d61d09881f2f6d955b71f07a36b096e8210bc13e61d2

Observation 36891865-dbf9-4bc1-8cbc-103fa35f04f4 · outbound

This paper cites Temporal pyramid network for action recognition.

When Spatial meets Temporal in Action Recognition Temporal pyramid network for action recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:22.945017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:22.945017Z digest=sha256:800ff13ac0b649a31f643ce1fd7035ad74a121e1b7bd71c612647cd947b3fa74

Observation 862c278e-9696-404b-a4e6-4c40eacfd801 · outbound

This paper cites Advancing video anomaly detection: A concise re- view and a new dataset.

When Spatial meets Temporal in Action Recognition Advancing video anomaly detection: A concise re- view and a new dataset

Reference 41

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:37:23.043623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T14:37:22.950052Z digest=sha256:9c36a03d8f932ef1b58350d5a4ec637b89021c92b6abfde8650105797a655641

Pith citing papers

Observation 90cbbc31-e234-4a38-9c94-fefe1dcff484 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? When Spatial meets Temporal in Action Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.062406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.062406Z digest=sha256:6ee86f8bf491e7121bf9b7f693a66885172e0354d765f0ab684e55669b037525

Observation 8c0037c5-c705-4285-bb38-641065632d4a · inbound

Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight cites this paper.

Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight When Spatial meets Temporal in Action Recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:52:58.194513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:52:58.194513Z digest=sha256:8357dc6e622992d6704d6b2ac11742dbc2ec53431626fbcec282d4ba58080d1c

Observation c045b02e-4aad-4f33-b797-c7fb7cede33d · inbound

Evolving Skeletons: Motion Dynamics in Action Recognition cites this paper.

Evolving Skeletons: Motion Dynamics in Action Recognition When Spatial meets Temporal in Action Recognition

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:13:04.038297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T22:12:58.294767Z digest=sha256:99ec1f98ed1ba080b26e1a99a125dbcb02710bc372b0c395219e7be16b00b91a