Pith. sign in

REVIEW 6 cited by

A Closer Look at Spatiotemporal Convolutions for Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.11248 v3 pith:DD2UQCW3 submitted 2017-11-30 cs.CV

classification cs.CV
keywords cnnsactionrecognitionspatiotemporalaccuracyadvantagesconvolutionalconvolutions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition. Our motivation stems from the observation that 2D CNNs applied to individual frames of the video have remained solid performers in action recognition. In this work we empirically demonstrate the accuracy advantages of 3D CNNs over 2D CNNs within the framework of residual learning. Furthermore, we show that factorizing the 3D convolutional filters into separate spatial and temporal components yields significantly advantages in accuracy. Our empirical study leads to the design of a new spatiotemporal convolutional block "R(2+1)D" which gives rise to CNNs that achieve results comparable or superior to the state-of-the-art on Sports-1M, Kinetics, UCF101 and HMDB51.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 215 citations worldwide. Full citation record

  1. Prithvi-Precip: Integrating Satellite Observations into an Atmospheric AI Foundation Model for Precipitation Forecasting

    physics.ao-ph 2026-08 conditional novelty 6.0 of 10

    Training on IMERG satellite precipitation and directly ingesting satellite observations improves medium-range AI precipitation forecasts, with the largest gains at lead times under 40 hours.

  2. The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction

    cs.CV 2026-03 conditional novelty 6.0 of 10

    In a cross-continental benchmark, breast tumor segmentation generalized reasonably across institutions, but predicting pathologic complete response from pre-treatment DCE-MRI alone was no better than random for nearly...

  3. SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SPACT18 is claimed to be the first action recognition dataset captured with a spike camera, paired with synchronized RGB and thermal video.

  4. Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new video dataset with per-frame bird species and behavior annotations, covering 13 species and 7 behaviors, is released with baseline recognition results.

  5. Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines

    cs.LG 2026-08 conditional novelty 5.0 of 10

    An equipment-centric vision system tracks cranes with fixed cameras and finite state machines, inferring hot-forging workpiece locations with 317.8 mm mean error and 100% event detection within a 33-second tolerance.

  6. Spatiotemporal Analysis of Forest Machine Operations Using 3D Video Classification

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A 3D ResNet-50 video classifier trained on a small dashcam dataset reaches 0.88 validation F1 for four forestry work elements, with acknowledged overfitting and limited data.

Pith tools