REVIEW 6 cited by
A Closer Look at Spatiotemporal Convolutions for Action Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition. Our motivation stems from the observation that 2D CNNs applied to individual frames of the video have remained solid performers in action recognition. In this work we empirically demonstrate the accuracy advantages of 3D CNNs over 2D CNNs within the framework of residual learning. Furthermore, we show that factorizing the 3D convolutional filters into separate spatial and temporal components yields significantly advantages in accuracy. Our empirical study leads to the design of a new spatiotemporal convolutional block "R(2+1)D" which gives rise to CNNs that achieve results comparable or superior to the state-of-the-art on Sports-1M, Kinetics, UCF101 and HMDB51.
Forward citations
Cited by 6 Pith papers
-
Prithvi-Precip: Integrating Satellite Observations into an Atmospheric AI Foundation Model for Precipitation Forecasting
Training on IMERG satellite precipitation and directly ingesting satellite observations improves medium-range AI precipitation forecasts, with the largest gains at lead times under 40 hours.
-
The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction
In a cross-continental benchmark, breast tumor segmentation generalized reasonably across institutions, but predicting pathologic complete response from pre-treatment DCE-MRI alone was no better than random for nearly...
-
SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
SPACT18 is claimed to be the first action recognition dataset captured with a spike camera, paired with synchronized RGB and thermal video.
-
Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos
A new video dataset with per-frame bird species and behavior annotations, covering 13 species and 7 behaviors, is released with baseline recognition results.
-
Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
An equipment-centric vision system tracks cranes with fixed cameras and finite state machines, inferring hot-forging workpiece locations with 317.8 mm mean error and 100% event detection within a 33-second tolerance.
-
Spatiotemporal Analysis of Forest Machine Operations Using 3D Video Classification
A 3D ResNet-50 video classifier trained on a small dashcam dataset reaches 0.88 validation F1 for four forestry work elements, with acknowledged overfitting and limited data.
Discussion (0). Continue with ORCID to comment.