REVIEW 2 cited by
Benchmarking the Robustness of Spatial-Temporal Models Against Corruptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The state-of-the-art deep neural networks are vulnerable to common corruptions (e.g., input data degradations, distortions, and disturbances caused by weather changes, system error, and processing). While much progress has been made in analyzing and improving the robustness of models in image understanding, the robustness in video understanding is largely unexplored. In this paper, we establish a corruption robustness benchmark, Mini Kinetics-C and Mini SSV2-C, which considers temporal corruptions beyond spatial corruptions in images. We make the first attempt to conduct an exhaustive study on the corruption robustness of established CNN-based and Transformer-based spatial-temporal models. The study provides some guidance on robust model design and training: Transformer-based model performs better than CNN-based models on corruption robustness; the generalization ability of spatial-temporal models implies robustness against temporal corruptions; model corruption robustness (especially robustness in the temporal domain) enhances with computational cost and model capacity, which may contradict the current trend of improving the computational efficiency of models. Moreover, we find the robustness intervention for image-related tasks (e.g., training models with noise) may not work for spatial-temporal models.
Forward citations
Cited by 2 Pith papers
-
EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
A new egocentric safety benchmark shows current video-language models can describe scenes well but fail at multi-step causal reasoning about blind spots and covert actions.
-
SHARDeg: A Benchmark for Skeletal Human Action Recognition in Degraded Scenarios
A new benchmark degrades NTU-120 skeleton data three ways, shows degradation type strongly affects accuracy, and finds LogSigRNN overtakes DeGCN at 3 FPS once missing frames are interpolated.
Discussion (0). Continue with ORCID to comment.