REVIEW 4 major objections 5 minor 53 references
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Temporal action localization improves when classification and localization are fed features from different pyramid layers, and the paper reports state-of-the-art mAP on five benchmarks with this design.
desk verdict Solid TAL architecture paper with a suspicious MultiTHUMOS SOTA margin that needs verification before the five-benchmark claim is trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The CLTDR decoder is the central mechanism: at each intermediate pyramid layer $l$, a transposed-convolution-upsampled feature $P^{l+1}$ is gated by an attention weight $W_{l+1}$ and added to $P^l$ to form the classification feature $f^l_{cls}$, while a downsampled $P^{l-1}$ is gated by $W_{l-1}$ and added to form the regression feature $f^l_{reg}$ (Equations 5-8). The RefineHead then fuses $P^{l-1}$, $P^l$, and $P^{l+1}$ into $f^l_c$ and predicts a classification scaling vector $R^l_{co}$ and boundary offsets $R^l_{so}$ and $R^l_{eo}$ used to adjust the coarse predictions. The GMG encoder is the supporting object: a three-branch module where a fully-connected branch captures instant information, a 1D depth-wise convolution captures local context, and an FFT branch with a learnable global filter $W_\varphi$ captures global dependencies, with a ReLU-gated fusion to suppress redundancy.
What would settle it
Run the CLTDR ablation on a dataset of semantically defined actions (e.g., meetings or assembly steps) using the same features: if classification improves more when fed the lower-layer fine feature (or regression improves more with the higher-layer coarse feature), or if replacing the cross-layer input with a same-layer feature matches the reported gains, the claimed feature-utility gradient fails. A cheaper check is the paper's own Table 9, which already shows that adding both higher and lower features can slightly reduce mAP.
Extended reading notes
Core claim
The central claim is that the conventional practice of giving classification and localization heads identical input features is a bottleneck, and that the fix is to route each task to the pyramid layer whose feature scale suits it. For classification, the paper combines each layer's feature with the temporally coarser, semantically richer feature from the layer above, selected by a learned attention weight; for regression, it combines with the temporally finer, boundary-detailed feature from the layer below. The decoupled outputs are then refined by a RefineHead that fuses higher and lower layer features and produces multiplicative classification adjustments and additive boundary offsets. On five benchmarks the full system reports top average mAP, and ablations attribute the gains to both the cross-layer routing and the multi-granularity encoder. The paper further claims the FFT-based global branch matches vanilla self-attention accuracy at roughly one-third the parameters.
Load-bearing premise
The load-bearing premise is that at every pyramid layer, the immediately higher layer's coarser feature always helps classification and the immediately lower layer's finer feature always helps localization; if that ordering is wrong for some action distribution, the cross-layer routing can hurt rather than help.
Editorial extensions
If this is right
- Classification and localization heads in TAL no longer need to consume identical features; feature-level decoupling becomes a transferable design choice.
- On THUMOS14, CLTDR-GMG reports average mAP 69.9% with I3D features, 71.8% with VideoMAEv2, and 74.3% with InterVideo2-6B, each above prior state of the art; on MultiTHUMOS the advantage over TriDet grows to roughly 6-7 average mAP points.
- The GMG encoder can be dropped into other TAL decoders: it improves ActionFormer's decoder by 0.8 and TriDet's by 0.4 average mAP under the paper's settings, while the CLTDR decoder improves ActionFormer, TemporalMaxer, TriDet, and ActionMamba encoders by 0.6-0.8.
- An FFT-based global filter with a learnable filter achieves accuracy comparable to vanilla self-attention at 9.5M vs 26.1M parameters in the GMG comparison.
- Because CLTDR is applied only to intermediate pyramid layers, adding more layers eventually hurts; the paper observes performance peaks at 6 layers on three datasets and 7 layers on two.
Reading between the lines
- The universal feature-utility gradient (higher = semantic, lower = boundary) is assumed per layer; a natural test would be to swap the routing on datasets where boundaries are semantically defined, and the paper's own Table 9 shows that adding both cross-layer inputs can slightly reduce mAP, suggesting the two routes carry partially redundant information.
- The gains from the FFT global branch suggest that other token-mixing operators (e.g., Mamba-style state spaces) could be plugged into GMG, though the paper does not test them.
- The method's sensitivity to feature quality (I3D vs VideoMAEv2 vs InterVideo2-6B) indicates that part of the reported SOTA is inherited from stronger pretrained features; a controlled comparison with fixed features would isolate the architectural contribution.
- Since RefineHead only sees the immediate neighbors of a layer, using a wider or learnable cross-layer span might further improve long-range boundary consistency, but the paper does not explore it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CLTDR-GMG, a one-stage temporal action localization model. GMG is an encoder module that aggregates instant, local, and global temporal information using pointwise/depthwise convolutions and an FFT-based global filter with a gating mechanism. CLTDR is a decoder that, at each pyramid layer, combines the current layer with the next higher layer's semantically coarse feature for classification and with the next lower layer's finer feature for regression, then refines both outputs through a RefineHead. The method is evaluated on THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS, and the authors report state-of-the-art average mAP on all five benchmarks, with ablations on THUMOS14 using InterVideo2-6B features.
Significance. If substantiated, the paper makes a useful empirical contribution: it shows a convolutional/FFT encoder can replace self-attention in a TAL decoder at lower parameter cost (Table 8), and it provides a clean study of cross-layer feature selection for task decoupling. The release of code and pre-trained models, the plug-and-play experiments with ActionFormer/TriDet encoders/decoders (Supplementary E.1–E.2), and the error diagnosis are strengths. However, the headline claim rests on small margins on most benchmarks and on a single large MultiTHUMOS margin that is not yet explained; the core design assumption is only ablated on one dataset. The contribution is therefore promising but needs verification before the SOTA claim can be accepted.
major comments (4)
- [MultiTHUMOS (Table 2)] Table 2 and the 'MultiTHUMOS' paragraph: the gains over TriDet are 6.4, 6.5, and 7.4 average mAP for I3D RGB, I3D RGB+Flow, and VideoMAEv2 features, respectively, whereas the same modules on THUMOS14 (Table 1) improve over TriDet by only 0.6, 1.7, and 1.3 average mAP. The manuscript does not establish that the quoted TriDet and ActionFormer rows for MultiTHUMOS were produced with identical feature files, temporal strides, NMS, and evaluation scripts; the statement in Supplementary Section D that feature extraction is 'consistent with THUMOS14' concerns the authors' own features and does not address baseline parity. Please provide a protocol-matched comparison (e.g., rerun TriDet/ActionFormer under the exact same pipeline), report per-threshold numbers, and either explain the outlier or remove MultiTHUMOS from the SOTA claim.
- [Ablation of CLTDR (Table 9)] The CLTDR design in Eqs. (5)–(8) assumes a universal feature-utility gradient: higher-layer features always help classification and lower-layer features always help regression. The only ablation for this choice, Table 9 on THUMOS14 with InterVideo2-6B, does not fully support that universality: adding both higher and lower layers can slightly degrade performance (e.g., regression mAP 73.2 with all three sources versus 74.3 for the best two-source configuration), and the beneficial direction is not shown to transfer to other datasets/features. Please add ablations in at least one additional dataset/feature setting and report variance, or revise the claim to describe CLTDR as a dataset-tuned design rather than a universal principle.
- [Tables 1, 4, 5 and Implementation Details] The margins that support the SOTA claim are small on several benchmarks: 0.6 average mAP on THUMOS14 with I3D (Table 1), 0.3 on ActivityNet-1.3 (Table 4), and 0.4 on HACS with I3D (Table 5). No standard deviations or multiple seeds are reported anywhere, and for the largest claimed gains (MultiTHUMOS) the protocol mismatch described above applies. Please report mean±std over at least three seeds for the proposed method and for the closest baselines, and state whether the comparison rows are from the authors' runs or published numbers.
- [Experimental protocol] Several comparisons are based on the authors' own re-implementations (Table 1 marks TemporalMaxer § and TriDet § with ∓; Supplementary Tables 12–13 are entirely re-implementations), but the main tables do not indicate which rows on MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet, and HACS are re-implemented, nor are the re-implementation details (hyper-parameters, feature preprocessing, NMS) provided. To make the comparisons reproducible and fair, please mark all re-implemented rows and provide configuration files or a link to the code with the exact evaluation protocol.
minor comments (5)
- [Eq. (1)] The FFT formula uses N in the exponent but the summation is over T, and the range '0 < u < T−1' should be '0 ≤ u < T'; please fix the notation.
- [Eq. (4)] The term ReLU(flocal)⊗FC(x) is labeled finstant and ReLU(fglobal)⊗Conv(x) is labeled flocal; these labels appear swapped relative to the text, please check.
- [HACS results] The text says the method achieves 39.2% with SlowFast features, while Table 5 reports 39.3; please align the numbers.
- [Supplementary tables and text] Table 14 header contains 'PIC-KITCHENS-100', which should be 'EPIC-KITCHENS-100', and Supplementary Section C contains the typo 'respecviely'.
- [Supplementary Section E.1] The phrase 'This can be attribute to' should be 'This can be attributed to'.
Circularity Check
No circularity: the SOTA claim is an external benchmark comparison and the module designs are validated by controlled ablations, not by fitting, definitional identity, or load-bearing self-citation.
full rationale
I walked the paper's derivation chain and found no step that reduces to its own inputs. The GMG encoder is defined by explicit equations (Eqs. 1-4) that transform input features via FFT, depth-wise convolution, and gated fusion; the CLTDR module is defined by Eqs. 5-12, which combine pyramid layers with learned attention weights and refine the resulting classification and regression outputs. These are new architectural constructions, not definitions that presuppose the claimed performance. The central claim, state-of-the-art results on five benchmarks, is an empirical comparison against externally published baselines, and no reported test number is used to set a constant in the model; per-dataset hyperparameters in the supplementary material are ordinary training choices, not fitted inputs renamed as predictions. The ablation studies (Tables 6-11 and supplementary Tables 12-13) support the efficacy claims by controlled comparison of variants, rather than by assuming the modules work. The only self-citation I could identify is Chen et al. 2021 (DDOD), whose author list includes Qiang Li; it appears in Related Work as background on task decoupling in object detection and is not a load-bearing premise for the present method. There is no uniqueness theorem imported from the authors, no ansatz smuggled in via citation, and no known empirical result merely renamed. Any concern about the MultiTHUMOS comparison protocol would be a correctness or evaluation-fairness issue, not circularity. The derivation chain is therefore self-contained with respect to circularity.
Assumptions & free parameters
free parameters (3)
- Feature pyramid depth L =
6 (THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100), 7 (ActivityNet-1.3, HACS)
- Trident-head bins B =
16, 16, 16, 12, 14 for THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, HACS
- Per-dataset training hyperparameters =
learning rates 1e-4 to 1e-3, epochs 15 to 62, warmups 10 to 32, weight decay 0.03 to 0.055 depending on dataset and…
assumptions (4)
- domain assumption The extracted features from pretrained backbones (I3D, VideoMAEv2, InterVideo2, SlowFast, TSP) are sufficient input representations for temporal action localization.
- domain assumption Classification is better served by semantically strong, temporally coarse higher-pyramid features, while localization is better served by detailed lower-pyramid features.
- standard math Hadamard product in the Fourier domain corresponds to circular convolution in the time domain, so the FFT-based global branch can be interpreted as a convolution.
- domain assumption The relative boundary modeling in the Trident-head from TriDet is a valid component for regression.
Cite this review
Pith. "Pith review of Temporal Action Localization with Cross Layer Task Decoupling and Refinement." pith.science (2026). https://pith.science/paper/VVP33BXV
@misc{pith2026241209202,
author = {Pith},
title = {Pith review of: Temporal Action Localization with Cross Layer Task Decoupling and Refinement},
year = {2026},
howpublished = {\url{https://pith.science/paper/VVP33BXV}},
note = {Machine review of arXiv:2412.09202}
}
read the original abstract
Temporal action localization (TAL) involves dual tasks to classify and localize actions within untrimmed videos. However, the two tasks often have conflicting requirements for features. Existing methods typically employ separate heads for classification and localization tasks but share the same input feature, leading to suboptimal performance. To address this issue, we propose a novel TAL method with Cross Layer Task Decoupling and Refinement (CLTDR). Based on the feature pyramid of video, CLTDR strategy integrates semantically strong features from higher pyramid layers and detailed boundary-aware boundary features from lower pyramid layers to effectively disentangle the action classification and localization tasks. Moreover, the multiple features from cross layers are also employed to refine and align the disentangled classification and regression results. At last, a lightweight Gated Multi-Granularity (GMG) module is proposed to comprehensively extract and aggregate video features at instant, local, and global temporal granularities. Benefiting from the CLTDR and GMG modules, our method achieves state-of-the-art performance on five challenging benchmarks: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS. Our code and pre-trained models are publicly available at: https://github.com/LiQiang0307/CLTDR-GMG.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Alwassel, H.; Giancola, S.; and Ghanem, B. 2021. TSP: Temporally-Sensitive Pretraining of Video Encoders for Localization Tasks. In IEEE/CVF International Conference on Computer Vision Workshops, ICCVW 2021, Montreal, BC, Canada, October 11-17, 2021 , 3166--3176. IEEE
work page 2021
-
[4]
C.; Escorcia, V.; and Ghanem, B
Alwassel, H.; Heilbron, F. C.; Escorcia, V.; and Ghanem, B. 2018. Diagnosing Error in Temporal Action Detectors. In Ferrari, V.; Hebert, M.; Sminchisescu, C.; and Weiss, Y., eds., Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part III , volume 11207 of Lecture Notes in Computer Science, 264--28...
work page 2018
-
[5]
Bai, Y.; Wang, Y.; Tong, Y.; Yang, Y.; Liu, Q.; and Liu, J. 2020. Boundary Content Graph Neural Network for Temporal Action Proposal Generation. In Vedaldi, A.; Bischof, H.; Brox, T.; and Frahm, J., eds., Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXVIII , volume 12373 of Lecture Notes in Com...
work page 2020
-
[6]
Bodla, N.; Singh, B.; Chellappa, R.; and Davis, L. S. 2017. Soft-NMS - Improving Object Detection with One Line of Code. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 , 5562--5570. IEEE Computer Society
work page 2017
-
[7]
Buch, S.; Escorcia, V.; Ghanem, B.; Fei - Fei, L.; and Niebles, J. C. 2017. End-to-End, Single-Stream Temporal Action Detection in Untrimmed Videos. In British Machine Vision Conference 2017, BMVC 2017, London, UK, September 4-7, 2017 . BMVA Press
work page 2017
-
[8]
Carreira, J.; and Zisserman, A. 2017. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 , 4724--4733. IEEE Computer Society
work page 2017
Show all 53 references
-
[9]
Chen, G.; Huang, Y.; Xu, J.; Pei, B.; Chen, Z.; Li, Z.; Wang, J.; Li, K.; Lu, T.; and Wang, L. 2024. Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding. arXiv:2403.09626
2024 arXiv
-
[10]
Chen, Z.; Yang, C.; Li, Q.; Zhao, F.; Zha, Z.; and Wu, F. 2021. Disentangle Your Dense Object Detector. In Shen, H. T.; Zhuang, Y.; Smith, J. R.; Yang, Y.; C \' e sar, P.; Metze, F.; and Prabhakaran, B., eds., MM '21: ACM Multimedia Conference, Virtual Event, China, October 20...
2021
-
[11]
Cheng, F.; and Bertasius, G. 2022. TallFormer: Temporal Action Localization with a Long-Memory Transformer. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October...
2022
-
[12]
M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M
Damen, D.; Doughty, H.; Farinella, G. M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M. 2022. Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100 . Int. J. Comput. Vis., 130(1): 33--55
2022
-
[13]
Dong, W.; Zhang, Z.; Song, C.; and Tan, T. 2022. Identifying the key frames: An attention-aware sampling method for action recognition. Pattern Recognition, 130: 108797
2022
-
[14]
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019. SlowFast Networks for Video Recognition. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 , 6201--6210. IEEE
2019
-
[15]
Ge, Z.; Liu, S.; Wang, F.; Li, Z.; and Sun, J. 2021. YOLOX: Exceeding YOLO Series in 2021. arXiv:2107.08430
2021 arXiv
-
[16]
Gu, A.; and Dao, T. 2024. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752
2024 arXiv
-
[17]
Guibas, J.; Mardani, M.; Li, Z.; Tao, A.; Anandkumar, A.; and Catanzaro, B. 2021. Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators. In International Conference on Learning Representations
2021
-
[18]
C.; Escorcia, V.; Ghanem, B.; and Niebles, J
Heilbron, F. C.; Escorcia, V.; Ghanem, B.; and Niebles, J. C. 2015. ActivityNet: A large-scale video benchmark for human activity understanding. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015 , 961--970. IEEE Computer Society
2015
-
[19]
C.; Niebles, J
Heilbron, F. C.; Niebles, J. C.; and Ghanem, B. 2016. Fast Temporal Activity Proposals for Efficient Detection of Human Actions in Untrimmed Videos. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 , 1914--1923...
2016
-
[20]
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; Suleyman, M.; and Zisserman, A. 2017. The Kinetics Human Action Video Dataset. arXiv:1705.06950
2017 arXiv
-
[21]
Lee - Thorp, J.; Ainslie, J.; Eckstein, I.; and Onta \ n \' o n, S. 2022. FNet: Mixing Tokens with Fourier Transforms. In Carpuat, M.; de Marneffe, M.; and Ru \' z, I. V. M., eds., Proceedings of the 2022 Conference of the North American Chapter of the Association for Computat...
2022
-
[22]
Lin, C.; Xu, C.; Luo, D.; Wang, Y.; Tai, Y.; Wang, C.; Li, J.; Huang, F.; and Fu, Y. 2021. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , 3320...
2021
-
[23]
Lin, T.; Zhao, X.; and Shou, Z. 2017. Single Shot Temporal Action Detection. In Liu, Q.; Lienhart, R.; Wang, H.; Chen, S. K.; Boll, S.; Chen, Y. P.; Friedland, G.; Li, J.; and Yan, S., eds., Proceedings of the 2017 ACM on Multimedia Conference, MM 2017, Mountain View, CA, USA,...
2017
-
[24]
Lin, T.; Zhao, X.; Su, H.; Wang, C.; and Yang, M. 2018. BSN: Boundary Sensitive Network for Temporal Action Proposal Generation. In Ferrari, V.; Hebert, M.; Sminchisescu, C.; and Weiss, Y., eds., Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, Septembe...
2018
-
[25]
Liu, X.; Wang, Q.; Hu, Y.; Tang, X.; Zhang, S.; Bai, S.; and Bai, X. 2022. End-to-End Temporal Action Detection With Transformer. IEEE Trans. Image Process. , 31: 5427--5441
2022
-
[26]
Long, F.; Yao, T.; Qiu, Z.; Tian, X.; Luo, J.; and Mei, T. 2019. Gaussian Temporal Awareness Networks for Action Localization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , 344--353. Computer Vision Foundation / IEEE
2019
-
[27]
Nag, S.; Zhu, X.; Song, Y.; and Xiang, T. 2022. Semi-supervised Temporal Action Detection with Proposal-Free Masking. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israe...
2022
-
[28]
Oppenheim, A. V. 1999. Discrete-time signal processing. Pearson Education India
1999
-
[29]
Qing, Z.; Su, H.; Gan, W.; Wang, D.; Wu, W.; Wang, X.; Qiao, Y.; Yan, J.; Gao, C.; and Sang, N. 2021 a . Temporal Context Aggregation Network for Temporal Action Proposal Refinement. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25,...
2021
-
[30]
Qing, Z.; Su, H.; Gan, W.; Wang, D.; Wu, W.; Wang, X.; Qiao, Y.; Yan, J.; Gao, C.; and Sang, N. 2021 b . Temporal Context Aggregation Network for Temporal Action Proposal Refinement. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25,...
2021
-
[31]
Rao, Y.; Zhao, W.; Zhu, Z.; Zhou, J.; and Lu, J. 2023. GFNet: Global Filter Networks for Visual Recognition. IEEE Trans. Pattern Anal. Mach. Intell. , 45(9): 10960--10973
2023
-
[32]
D.; and Savarese, S
Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I. D.; and Savarese, S. 2019. Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-...
2019
-
[33]
Shi, D.; Zhong, Y.; Cao, Q.; Ma, L.; Lit, J.; and Tao, D. 2023. TriDet: Temporal Action Detection with Relative Boundary Modeling. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , 18857--18866. IEEE
2023
-
[34]
Shi, D.; Zhong, Y.; Cao, Q.; Zhang, J.; Ma, L.; Li, J.; and Tao, D. 2022. ReAct: Temporal Action Detection with Relational Queries. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, T...
2022
-
[35]
Shou, Z.; Wang, D.; and Chang, S. 2016. Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 , 1049--1058. IEEE Computer Society
2016
-
[36]
Tan, J.; Zhao, X.; Shi, X.; Kang, B.; and Wang, L. 2022. PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points. In NeurIPS
2022
-
[37]
N.; Kim, K.; and Sohn, K
Tang, T. N.; Kim, K.; and Sohn, K. 2023. TemporalMaxer: Maximize Temporal Context with only Max Pooling for Temporal Action Localization. arXiv:2303.09055
2023 arXiv
-
[38]
N.; Kaiser, L.; and Polosukhin, I
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In Guyon, I.; von Luxburg, U.; Bengio, S.; Wallach, H. M.; Fergus, R.; Vishwanathan, S. V. N.; and Garnett, R., eds., Advances in Neura...
2017
-
[39]
Wang, L.; Huang, B.; Zhao, Z.; Tong, Z.; He, Y.; Wang, Y.; Wang, Y.; and Qiao, Y. 2023. VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , 145...
2023
-
[40]
Wang, Y.; Li, K.; Li, X.; Yu, J.; He, Y.; Wang, C.; Chen, G.; Pei, B.; Yan, Z.; Zheng, R.; Xu, J.; Wang, Z.; Shi, Y.; Jiang, T.; Li, S.; Zhang, H.; Huang, Y.; Qiao, Y.; Wang, Y.; and Wang, L. 2024. InternVideo2: Scaling Foundation Models for Multimodal Video Understanding. arX...
2024 arXiv
-
[41]
Wu, Y.; Chen, Y.; Yuan, L.; Liu, Z.; Wang, L.; Li, H.; and Fu, Y. 2020. Rethinking classification and localization for object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10186--10195
2020
-
[42]
S.; Thabet, A
Xu, M.; Zhao, C.; Rojas, D. S.; Thabet, A. K.; and Ghanem, B. 2020. G-TAD: Sub-Graph Localization for Temporal Action Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020 , 10153--10162. Computer Visio...
2020
-
[43]
Z.; Qin, H.; Feng, M.; Zhang, L.; and Mian, A
Yan, X.; Gilani, S. Z.; Qin, H.; Feng, M.; Zhang, L.; and Mian, A. 2018. Deep Keyframe Detection in Human Action Videos. arXiv:1804.10021
2018 arXiv
-
[44]
Yang, J.; Wei, P.; Ren, Z.; and Zheng, N. 2024. Gated Multi-Scale Transformer for Temporal Action Localization. IEEE Transactions on Multimedia, 26: 5705--5717
2024
-
[45]
Yang, J.; Wei, P.; and Zheng, N. 2024. Cross Time-Frequency Transformer for Temporal Action Localization. IEEE Transactions on Circuits and Systems for Video Technology, 34(6): 4625--4638
2024
-
[46]
Yang, L.; Peng, H.; Zhang, D.; Fu, J.; and Han, J. 2020. Revisiting Anchor Mechanisms for Temporal Action Localization. IEEE Trans. Image Process. , 29: 8535--8548
2020
-
[47]
Yeung, S.; Russakovsky, O.; Jin, N.; Andriluka, M.; Mori, G.; and Fei - Fei, L. 2018 a . Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos. Int. J. Comput. Vis., 126(2-4): 375--389
2018
-
[48]
Yeung, S.; Russakovsky, O.; Jin, N.; Andriluka, M.; Mori, G.; and Fei - Fei, L. 2018 b . Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos. Int. J. Comput. Vis., 126(2-4): 375--389
2018
-
[49]
Zhang, C.; Wu, J.; and Li, Y. 2022. ActionFormer: Localizing Moments of Actions with Transformers. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2...
2022
-
[50]
Zhang, H.; Wang, Y.; Dayoub, F.; and Sunderhauf, N. 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8514--8523
2021
-
[51]
K.; and Ghanem, B
Zhao, C.; Thabet, A. K.; and Ghanem, B. 2021. Video Self-Stitching Graph Network for Temporal Action Localization. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021 , 13638--13647. IEEE
2021
-
[52]
Zhao, H.; Torralba, A.; Torresani, L.; and Yan, Z. 2019. HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 , 8667...
2019
-
[53]
Zhuang, J.; Qin, Z.; Yu, H.; and Chen, X. 2023. Task-Specific Context Decoupling for Object Detection. arXiv:2303.01047
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.