Pith. sign in

REVIEW 4 major objections 5 minor 53 references

Temporal Action Localization with Cross Layer Task Decoupling and Refinement

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Temporal action localization improves when classification and localization are fed features from different pyramid layers, and the paper reports state-of-the-art mAP on five benchmarks with this design.

desk verdict Solid TAL architecture paper with a suspicious MultiTHUMOS SOTA margin that needs verification before the five-benchmark claim is trusted. read the letter →

arxiv 2412.09202 v2 pith:VVP33BXV submitted 2024-12-12 cs.CV

classification cs.CV
keywords temporalactionlocalizationcross-layertaskdecouplingfeaturepyramidgatedmulti-granularityFFTglobalfilterclassificationboundaryregressionvideounderstanding
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that temporal action localization improves when the classification and localization heads stop sharing the same input features. It proposes CLTDR, which feeds each pyramid layer's classification head with the semantically stronger feature from the next-higher layer and its regression head with the boundary-richer feature from the next-lower layer, gated by attention weights. A refinement head then fuses both cross-layer features to align the two outputs. Together with a lightweight Gated Multi-Granularity encoder that captures instant, local, and global temporal context via an FFT-based global filter, the method reports state-of-the-art average mAP on THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS. If correct, this shows that feature-level task decoupling, not just parameter-level head separation, is a reliable source of TAL gains.

What carries the argument

The CLTDR decoder is the central mechanism: at each intermediate pyramid layer $l$, a transposed-convolution-upsampled feature $P^{l+1}$ is gated by an attention weight $W_{l+1}$ and added to $P^l$ to form the classification feature $f^l_{cls}$, while a downsampled $P^{l-1}$ is gated by $W_{l-1}$ and added to form the regression feature $f^l_{reg}$ (Equations 5-8). The RefineHead then fuses $P^{l-1}$, $P^l$, and $P^{l+1}$ into $f^l_c$ and predicts a classification scaling vector $R^l_{co}$ and boundary offsets $R^l_{so}$ and $R^l_{eo}$ used to adjust the coarse predictions. The GMG encoder is the supporting object: a three-branch module where a fully-connected branch captures instant information, a 1D depth-wise convolution captures local context, and an FFT branch with a learnable global filter $W_\varphi$ captures global dependencies, with a ReLU-gated fusion to suppress redundancy.

What would settle it

Run the CLTDR ablation on a dataset of semantically defined actions (e.g., meetings or assembly steps) using the same features: if classification improves more when fed the lower-layer fine feature (or regression improves more with the higher-layer coarse feature), or if replacing the cross-layer input with a same-layer feature matches the reported gains, the claimed feature-utility gradient fails. A cheaper check is the paper's own Table 9, which already shows that adding both higher and lower features can slightly reduce mAP.

Watch

Extended reading notes

Core claim

The central claim is that the conventional practice of giving classification and localization heads identical input features is a bottleneck, and that the fix is to route each task to the pyramid layer whose feature scale suits it. For classification, the paper combines each layer's feature with the temporally coarser, semantically richer feature from the layer above, selected by a learned attention weight; for regression, it combines with the temporally finer, boundary-detailed feature from the layer below. The decoupled outputs are then refined by a RefineHead that fuses higher and lower layer features and produces multiplicative classification adjustments and additive boundary offsets. On five benchmarks the full system reports top average mAP, and ablations attribute the gains to both the cross-layer routing and the multi-granularity encoder. The paper further claims the FFT-based global branch matches vanilla self-attention accuracy at roughly one-third the parameters.

Load-bearing premise

The load-bearing premise is that at every pyramid layer, the immediately higher layer's coarser feature always helps classification and the immediately lower layer's finer feature always helps localization; if that ordering is wrong for some action distribution, the cross-layer routing can hurt rather than help.

Editorial extensions

If this is right

  • Classification and localization heads in TAL no longer need to consume identical features; feature-level decoupling becomes a transferable design choice.
  • On THUMOS14, CLTDR-GMG reports average mAP 69.9% with I3D features, 71.8% with VideoMAEv2, and 74.3% with InterVideo2-6B, each above prior state of the art; on MultiTHUMOS the advantage over TriDet grows to roughly 6-7 average mAP points.
  • The GMG encoder can be dropped into other TAL decoders: it improves ActionFormer's decoder by 0.8 and TriDet's by 0.4 average mAP under the paper's settings, while the CLTDR decoder improves ActionFormer, TemporalMaxer, TriDet, and ActionMamba encoders by 0.6-0.8.
  • An FFT-based global filter with a learnable filter achieves accuracy comparable to vanilla self-attention at 9.5M vs 26.1M parameters in the GMG comparison.
  • Because CLTDR is applied only to intermediate pyramid layers, adding more layers eventually hurts; the paper observes performance peaks at 6 layers on three datasets and 7 layers on two.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The universal feature-utility gradient (higher = semantic, lower = boundary) is assumed per layer; a natural test would be to swap the routing on datasets where boundaries are semantically defined, and the paper's own Table 9 shows that adding both cross-layer inputs can slightly reduce mAP, suggesting the two routes carry partially redundant information.
  • The gains from the FFT global branch suggest that other token-mixing operators (e.g., Mamba-style state spaces) could be plugged into GMG, though the paper does not test them.
  • The method's sensitivity to feature quality (I3D vs VideoMAEv2 vs InterVideo2-6B) indicates that part of the reported SOTA is inherited from stronger pretrained features; a controlled comparison with fixed features would isolate the architectural contribution.
  • Since RefineHead only sees the immediate neighbors of a layer, using a wider or learnable cross-layer span might further improve long-range boundary consistency, but the paper does not explore it.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes CLTDR-GMG, a one-stage temporal action localization model. GMG is an encoder module that aggregates instant, local, and global temporal information using pointwise/depthwise convolutions and an FFT-based global filter with a gating mechanism. CLTDR is a decoder that, at each pyramid layer, combines the current layer with the next higher layer's semantically coarse feature for classification and with the next lower layer's finer feature for regression, then refines both outputs through a RefineHead. The method is evaluated on THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS, and the authors report state-of-the-art average mAP on all five benchmarks, with ablations on THUMOS14 using InterVideo2-6B features.

Significance. If substantiated, the paper makes a useful empirical contribution: it shows a convolutional/FFT encoder can replace self-attention in a TAL decoder at lower parameter cost (Table 8), and it provides a clean study of cross-layer feature selection for task decoupling. The release of code and pre-trained models, the plug-and-play experiments with ActionFormer/TriDet encoders/decoders (Supplementary E.1–E.2), and the error diagnosis are strengths. However, the headline claim rests on small margins on most benchmarks and on a single large MultiTHUMOS margin that is not yet explained; the core design assumption is only ablated on one dataset. The contribution is therefore promising but needs verification before the SOTA claim can be accepted.

major comments (4)
  1. [MultiTHUMOS (Table 2)] Table 2 and the 'MultiTHUMOS' paragraph: the gains over TriDet are 6.4, 6.5, and 7.4 average mAP for I3D RGB, I3D RGB+Flow, and VideoMAEv2 features, respectively, whereas the same modules on THUMOS14 (Table 1) improve over TriDet by only 0.6, 1.7, and 1.3 average mAP. The manuscript does not establish that the quoted TriDet and ActionFormer rows for MultiTHUMOS were produced with identical feature files, temporal strides, NMS, and evaluation scripts; the statement in Supplementary Section D that feature extraction is 'consistent with THUMOS14' concerns the authors' own features and does not address baseline parity. Please provide a protocol-matched comparison (e.g., rerun TriDet/ActionFormer under the exact same pipeline), report per-threshold numbers, and either explain the outlier or remove MultiTHUMOS from the SOTA claim.
  2. [Ablation of CLTDR (Table 9)] The CLTDR design in Eqs. (5)–(8) assumes a universal feature-utility gradient: higher-layer features always help classification and lower-layer features always help regression. The only ablation for this choice, Table 9 on THUMOS14 with InterVideo2-6B, does not fully support that universality: adding both higher and lower layers can slightly degrade performance (e.g., regression mAP 73.2 with all three sources versus 74.3 for the best two-source configuration), and the beneficial direction is not shown to transfer to other datasets/features. Please add ablations in at least one additional dataset/feature setting and report variance, or revise the claim to describe CLTDR as a dataset-tuned design rather than a universal principle.
  3. [Tables 1, 4, 5 and Implementation Details] The margins that support the SOTA claim are small on several benchmarks: 0.6 average mAP on THUMOS14 with I3D (Table 1), 0.3 on ActivityNet-1.3 (Table 4), and 0.4 on HACS with I3D (Table 5). No standard deviations or multiple seeds are reported anywhere, and for the largest claimed gains (MultiTHUMOS) the protocol mismatch described above applies. Please report mean±std over at least three seeds for the proposed method and for the closest baselines, and state whether the comparison rows are from the authors' runs or published numbers.
  4. [Experimental protocol] Several comparisons are based on the authors' own re-implementations (Table 1 marks TemporalMaxer § and TriDet § with ∓; Supplementary Tables 12–13 are entirely re-implementations), but the main tables do not indicate which rows on MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet, and HACS are re-implemented, nor are the re-implementation details (hyper-parameters, feature preprocessing, NMS) provided. To make the comparisons reproducible and fair, please mark all re-implemented rows and provide configuration files or a link to the code with the exact evaluation protocol.
minor comments (5)
  1. [Eq. (1)] The FFT formula uses N in the exponent but the summation is over T, and the range '0 < u < T−1' should be '0 ≤ u < T'; please fix the notation.
  2. [Eq. (4)] The term ReLU(flocal)⊗FC(x) is labeled finstant and ReLU(fglobal)⊗Conv(x) is labeled flocal; these labels appear swapped relative to the text, please check.
  3. [HACS results] The text says the method achieves 39.2% with SlowFast features, while Table 5 reports 39.3; please align the numbers.
  4. [Supplementary tables and text] Table 14 header contains 'PIC-KITCHENS-100', which should be 'EPIC-KITCHENS-100', and Supplementary Section C contains the typo 'respecviely'.
  5. [Supplementary Section E.1] The phrase 'This can be attribute to' should be 'This can be attributed to'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the SOTA claim is an external benchmark comparison and the module designs are validated by controlled ablations, not by fitting, definitional identity, or load-bearing self-citation.

full rationale

I walked the paper's derivation chain and found no step that reduces to its own inputs. The GMG encoder is defined by explicit equations (Eqs. 1-4) that transform input features via FFT, depth-wise convolution, and gated fusion; the CLTDR module is defined by Eqs. 5-12, which combine pyramid layers with learned attention weights and refine the resulting classification and regression outputs. These are new architectural constructions, not definitions that presuppose the claimed performance. The central claim, state-of-the-art results on five benchmarks, is an empirical comparison against externally published baselines, and no reported test number is used to set a constant in the model; per-dataset hyperparameters in the supplementary material are ordinary training choices, not fitted inputs renamed as predictions. The ablation studies (Tables 6-11 and supplementary Tables 12-13) support the efficacy claims by controlled comparison of variants, rather than by assuming the modules work. The only self-citation I could identify is Chen et al. 2021 (DDOD), whose author list includes Qiang Li; it appears in Related Work as background on task decoupling in object detection and is not a load-bearing premise for the present method. There is no uniqueness theorem imported from the authors, no ansatz smuggled in via citation, and no known empirical result merely renamed. Any concern about the MultiTHUMOS comparison protocol would be a correctness or evaluation-fairness issue, not circularity. The derivation chain is therefore self-contained with respect to circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The method is a neural architecture; the central claim rests on the sufficiency of frozen pretrained features, the cross-layer feature-utility assumption for classification versus localization, standard Fourier convolution equivalence, and the borrowed Trident-head regression component. No new physical or conceptual entities are postulated. The listed free parameters are ordinary per-dataset hyperparameters and architecture choices.

free parameters (3)
  • Feature pyramid depth L = 6 (THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100), 7 (ActivityNet-1.3, HACS)
    Selected by validation performance; Table 14 shows the chosen values are optimal in the evaluated sweep.
  • Trident-head bins B = 16, 16, 16, 12, 14 for THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, HACS
    Set following TriDet and tuned per dataset in Supplementary C.
  • Per-dataset training hyperparameters = learning rates 1e-4 to 1e-3, epochs 15 to 62, warmups 10 to 32, weight decay 0.03 to 0.055 depending on dataset and…
    Tuned on each dataset and feature type; these choices affect the reported mAP but are standard optimization tuning.
assumptions (4)
  • domain assumption The extracted features from pretrained backbones (I3D, VideoMAEv2, InterVideo2, SlowFast, TSP) are sufficient input representations for temporal action localization.
    Invoked throughout the Method and Experiments; the entire pipeline operates on frozen feature sequences and does not learn from raw video.
  • domain assumption Classification is better served by semantically strong, temporally coarse higher-pyramid features, while localization is better served by detailed lower-pyramid features.
    This is the design premise of the decoupled classification and regression modules (Eqs. 5-8); the paper justifies it with examples and ablations on THUMOS14, but assumes it transfers to all five benchmarks.
  • standard math Hadamard product in the Fourier domain corresponds to circular convolution in the time domain, so the FFT-based global branch can be interpreted as a convolution.
    Invoked after Eq. 3, citing Oppenheim 1999; this is standard signal processing.
  • domain assumption The relative boundary modeling in the Trident-head from TriDet is a valid component for regression.
    The DRM adopts the Trident-head from Shi et al. 2023 rather than deriving it; the paper relies on that component's correctness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Temporal Action Localization with Cross Layer Task Decoupling and Refinement." pith.science (2026). https://pith.science/paper/VVP33BXV

@misc{pith2026241209202,
  author       = {Pith},
  title        = {Pith review of: Temporal Action Localization with Cross Layer Task Decoupling and Refinement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VVP33BXV}},
  note         = {Machine review of arXiv:2412.09202}
}
read the original abstract

Temporal action localization (TAL) involves dual tasks to classify and localize actions within untrimmed videos. However, the two tasks often have conflicting requirements for features. Existing methods typically employ separate heads for classification and localization tasks but share the same input feature, leading to suboptimal performance. To address this issue, we propose a novel TAL method with Cross Layer Task Decoupling and Refinement (CLTDR). Based on the feature pyramid of video, CLTDR strategy integrates semantically strong features from higher pyramid layers and detailed boundary-aware boundary features from lower pyramid layers to effectively disentangle the action classification and localization tasks. Moreover, the multiple features from cross layers are also employed to refine and align the disentangled classification and regression results. At last, a lightweight Gated Multi-Granularity (GMG) module is proposed to comprehensively extract and aggregate video features at instant, local, and global temporal granularities. Benefiting from the CLTDR and GMG modules, our method achieves state-of-the-art performance on five challenging benchmarks: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3, and HACS. Our code and pre-trained models are publicly available at: https://github.com/LiQiang0307/CLTDR-GMG.

Figures

Figures reproduced from arXiv: 2412.09202 by the authors.

Figure 1
Figure 1. Comparison of different task decoupling. (a) Pre [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. An illustration of our method. We build a feature pyramid with GMG module. The CLTDR decoder at the [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of GMG. Zoom in for better view. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Illustration of Decoupled Classification Module and Decoupled Regression Module. [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: The visual comparison between CLTDR-GMG and other two methods before NMS. Each point denotes the classifi [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: The false positive profile obtained by our methods and the impact of error types. [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: The average-mAP of our method for different action metrics, where [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: The average false negative rate of our method for various characteristics of actions. [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 42 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Alwassel, H.; Giancola, S.; and Ghanem, B. 2021. TSP: Temporally-Sensitive Pretraining of Video Encoders for Localization Tasks. In IEEE/CVF International Conference on Computer Vision Workshops, ICCVW 2021, Montreal, BC, Canada, October 11-17, 2021 , 3166--3176. IEEE

  4. [4]

    C.; Escorcia, V.; and Ghanem, B

    Alwassel, H.; Heilbron, F. C.; Escorcia, V.; and Ghanem, B. 2018. Diagnosing Error in Temporal Action Detectors. In Ferrari, V.; Hebert, M.; Sminchisescu, C.; and Weiss, Y., eds., Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part III , volume 11207 of Lecture Notes in Computer Science, 264--28...

  5. [5]

    Bai, Y.; Wang, Y.; Tong, Y.; Yang, Y.; Liu, Q.; and Liu, J. 2020. Boundary Content Graph Neural Network for Temporal Action Proposal Generation. In Vedaldi, A.; Bischof, H.; Brox, T.; and Frahm, J., eds., Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXVIII , volume 12373 of Lecture Notes in Com...

  6. [6]

    Bodla, N.; Singh, B.; Chellappa, R.; and Davis, L. S. 2017. Soft-NMS - Improving Object Detection with One Line of Code. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 , 5562--5570. IEEE Computer Society

  7. [7]

    Buch, S.; Escorcia, V.; Ghanem, B.; Fei - Fei, L.; and Niebles, J. C. 2017. End-to-End, Single-Stream Temporal Action Detection in Untrimmed Videos. In British Machine Vision Conference 2017, BMVC 2017, London, UK, September 4-7, 2017 . BMVA Press

  8. [8]

    Carreira, J.; and Zisserman, A. 2017. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017 , 4724--4733. IEEE Computer Society

Show all 53 references
  1. [9]

    Chen, G.; Huang, Y.; Xu, J.; Pei, B.; Chen, Z.; Li, Z.; Wang, J.; Li, K.; Lu, T.; and Wang, L. 2024. Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding. arXiv:2403.09626

  2. [10]

    Chen, Z.; Yang, C.; Li, Q.; Zhao, F.; Zha, Z.; and Wu, F. 2021. Disentangle Your Dense Object Detector. In Shen, H. T.; Zhuang, Y.; Smith, J. R.; Yang, Y.; C \' e sar, P.; Metze, F.; and Prabhakaran, B., eds., MM '21: ACM Multimedia Conference, Virtual Event, China, October 20...

  3. [11]

    Cheng, F.; and Bertasius, G. 2022. TallFormer: Temporal Action Localization with a Long-Memory Transformer. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October...

  4. [12]

    M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M

    Damen, D.; Doughty, H.; Farinella, G. M.; Furnari, A.; Kazakos, E.; Ma, J.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M. 2022. Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100 . Int. J. Comput. Vis., 130(1): 33--55

  5. [13]

    Dong, W.; Zhang, Z.; Song, C.; and Tan, T. 2022. Identifying the key frames: An attention-aware sampling method for action recognition. Pattern Recognition, 130: 108797

  6. [14]

    Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019. SlowFast Networks for Video Recognition. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 , 6201--6210. IEEE

  7. [15]

    Ge, Z.; Liu, S.; Wang, F.; Li, Z.; and Sun, J. 2021. YOLOX: Exceeding YOLO Series in 2021. arXiv:2107.08430

  8. [16]

    Gu, A.; and Dao, T. 2024. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752

  9. [17]

    Guibas, J.; Mardani, M.; Li, Z.; Tao, A.; Anandkumar, A.; and Catanzaro, B. 2021. Efficient Token Mixing for Transformers via Adaptive Fourier Neural Operators. In International Conference on Learning Representations

  10. [18]

    C.; Escorcia, V.; Ghanem, B.; and Niebles, J

    Heilbron, F. C.; Escorcia, V.; Ghanem, B.; and Niebles, J. C. 2015. ActivityNet: A large-scale video benchmark for human activity understanding. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015 , 961--970. IEEE Computer Society

  11. [19]

    C.; Niebles, J

    Heilbron, F. C.; Niebles, J. C.; and Ghanem, B. 2016. Fast Temporal Activity Proposals for Efficient Detection of Human Actions in Untrimmed Videos. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 , 1914--1923...

  12. [20]

    Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; Suleyman, M.; and Zisserman, A. 2017. The Kinetics Human Action Video Dataset. arXiv:1705.06950

  13. [21]

    Lee - Thorp, J.; Ainslie, J.; Eckstein, I.; and Onta \ n \' o n, S. 2022. FNet: Mixing Tokens with Fourier Transforms. In Carpuat, M.; de Marneffe, M.; and Ru \' z, I. V. M., eds., Proceedings of the 2022 Conference of the North American Chapter of the Association for Computat...

  14. [22]

    Lin, C.; Xu, C.; Luo, D.; Wang, Y.; Tai, Y.; Wang, C.; Li, J.; Huang, F.; and Fu, Y. 2021. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , 3320...

  15. [23]

    Lin, T.; Zhao, X.; and Shou, Z. 2017. Single Shot Temporal Action Detection. In Liu, Q.; Lienhart, R.; Wang, H.; Chen, S. K.; Boll, S.; Chen, Y. P.; Friedland, G.; Li, J.; and Yan, S., eds., Proceedings of the 2017 ACM on Multimedia Conference, MM 2017, Mountain View, CA, USA,...

  16. [24]

    Lin, T.; Zhao, X.; Su, H.; Wang, C.; and Yang, M. 2018. BSN: Boundary Sensitive Network for Temporal Action Proposal Generation. In Ferrari, V.; Hebert, M.; Sminchisescu, C.; and Weiss, Y., eds., Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, Septembe...

  17. [25]

    Liu, X.; Wang, Q.; Hu, Y.; Tang, X.; Zhang, S.; Bai, S.; and Bai, X. 2022. End-to-End Temporal Action Detection With Transformer. IEEE Trans. Image Process. , 31: 5427--5441

  18. [26]

    Long, F.; Yao, T.; Qiu, Z.; Tian, X.; Luo, J.; and Mei, T. 2019. Gaussian Temporal Awareness Networks for Action Localization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , 344--353. Computer Vision Foundation / IEEE

  19. [27]

    Nag, S.; Zhu, X.; Song, Y.; and Xiang, T. 2022. Semi-supervised Temporal Action Detection with Proposal-Free Masking. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israe...

  20. [28]

    Oppenheim, A. V. 1999. Discrete-time signal processing. Pearson Education India

  21. [29]

    Qing, Z.; Su, H.; Gan, W.; Wang, D.; Wu, W.; Wang, X.; Qiao, Y.; Yan, J.; Gao, C.; and Sang, N. 2021 a . Temporal Context Aggregation Network for Temporal Action Proposal Refinement. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25,...

  22. [30]

    Qing, Z.; Su, H.; Gan, W.; Wang, D.; Wu, W.; Wang, X.; Qiao, Y.; Yan, J.; Gao, C.; and Sang, N. 2021 b . Temporal Context Aggregation Network for Temporal Action Proposal Refinement. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25,...

  23. [31]

    Rao, Y.; Zhao, W.; Zhu, Z.; Zhou, J.; and Lu, J. 2023. GFNet: Global Filter Networks for Visual Recognition. IEEE Trans. Pattern Anal. Mach. Intell. , 45(9): 10960--10973

  24. [32]

    D.; and Savarese, S

    Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I. D.; and Savarese, S. 2019. Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-...

  25. [33]

    Shi, D.; Zhong, Y.; Cao, Q.; Ma, L.; Lit, J.; and Tao, D. 2023. TriDet: Temporal Action Detection with Relative Boundary Modeling. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , 18857--18866. IEEE

  26. [34]

    Shi, D.; Zhong, Y.; Cao, Q.; Zhang, J.; Ma, L.; Li, J.; and Tao, D. 2022. ReAct: Temporal Action Detection with Relational Queries. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, T...

  27. [35]

    Shou, Z.; Wang, D.; and Chang, S. 2016. Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 , 1049--1058. IEEE Computer Society

  28. [36]

    Tan, J.; Zhao, X.; Shi, X.; Kang, B.; and Wang, L. 2022. PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points. In NeurIPS

  29. [37]

    N.; Kim, K.; and Sohn, K

    Tang, T. N.; Kim, K.; and Sohn, K. 2023. TemporalMaxer: Maximize Temporal Context with only Max Pooling for Temporal Action Localization. arXiv:2303.09055

  30. [38]

    N.; Kaiser, L.; and Polosukhin, I

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In Guyon, I.; von Luxburg, U.; Bengio, S.; Wallach, H. M.; Fergus, R.; Vishwanathan, S. V. N.; and Garnett, R., eds., Advances in Neura...

  31. [39]

    Wang, L.; Huang, B.; Zhao, Z.; Tong, Z.; He, Y.; Wang, Y.; Wang, Y.; and Qiao, Y. 2023. VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 , 145...

  32. [40]

    Wang, Y.; Li, K.; Li, X.; Yu, J.; He, Y.; Wang, C.; Chen, G.; Pei, B.; Yan, Z.; Zheng, R.; Xu, J.; Wang, Z.; Shi, Y.; Jiang, T.; Li, S.; Zhang, H.; Huang, Y.; Qiao, Y.; Wang, Y.; and Wang, L. 2024. InternVideo2: Scaling Foundation Models for Multimodal Video Understanding. arX...

  33. [41]

    Wu, Y.; Chen, Y.; Yuan, L.; Liu, Z.; Wang, L.; Li, H.; and Fu, Y. 2020. Rethinking classification and localization for object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 10186--10195

  34. [42]

    S.; Thabet, A

    Xu, M.; Zhao, C.; Rojas, D. S.; Thabet, A. K.; and Ghanem, B. 2020. G-TAD: Sub-Graph Localization for Temporal Action Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020 , 10153--10162. Computer Visio...

  35. [43]

    Z.; Qin, H.; Feng, M.; Zhang, L.; and Mian, A

    Yan, X.; Gilani, S. Z.; Qin, H.; Feng, M.; Zhang, L.; and Mian, A. 2018. Deep Keyframe Detection in Human Action Videos. arXiv:1804.10021

  36. [44]

    Yang, J.; Wei, P.; Ren, Z.; and Zheng, N. 2024. Gated Multi-Scale Transformer for Temporal Action Localization. IEEE Transactions on Multimedia, 26: 5705--5717

  37. [45]

    Yang, J.; Wei, P.; and Zheng, N. 2024. Cross Time-Frequency Transformer for Temporal Action Localization. IEEE Transactions on Circuits and Systems for Video Technology, 34(6): 4625--4638

  38. [46]

    Yang, L.; Peng, H.; Zhang, D.; Fu, J.; and Han, J. 2020. Revisiting Anchor Mechanisms for Temporal Action Localization. IEEE Trans. Image Process. , 29: 8535--8548

  39. [47]

    Yeung, S.; Russakovsky, O.; Jin, N.; Andriluka, M.; Mori, G.; and Fei - Fei, L. 2018 a . Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos. Int. J. Comput. Vis., 126(2-4): 375--389

  40. [48]

    Yeung, S.; Russakovsky, O.; Jin, N.; Andriluka, M.; Mori, G.; and Fei - Fei, L. 2018 b . Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos. Int. J. Comput. Vis., 126(2-4): 375--389

  41. [49]

    Zhang, C.; Wu, J.; and Li, Y. 2022. ActionFormer: Localizing Moments of Actions with Transformers. In Avidan, S.; Brostow, G. J.; Ciss \' e , M.; Farinella, G. M.; and Hassner, T., eds., Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2...

  42. [50]

    Zhang, H.; Wang, Y.; Dayoub, F.; and Sunderhauf, N. 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8514--8523

  43. [51]

    K.; and Ghanem, B

    Zhao, C.; Thabet, A. K.; and Ghanem, B. 2021. Video Self-Stitching Graph Network for Temporal Action Localization. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021 , 13638--13647. IEEE

  44. [52]

    Zhao, H.; Torralba, A.; Torresani, L.; and Yan, Z. 2019. HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 , 8667...

  45. [53]

    Zhuang, J.; Qin, Z.; Yu, H.; and Chen, X. 2023. Task-Specific Context Decoupling for Object Detection. arXiv:2303.01047

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.