Pith. sign in

REVIEW 64 references

Action Anticipation from SoccerNet Football Video Broadcasts

T0 review · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A new action anticipation benchmark for football broadcasts, with a dataset, metrics, and a baseline model that predicts ball-related actions up to ten seconds ahead.

arxiv 2504.12021 v1 pith:EHUVOQVG submitted 2025-04-16 cs.CV

classification cs.CV
keywords anticipationactionactionsfootballfuturevideosdatasetsoccernet
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper defines a new computer vision task: given a short clip of football broadcast video, predict which ball-related actions will happen in the next five or ten seconds, and when. This differs from action spotting, which detects actions after they appear in the video. The authors build the dataset by adapting the existing SoccerNet Ball Action Spotting dataset, keeping ten action classes such as passes, drives, headers, and shots. They exclude goals and free kicks because those classes are too rare in the test split, which would make the evaluation metrics unstable.

The paper proposes a baseline model called FAANTRA. It extracts per-frame features with an efficient backbone network, processes them with a transformer encoder, and uses learnable queries in a transformer decoder to predict action presence, action class, and timing. The model is trained with a multi-task loss that also includes an auxiliary action segmentation task within the observed context window. For evaluation, the authors adapt the standard mean Average Precision metric to anticipation, using several temporal tolerances, so that a prediction counts as correct if it falls within a few seconds of the true action time.

The main results show that the task is hard. The best baseline reaches an average mAP of about 24 on a 0-100 scale for five-second anticipation, while an action spotting model given the future frames as input reaches about 62 on the same metric. The gap illustrates the difficulty of predicting spontaneous football actions from past context alone. Ablations show that spatial resolution and the auxiliary segmentation task matter a lot.

Extended reading notes

Core claim

The paper claims to introduce the first structured benchmark for action anticipation in football broadcast video, including a new dataset (SN-BAA), new metrics (mAP@delta), and a baseline model (FAANTRA). The central quantitative claim is that FAANTRA achieves average mAP of 24.08 at delta tolerances {1,2,3,4,5,infinity} for a five-second anticipation window, trained jointly on SN-BAA and SN-AS with RegNetY-400MF. The paper also claims this task is feasible yet challenging, evidenced by the large gap between FAANTRA and T-DEED used as an upper bound with access to future frames (53.19 average mAP for T-DEED trained on SN-BAA at 200MF).

Load-bearing premise

The dataset construction assumes that all actions in the anticipation window are completely predictable from the context window alone, without any use of future frames. This is the defining assumption of the task. More specifically, the evaluation clips test videos into 30-second segments with a sliding window stride of Ta seconds, and the paper assumes this clipping ensures all actions are evaluated without double counting or missing actions. The validity of the metric depends on this assumption. Additionally, the paper assumes that excluding goals and free-kicks is harmless for the central claim, because these classes are too rare in the test split. If the metric or the clipping were flawed, the reported results would not be comparable across methods. The model also assumes that the auxiliary segmentation task on observed frames transfers useful semantics to the unobserved anticipation window, an assumption supported by their ablation but not by a theoretical argument.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The paper is an empirical benchmark paper rather than a derivation. There are no invented physical entities or new theoretical constructs. The main load-bearing choices are the metric definition, the dataset split and clipping, and the multi-task loss weights. These are domain assumptions and tuned hyperparameters, which are listed above. The absence of invented entities is expected for a benchmark paper.

free parameters (4)
  • Loss weights lambda_D, lambda_C, lambda_T, lambda_S = lambda_D=1, lambda_C=1, lambda_T=10, lambda_S=1
    Chosen by hand without a sensitivity analysis or a stated selection criterion. The temporal loss weight differs from the others by an order of magnitude, which likely affects the localization performance.
  • Number of queries q = q=8 for Ta=5, q=16 for Ta=10
    Set to the maximum number of actions in a window for Ta=5, and doubled for Ta=10. This is a task-specific choice that could be considered a tuned hyperparameter, though it has a plausible rationale.
  • Context window length Tc = Tc=5 seconds
    Selected based on an ablation (Fig. 4) that shows performance plateaus beyond Tc=5. This is a data-driven choice that affects the central results.
  • Attention span k and layer counts lE, lD = k=15, lE=4, lD=2
    Selected based on ablations in Tab. 5. These are standard hyperparameters, but they are chosen by test-set performance, which carries some risk of overfitting to this benchmark.
assumptions (3)
  • domain assumption mAP@delta with the given tolerance windows is a valid and informative measure of anticipation quality.
    The metric is adapted from action spotting and assumes that a prediction within delta/2 seconds of the ground truth is correct. This is a reasonable but non-trivial choice, because the paper does not analyze the sensitivity of conclusions to the tolerance values.
  • domain assumption Clipping test videos into 30-second clips with stride Ta covers all actions without distorting the evaluation.
    This assumption is stated in the dataset section without a formal proof or empirical check. If the clipping misses some action configurations, the reported scores could be biased.
  • domain assumption The auxiliary segmentation task on the context window provides supervision that transfers to the anticipation window.
    The paper validates this empirically in an ablation, but the transfer mechanism is not explained or proven. It is a background assumption about representation learning, not a mathematical axiom.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Action Anticipation from SoccerNet Football Video Broadcasts." pith.science (2026). https://pith.science/paper/EHUVOQVG

@misc{pith2026250412021,
  author       = {Pith},
  title        = {Pith review of: Action Anticipation from SoccerNet Football Video Broadcasts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EHUVOQVG}},
  note         = {Machine review of arXiv:2504.12021}
}
abstract

Artificial intelligence has revolutionized the way we analyze sports videos, whether to understand the actions of games in long untrimmed videos or to anticipate the player's motion in future frames. Despite these efforts, little attention has been given to anticipating game actions before they occur. In this work, we introduce the task of action anticipation for football broadcast videos, which consists in predicting future actions in unobserved future frames, within a five- or ten-second anticipation window. To benchmark this task, we release a new dataset, namely the SoccerNet Ball Action Anticipation dataset, based on SoccerNet Ball Action Spotting. Additionally, we propose a Football Action ANticipation TRAnsformer (FAANTRA), a baseline method that adapts FUTR, a state-of-the-art action anticipation model, to predict ball-related actions. To evaluate action anticipation, we introduce new metrics, including mAP@$\delta$, which evaluates the temporal precision of predicted future actions, as well as mAP@$\infty$, which evaluates their occurrence within the anticipation window. We also conduct extensive ablation studies to examine the impact of various task settings, input configurations, and model architectures. Experimental results highlight both the feasibility and challenges of action anticipation in football videos, providing valuable insights into the design of predictive models for sports analytics. By forecasting actions before they unfold, our work will enable applications in automated broadcasting, tactical analysis, and player decision-making. Our dataset and code are publicly available at https://github.com/MohamadDalal/FAANTRA.

Figures

Figures reproduced from arXiv: 2504.12021 by the authors.

Figure 1
Figure 1. Overview of our new action anticipation task for sports. Action anticipation aims to predict and temporally local￾ize future actions in an anticipation window of Ta seconds using information from a preceding observed context window of Tc sec￾onds. Unlike action spotting, where models can access the entire video sequence to detect actions, action anticipation requires pre￾dicting future events without access to futur… view at source ↗
Figure 2
Figure 2. FAANTRA Architecture Overview. FAANTRA processes context video frames by extracting per-frame representations through a backbone (BB). These features are fed into a transformer encoder to capture temporal dependencies. A set of learnable queries, repre￾senting action predictions, are initialized in the transformer decoder and refined through multiple layers, leveraging information from the encoder. Each refined quer… view at source ↗
Figure 3
Figure 3. Anticipation window length analysis: Performance eval [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Context window length analysis: Performance evalua [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Example of a pass action [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: Example of a drive action [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Example of a header action [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Example of a high pass action [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Example of an out action 14 [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Example of a throw in action [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]
Figure 11
Figure 11. Figure 11: Example of a cross action [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Example of a ball player block action [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: Example of a shot action [PITH_FULL_IMAGE:figures/full_fig_p015_13.png]
Figure 14
Figure 14. Figure 14: Example of a player successful tackle action [PITH_FULL_IMAGE:figures/full_fig_p015_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 60 canonical work pages

  1. [1]

    Uncertainty-aware an- ticipation of activities

    Yazan Abu Farha and Juergen Gall. Uncertainty-aware an- ticipation of activities. In IEEE/CVF Int. Conf. Comput. Vis. Work. (ICCV Work.), pages 1197–1204, Seoul, South Korea,

  2. [2]

    Long-term anticipation of activities with cycle consis- tency

    Yazan Abu Farha, Qiuhong Ke, Bernt Schiele, and Juergen Gall. Long-term anticipation of activities with cycle consis- tency. In DAGM German Conference on Pattern Recogni- tion, pages 159–173, 2021. 2, 5

  3. [3]

    Solution for SoccerNet ball action spot- ting challenge 2023

    Ruslan Baikulov. Solution for SoccerNet ball action spot- ting challenge 2023. https://github.com/lRomul/ ball-action-spotting, 2023. 2

  4. [4]

    Where will players move next? dynamic graphs and hier- archical fusion for movement forecasting in badminton

    Kai-Shiang Chang, Wei-Yao Wang, and Wen-Chih Peng. Where will players move next? dynamic graphs and hier- archical fusion for movement forecasting in badminton. In AAAI, pages 6998–7005, Washington, D.C., USA, 2023. 1, 3

  5. [5]

    Moeslund

    Anthony Cioppa, Adrien Deli `ege, Silvio Giancola, Bernard Ghanem, Marc Van Droogenbroeck, Rikke Gade, and Thomas B. Moeslund. A context-aware loss function for ac- tion spotting in soccer videos. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 13123–13133, Seattle, W A, USA, 2020. 2

  6. [6]

    Anthony Cioppa, Silvio Giancola, Vladimir Somers, Vic- tor Joos, Floriane Magera, Jan Held, Seyed Abolfazl Ghasemzadeh, Xin Zhou, Karolina Seweryn, Mateusz Kowalczyk, Zuzanna Mr´oz, Szymon Łukasik, Michał Hało´n, Hassan Mkhallati, Adrien Deli `ege, Carlos Hinojosa, Karen Sanchez, Amir M. Mansourian, Pierre Miralles, Olivier Barnich, Christophe De Vleescho...

  7. [7]

    Anthony Cioppa, Silvio Giancola, Vladimir Somers, Flori- ane Magera, Xin Zhou, Hassan Mkhallati, Adrien Deli `ege, Jan Held, Carlos Hinojosa, Amir M. Mansourian, Pierre Miralles, Olivier Barnich, Christophe De Vleeschouwer, Alexandre Alahi, Bernard Ghanem, Marc Van Droogen- broeck, Abdullah Kamal, Adrien Maglo, Albert Clap ´es, Amr Abdelaziz, Artur Xarles...

  8. [8]

    Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100. Int. J. Comput. Vis., 130(1):33–55, 2021. 2, 3

Show all 64 references
  1. [9]

    Online ac- tion detection

    Roeland De Geest, Efstratios Gavves, Amir Ghodrati, Zhenyang Li, Cees Snoek, and Tinne Tuytelaars. Online ac- tion detection. In Eur. Conf. Comput. Vis. (ECCV) , pages 269–284. 2016. 2, 3

  2. [10]

    Seikavandi, Jacob V

    Adrien Deli `ege, Anthony Cioppa, Silvio Giancola, Meisam J. Seikavandi, Jacob V . Dueholm, Kamal Nas- rollahi, Bernard Ghanem, Thomas B. Moeslund, and Marc Van Droogenbroeck. SoccerNet-v2: A dataset and bench- marks for holistic understanding of broadcast soccer videos. In IE...

  3. [11]

    COMEDIAN: Self-supervised learning and knowledge distillation for action spotting using transformers

    Julien Denize, Mykola Liashuha, Jaonary Rabarisoa, Astrid Orcesi, and Romain H ´erault. COMEDIAN: Self-supervised learning and knowledge distillation for action spotting using transformers. In IEEE/CVF Winter Conf. Appl. Comput. Vis. Work. (WACVW), pages 518–528, Waikoloa, HI,...

  4. [12]

    When will you do what? anticipating temporal occurrences of activities

    Yazan Abu Farha, Alexander Richard, and Juergen Gall. When will you do what? anticipating temporal occurrences of activities. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , pages 5343–5352, Salt Lake City, UT, USA, 2018. 2

  5. [13]

    What will happen next? forecasting player moves in sports videos

    Panna Felsen, Pulkit Agrawal, and Jitendra Malik. What will happen next? forecasting player moves in sports videos. 9 In IEEE Int. Conf. Comput. Vis. (ICCV) , pages 3362–3371, Venice, Italy, 2017. 3

  6. [14]

    Rolling- unrolling LSTMs for action anticipation from first-person video

    Antonino Furnari and Giovanni Maria Farinella. Rolling- unrolling LSTMs for action anticipation from first-person video. IEEE Trans. Pattern Anal. Mach. Intell. , 43(11): 4021–4036, 2021. 2

  7. [15]

    Forecasting future action sequences with neural memory networks

    Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes. Forecasting future action sequences with neural memory networks. In Br. Mach. Vis. Conf. (BMVC), pages 1–12, Cardiff, United Kingdom, 2019. 2

  8. [16]

    Temporally-aware feature pooling for action spotting in soccer broadcasts

    Silvio Giancola and Bernard Ghanem. Temporally-aware feature pooling for action spotting in soccer broadcasts. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW), pages 4485–4494, Nashville, TN, USA, 2021. 2

  9. [17]

    SoccerNet: A scalable dataset for action spotting in soccer videos

    Silvio Giancola, Mohieddine Amine, Tarek Dghaily, and Bernard Ghanem. SoccerNet: A scalable dataset for action spotting in soccer videos. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW), pages 1792–179210, Salt Lake City, UT, USA, 2018. 2, 3

  10. [18]

    Anthony Chan, He Zhu, Hongwei Kan, Jiaming Chu, Jianming Hu, Jianyang Gu, Jin Chen, Jo˜ao V

    Silvio Giancola, Anthony Cioppa, Adrien Deli `ege, Flori- ane Magera, Vladimir Somers, Le Kang, Xin Zhou, Olivier Barnich, Christophe De Vleeschouwer, Alexandre Alahi, Bernard Ghanem, Marc Van Droogenbroeck, Abdulrahman Darwish, Adrien Maglo, Albert Clap ´es, Andreas Luyts, An...

  11. [19]

    Deep learning for action spot- ting in association football videos

    Silvio Giancola, Anthony Cioppa, Bernard Ghanem, and Marc Van Droogenbroeck. Deep learning for action spot- ting in association football videos. arXiv, abs/2410.01304,

  12. [20]

    Latency matters: Real-time action fore- casting transformer

    Harshayu Girase, Nakul Agarwal, Chiho Choi, and Kart- tikeya Mangalam. Latency matters: Real-time action fore- casting transformer. In IEEE/CVF Conf. Comput. Vis. Pat- tern Recognit. (CVPR) , pages 18759–18769, Vancouver, Can., 2023. 2

  13. [21]

    Anticipative video transformer

    Rohit Girdhar and Kristen Grauman. Anticipative video transformer. In IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , pages 13485–13495, Montreal, QC, Canada, 2021. 2

  14. [22]

    What to do and where to go next? action prediction in soccer using multimodal co- attention transformer

    Ryota Goka, Yuya Moroto, Keisuke Maeda, Takahiro Ogawa, and Miki Haseyama. What to do and where to go next? action prediction in soccer using multimodal co- attention transformer. InInt. ACM Work. Multimedia Content Anal. Sports (MMSports) , pages 75–79, Melbourne, Victo- ria,...

  15. [23]

    Future transformer for long-term action an- ticipation

    Dayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha, and Minsu Cho. Future transformer for long-term action an- ticipation. In IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR) , pages 3042–3051, New Orleans, LA, USA,

  16. [24]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Mar- tin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Meng...

  17. [25]

    Ross, Carl V on- drick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijaya- narasimhan, George Toderici, Susanna Ricco, Rahul Suk- thankar, Cordelia Schmid, and Jitendra Malik

    Chunhui Gu, Chen Sun, David A. Ross, Carl V on- drick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijaya- narasimhan, George Toderici, Susanna Ricco, Rahul Suk- thankar, Cordelia Schmid, and Jitendra Malik. A V A: A video dataset of spatio-temporally localized atomic visual act...

  18. [26]

    V ARS: Video assistant referee system for automated soc- cer decision making from multiple views

    Jan Held, Anthony Cioppa, Silvio Giancola, Abdullah Hamdi, Bernard Ghanem, and Marc Van Droogenbroeck. V ARS: Video assistant referee system for automated soc- cer decision making from multiple views. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW), pages 5086–5...

  19. [27]

    X-V ARS: Introducing explainability in football refereeing with multi- modal large language models

    Jan Held, Hani Itani, Anthony Cioppa, Silvio Giancola, Bernard Ghanem, and Marc Van Droogenbroeck. X-V ARS: Introducing explainability in football refereeing with multi- modal large language models. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW) , pages 3267–32...

  20. [28]

    Spotting temporally pre- cise, fine-grained events in video

    James Hong, Haotian Zhang, Micha ¨el Gharbi, Matthew 10 Fisher, and Kayvon Fatahalian. Spotting temporally pre- cise, fine-grained events in video. In Eur. Conf. Comput. Vis. (ECCV), pages 33–51, Tel Aviv, Isra¨el, 2022. 2, 4, 5

  21. [29]

    in the wild

    Haroon Idrees, Amir R. Zamir, Yu-Gang Jiang, Alex Gor- ban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah. The THUMOS challenge on action recognition for videos “ in the wild”. Comput. Vis. Image Underst., 155:1–23, 2017. 3

  22. [30]

    The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities

    Hilde Kuehne, Ali Arslan, and Thomas Serre. The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities. In IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 780–787, Columbus, OH, USA, 2014. 2, 3

  23. [31]

    Yin Li, Miao Liu, and James M. Rehg. In the eye of beholder: Joint learning of gaze and actions in first person video. In Eur. Conf. Comput. Vis. (ECCV), pages 639–655, 2018. 2, 3

  24. [32]

    MViTv2: Improved multiscale vision transformers for classification and detection

    Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Man- galam, Bo Xiong, Jitendra Malik, and Christoph Feichten- hofer. MViTv2: Improved multiscale vision transformers for classification and detection. In IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR), pages 4794–4804, Ne...

  25. [33]

    Peeking into the future: Pre- dicting future person activities and locations in videos

    Junwei Liang, Lu Jiang, Juan Carlos Niebles, Alexander Hauptmann, and Li Fei-Fei. Peeking into the future: Pre- dicting future person activities and locations in videos. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW), pages 2960–2963, Long Beach, CA, USA, 2019. 2

  26. [34]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Int. Conf. Learn. Represent., New Orleans, LA, USA, 2019. 6

  27. [35]

    Intention-conditioned long-term human egocentric action anticipation

    Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee. Intention-conditioned long-term human egocentric action anticipation. In IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), pages 6037–6046, Waikoloa, HI, USA, 2023. 2

  28. [36]

    Ego-topo: Environment affordances from egocentric video

    Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman. Ego-topo: Environment affordances from egocentric video. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 160–169, Seattle, W A, USA, 2020. 2

  29. [37]

    Re- thinking learning approaches for long-term action anticipa- tion

    Megha Nawhal, Akash Abdu Jyothi, and Greg Mori. Re- thinking learning approaches for long-term action anticipa- tion. In Eur. Conf. Comput. Vis. (ECCV) , pages 558–576,

  30. [38]

    Designing network design spaces

    Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollar. Designing network design spaces. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 10425–10433, Seattle, W A, USA, 2020. 4

  31. [39]

    Temporal aggregate representations for long-range video understand- ing

    Fadime Sener, Dipika Singhania, and Angela Yao. Temporal aggregate representations for long-range video understand- ing. In Eur. Conf. Comput. Vis. (ECCV) , pages 154–171,

  32. [40]

    As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

    Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities. In IEEE/CVF Conf. Com- put. Vis. Pattern Recognit. (CVPR) , pages 21064–2...

  33. [41]

    Survey of action recognition, spotting and spatio-temporal localization in soccer – current trends and research perspec- tives

    Karolina Seweryn, Anna Wr ´oblewska, and Szymon Łukasik. Survey of action recognition, spotting and spatio-temporal localization in soccer – current trends and research perspec- tives. arXiv, abs/2309.12067, 2023. 2

  34. [42]

    Sigurdsson, G ¨ul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta

    Gunnar A. Sigurdsson, G ¨ul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activity under- standing. In Eur. Conf. Comput. Vis. (ECCV) , pages 510– 526, 2016. 3

  35. [43]

    Jo ˜ao V . B. Soares, Avijit Shah, and Topojoy Biswas. Tem- porally precise action spotting in soccer videos using dense detection anchors. In IEEE Int. Conf. Image Process. (ICIP), pages 2796–2800, Bordeaux, France, 2022. 2, 8

  36. [44]

    Sebastian Stein and Stephen J. McKenna. Combining em- bedded accelerometers with computer vision for recognizing food preparation activities. In ACM Int. Jt. Conf. Pervasive Ubiquitous Comput. , pages 729–738, Zurich, Switzerland,

  37. [45]

    Gate-shift-fuse for video action recognition

    Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz. Gate-shift-fuse for video action recognition. IEEE Trans. Pattern Anal. Mach. Intell., 45(9):10913–10928, 2023. 4

  38. [46]

    Relational action forecasting

    Chen Sun, Abhinav Shrivastava, Carl V ondrick, Rahul Suk- thankar, Kevin Murphy, and Cordelia Schmid. Relational action forecasting. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 273–283, Long Beach, CA, USA,

  39. [47]

    RMS-net: Regression and masking for soccer event spotting

    Matteo Tomei, Lorenzo Baraldi, Simone Calderara, Simone Bronzin, and Rita Cucchiara. RMS-net: Regression and masking for soccer event spotting. In IEEE Int. Conf. Pat- tern Recognit. (ICPR), pages 7699–7706, Milan, Italy, 2021. 2

  40. [48]

    AI-based video clip- ping of soccer events

    Joakim Valand, Haris Kadragic, Steven Hicks, Vajira Thambawita, Cise Midoglu, Tomas Kupka, Dag Johansen, Michael Riegler, and P ˚al Halvorsen. AI-based video clip- ping of soccer events. Mach. Learn. & Knowl. Extr. , 3(4): 1–19, 2021. 2

  41. [49]

    Automated clipping of soccer events using machine learning

    Joakim Valand, Haris Kadragic, Steven Hicks, Vajira Thambawita, Cise Midoglu, Tomas Kupka, Dag Johansen, Michael Riegler, and P˚al Halvorsen. Automated clipping of soccer events using machine learning. In Int. Symp. Multi- media (ISM), pages 210–214, Naple, Italy, 2021. 2

  42. [50]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Adv. Neural Inf. Process. Syst. (NeurIPS) , pages 6000–6010, Long Beach, CA, USA, 2017. 5

  43. [51]

    Predicting shot locations in tennis using spatiotempo- ral data

    Xinyu Wei, Patrick Lucey, Stuart Morgan, and Sridha Srid- haran. Predicting shot locations in tennis using spatiotempo- ral data. In Digit. Image Comput.: Tech. Appl. , pages 1–8, Hobart, TAS, Australia, 2013. 1, 3

  44. [52]

    Forecasting events using an aug- mented hidden conditional random field

    Xinyu Wei, Patrick Lucey, Stephen Vidas, Stuart Morgan, and Sridha Sridharan. Forecasting events using an aug- mented hidden conditional random field. In Asian Conf. Comput. Vis. (ACCV), pages 569–582, 2015. 3

  45. [53]

    A survey on video action recognition in sports: Datasets, methods and 11 applications

    Fei Wu, Qingzhong Wang, Jiang Bian, Ning Ding, Feixiang Lu, Jun Cheng, Dejing Dou, and Haoyi Xiong. A survey on video action recognition in sports: Datasets, methods and 11 applications. IEEE Trans. Multimedia, 25:7943–7966, 2023. 1

  46. [54]

    Moeslund, and Al- bert Clap´es

    Artur Xarles, Sergio Escalera, Thomas B. Moeslund, and Al- bert Clap´es. ASTRA: An Action Spotting TRAnsformer for soccer videos. In Int. ACM Work. Multimedia Content Anal. Sports (MMSports) , page 93–102, Ottawa, Ontario, Can.,

  47. [55]

    Moeslund, and Al- bert Clap ´es

    Artur Xarles, Sergio Escalera, Thomas B. Moeslund, and Al- bert Clap ´es. T-DEED: Temporal-discriminability enhancer encoder-decoder for precise event spotting in sports videos. In IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW), pages 3410–3419, Seattle, W A, USA,...

  48. [56]

    Long short-term trans- former for online action detection

    Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto. Long short-term trans- former for online action detection. In Adv. Neural Inf. Pro- cess. Syst. (NeurIPS), pages 1086–1099. 2021. 2

  49. [57]

    Object-centric video representation for long-term action anticipation

    Ce Zhang, Changcheng Fu, Shijie Wang, Nakul Agarwal, Kwonjoon Lee, Chiho Choi, and Chen Sun. Object-centric video representation for long-term action anticipation. In IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), pages 6737–6747, Waikoloa, HI, USA, 2024. 2

  50. [58]

    AntGPT: Can large language models help long-term action anticipation from videos? In Int

    Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu, Minh Quan Do, Nakul Agarwal, Kwonjoon Lee, and Chen Sun. AntGPT: Can large language models help long-term action anticipation from videos? In Int. Conf. Learn. Repre- sent., Vienna, Austria, 2024. 2

  51. [59]

    Real-time online video detection with temporal smoothing transformers

    Yue Zhao and Philipp Kr ¨ahenb¨uhl. Real-time online video detection with temporal smoothing transformers. In Eur. Conf. Comput. Vis. (ECCV), pages 485–502. 2022. 2

  52. [60]

    Unsupervised learning for forecasting action representations

    Yi Zhong and Wei-Shi Zheng. Unsupervised learning for forecasting action representations. In IEEE Int. Conf. Image Process. (ICIP), pages 1073–1077, Athens, Greece, 2018. 2

  53. [61]

    A survey on deep learning techniques for action anticipation

    Zeyun Zhong, Manuel Martin, Michael V oit, Juergen Gall, and J ¨urgen Beyerer. A survey on deep learning techniques for action anticipation. arXiv, abs/2309.17257, 2023. 3

  54. [62]

    Anticipative feature fu- sion transformer for multi-modal action anticipation

    Zeyun Zhong, David Schneider, Michael V oit, Rainer Stiefelhagen, and Jurgen Beyerer. Anticipative feature fu- sion transformer for multi-modal action anticipation. In IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), pages 6057–6066, Waikoloa, HI, USA, 2023. 2

  55. [63]

    DiffAnt: Diffusion models for action anticipation

    Zeyun Zhong, Chengzhi Wu, Manuel Martin, Michael V oit, Juergen Gall, and J¨urgen Beyerer. DiffAnt: Diffusion models for action anticipation. arXiv, abs/2311.15991, 2023. 2, 5

  56. [64]

    Feature combination meets attention: Baidu soccer embed- dings and transformer based temporal detection

    Xin Zhou, Le Kang, Zhiyu Cheng, Bo He, and Jingyu Xin. Feature combination meets attention: Baidu soccer embed- dings and transformer based temporal detection. arXiv, abs/2106.14447, 2021. 2 12 Action Anticipation from SoccerNet Football Video Broadcasts Supplementary Material...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.