REVIEW 3 major objections 6 minor 98 references
Event Detection in Videos: A Framework for the Development of New Methods
T0 review · 3 major / 6 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read A three-pillar framework—datasets, probabilistic evaluation, and deployment scenarios—aims to make video event-detection methods comparable without the biases of small, narrow benchmarks.
desk verdict Solid three-pillar methodology paper: new multi-environment datasets and scenario checklist are the real additions; the ranking/Tile math is mostly restated prior work and the crisp-classification premise is a scope condition, not a flaw. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The probabilistic performance pipeline (performance as a probability measure on {tn, fp, fn, tp}, source-mixture summarization, and the two-parameter family of canonical ranking scores plotted as Tiles) that turns any well-specified random evaluation experiment into stable, preference-aware rankings.
What would settle it
Apply the Tile ranking pipeline to a task whose event boundaries remain contested (for example, “how many people entered” under two different elementary-event definitions) and check whether the resulting orderings of methods reverse or become unstable across those definitions.
Extended reading notes
Core claim
A rigorous framework for new video event-detection methods rests on three pillars: (1) a hierarchically structured, multi-environment, tag-annotated large-scale collection of videos that includes two new public/private datasets; (2) a probabilistic evaluation pipeline that represents performance as a probability measure on the four crisp outcomes, summarizes by source mixtures, analyzes with importance-parameterized ranking scores displayed as Tiles, and ranks methods stably; and (3) explicit application scenarios that fix data access, prior knowledge, and evaluation conditions so methods can be compared fairly.
Load-bearing premise
Every interesting video event can be turned, without leftover ambiguity about spatial or temporal boundaries, into a single two-class crisp classification whose four outcomes are fully defined by one random experiment that stays linear in both data sources and methods.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a three-pillar framework for developing and comparing event-detection methods in videos: (i) a hierarchical, tag-based organization of event-monitoring video data spanning urban, natural, maritime, and underwater environments, including two new datasets (FSD from public IP cameras and synthetic SUC from CARLA); (ii) a probabilistic evaluation pipeline that casts detection as two-class crisp classification, averages performances via mixture-linear confusion matrices, and ranks methods with application-parameterized canonical ranking scores visualized as Tiles; and (iii) an explicit checklist of application-scenario characteristics (data, methods, evaluation, operational constraints) so that methods can be compared under stated conditions. The evaluation mathematics is grounded in the authors’ prior axiomatic work on ranking scores and Tiles; the dataset pillar unifies 36 public sources plus FSD/SUC under environment–modality–tag metadata.
Significance. If adopted, the framework would reduce the long-standing practice of ranking background-subtraction and related detectors with ad-hoc score averages and poorly specified tasks, and would give challenge organizers a reproducible way to report full confusion matrices and multi-criterion rankings. The Tile geometry and the linearity-based summarization/ranking results are carefully referenced to peer-reviewed foundations and are a genuine strength. FSD (≈153k annotated pairs) and SUC (≈187.5k frames with instance masks and weather metadata) are concrete, usable resources. The scenario checklist is a practical contribution for reproducibility and for high-risk AI documentation. The main value is organizational and methodological rather than a new detector or a large empirical leaderboard.
major comments (3)
- [§III–V (esp. §IV pipeline and new datasets)] The central claim is that the three pillars form a usable development framework, yet the manuscript never runs the full pipeline end-to-end on FSD or SUC (or on a multi-environment subset of EMVD). Section IV illustrates Tiles on CDnet and cites IWDD, but there is no Entity/Value Tile, mixture summarization, or scenario-conditioned ranking produced from the new data. Without at least one complete worked example, the claim that the framework enables rigorous development remains aspirational rather than demonstrated.
- [§IV-A1–A4] §IV largely restates the probabilistic performance model, ranking scores R_I, canonical scores parameterized by (a,b), and Tile flavors from the authors’ prior work [13,14,39]. The event-detection-specific content is mainly the three random-evaluation experiments and the hardware/real-time metrics. The manuscript should state explicitly what is new for video event detection versus what is imported, and should show where the linearity-in-method / linearity-in-source assumptions fail for common tasks (e.g., multi-object matching with non-unique associations, or clip-level events with soft temporal bounds).
- [§I, §IV-A1, §V-C] The paper correctly flags Bertrand’s paradox and ill-posed spatial/temporal bounds (§I), then requires the designer to fix a single linear random evaluation experiment. That is a scope condition, not a contradiction, but it is load-bearing: many clip-level and action-spotting tasks do not admit an unambiguous exhaustive negative class or a unique matching rule. §V should require every scenario to publish the exact random experiment (or matching protocol) and to state whether linearity holds; otherwise Tile rankings lose their formal guarantees on those tasks.
minor comments (6)
- [Abstract, §I] Abstract and §I claim that lack of large-scale data and rigorous evaluation “have biased” comparisons; this is plausible but unsupported by a concrete before/after or citation analysis. Soften or cite evidence.
- [§III-B, Table I, Fig. 2] Fig. 1 and Table I are useful; ensure every dataset listed in §III-B appears consistently in the table and timeline (Fig. 2), including Audio-Visual Vehicle and FSD/SUC.
- [§III-C] FSD and SUC are partly private “to enable fair evaluation.” State clearly what is public now, what will be released, and how third parties can reproduce the framework without the private split.
- [§IV-D, §IV-E, §V-D] Hardware metrics (§IV-D–E) are sensible but disconnected from the probabilistic pipeline; a short note on how MEM_Δ / PFR_Δ / delay interact with scenario constraints would help.
- [Title page, throughout] Minor typos and formatting: “Deli `ege”, “Pi ´erard”, “Staff, IEEE,” trailing commas in author list; “W ACV” spacing; arXiv id in header is future-dated (2607).
- [§IV-B (G2)] Guideline (G2) urges full confusion matrices; the paper itself does not release matrices for FSD/SUC. Align practice with the guideline or mark them as forthcoming.
Circularity Check
Mild self-citation load-bearing for the evaluation pillar only; datasets, scenarios, and overall three-pillar proposal remain independent of any constructional reduction.
-
self citation load bearing
[Section II (Related Work) and Section IV-A (esp. IV-A1–A4)]
"The second limitation has been addressed in the works of Piérard et al. [13, 14] which have been applied successfully in the International Contest on Illegal Waste Dumping Detection (IWSS 2026) in conjunction with WACV 2026 [7]. These works have led us to propose the performance evaluation tools described in Section IV. ... we follow the framework of Piérard et al. [13] ... Piérard et al. [13] introduced the first axiomatic framework for performance-based rankings ... the canonical ranking scores are defined in [14, 39]"
One of the three main pillars (the ‘rigorous’ probabilistic evaluation pipeline, ranking scores, and Tile visualizations) is justified almost exclusively by citations to prior work whose author lists heavily overlap the present paper. The present text treats those results as established external foundations rather than re-deriving them, making the evaluation claims load-bearing on self-citation. The prior works themselves contain independent mathematical arguments, so the reduction is not total; the other two pillars (datasets, scenarios) do not depend on it.
full rationale
This is a framework/proposal paper, not a derivation of a numerical prediction or uniqueness theorem from first principles. The three claimed pillars (hierarchical tagged datasets including new FSD/SUC collections, a probabilistic evaluation pipeline with Tiles/ranking scores, and explicit application scenarios) do not reduce by construction to their own inputs. The evaluation pillar (Section IV) does rest load-bearingly on the authors’ own prior axiomatic ranking and Tile papers (Piérard et al. [13,14] and related), which share multiple co-authors and are invoked as the foundation for summarization, canonical ranking scores, Value/Entity Tiles, and stability arguments. Those citations supply independent mathematical content rather than fitted parameters or tautological redefinitions, so the circularity is limited to self-citation dependence for one pillar. No equation equates a claimed prediction to a fitted input; no uniqueness theorem is imported solely to forbid alternatives; Bertrand’s paradox is acknowledged as a scope condition rather than resolved circularly. Datasets and scenario language stand free of that chain. Score 2 reflects proportionate mild self-citation without forcing the central framework claim.
Assumptions & free parameters
free parameters (2)
- mixture weights λ_s over data sources
- importance pair (a,b) on the Tile
assumptions (3)
- domain assumption Any event-detection task of interest admits an exhaustive two-class crisp classification formulation whose four outcomes are completely defined by a single random evaluation experiment.
- domain assumption The chosen random evaluation experiment is linear with respect to both the data source and the evaluated method, so that mixtures of videos or of methods yield convex combinations of performances.
- standard math Standard probability measure theory on the finite sample space Ω = {tn,fp,fn,tp} is an adequate model of performance.
invented entities (3)
-
Foreground Segmentation Dataset (FSD)
-
Synthetic Urban Crossroad (SUC) dataset
-
Application scenario (formal checklist)
Cite this review
Pith. "Pith review of Event Detection in Videos: A Framework for the Development of New Methods." pith.science (2026). https://pith.science/paper/GNUQE5YX
@misc{pith2026260704372,
author = {Pith},
title = {Pith review of: Event Detection in Videos: A Framework for the Development of New Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/GNUQE5YX}},
note = {Machine review of arXiv:2607.04372}
}
read the original abstract
Event detection tasks in videos, the most important aspect of video surveillance, aim to detect events either at the pixel-level, frame-level, or clip-level. Plenty of methods intended for event detection in different environments, for various applications, and within different acquisition techniques were introduced. Naturally, the attempts were made as well to classify these algorithms in terms of detection of performance or in terms of real-time abilities. Nevertheless, the lack of a large-scale dataset as well as rigorous performance evaluation methods have biased such comparisons as well as the development of the methods. Given the diversity of existing approaches, we believe it is essential for researchers to position their work within such a rich landscape. Thus, we propose a rigorous framework for developing new methods in event detection for videos. Specifically, this framework is based on three main pillars: datasets, performance evaluation, and scenarios for deploying methods.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Changedetection.net: A new change detection benchmark dataset
Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, and Prakash Ishwar. “Changedetection.net: A new change detection benchmark dataset”. In:IEEE Int. Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Providence, RI, USA: IEEE, June 2012, pp. 1–8.URL: https://doi.org/10.1109/CVPRW.2012.6238919
-
[2]
CD- net 2014: An Expanded Change Detection Benchmark Dataset
Yi Wang, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, Yannick Benezeth, and Prakash Ishwar. “CD- net 2014: An Expanded Change Detection Benchmark Dataset”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Columbus, OH, USA: Inst. Electr. Electron. Eng. (IEEE), June 2014, pp. 393–400. URL: https://doi.org/10.1109/CVPRW.2014.126
-
[3]
SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos
Silvio Giancola, Mohieddine Amine, Tarek Dghaily, and Bernard Ghanem. “SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Salt Lake City, UT, USA: IEEE, June 2018, pp. 1792– 179210.URL: https : / / doi . org / 10 . 1109 / cvprw. 2018 . 00223
2018
-
[4]
Smoke detection in video using convolutional neural networks and efficient spatio-temporal features
Mahdi Hashemzadeh, Nacer Farajzadeh, and Milad Heydari. “Smoke detection in video using convolutional neural networks and efficient spatio-temporal features”. In:Appl. Soft Comput.128 (Oct. 2022), p. 109496.URL: https://doi.org/10.1016/j.asoc.2022.109496
-
[5]
VideoGasNet: Deep learning for natural gas methane leak classification using an infrared camera
Jingfan Wang, Jingwei Ji, Arvind P. Ravikumar, Silvio Savarese, and Adam R. Brandt. “VideoGasNet: Deep learning for natural gas methane leak classification using an infrared camera”. In:Energy238 (Jan. 2022), p. 121516.URL: https://doi.org/10.1016/j.energy.2021. 121516
-
[6]
Perimeter Intrusion Detection by Video Surveillance: A Survey
Devashish Lohani, Carlos Crispim-Junior, Quentin Barth´elemy, Sarah Bertrand, Lionel Robinault, and Laure Tougne Rodet. “Perimeter Intrusion Detection by Video Surveillance: A Survey”. In:Sensors22.9 (May 2022), pp. 1–28.URL: https://doi.org/10.3390/ s22093601
2022
-
[7]
Illegal waste dump- ing detection
Thierry Bouwmans, Antonio Greco, S ´ebastien Pi ´erard, Andrea Vincenzo Ricciardi, Carlo Sansone, Marc Van Droogenbroeck, and Bruno Vento. “Illegal waste dump- ing detection”. In:IEEE/CVF Winter Conf. Appl. Com- put. Vis. Work. (WACVW). Tucson, AZ, USA, Mar. 2026, pp. 539–548
2026
-
[8]
Gauthier- Villars et fils, 1889
Joseph Bertrand.Calcul des probabilit ´es. Gauthier- Villars et fils, 1889
Show all 98 references
-
[9]
A survey of approaches and trends in person re-identification
Apurva Bedagkar-Gala and Shishir K. Shah. “A survey of approaches and trends in person re-identification”. In:Image Vis. Comput.32.4 (Apr. 2014), pp. 270–286. URL: https://doi.org/10.1016/j.imavis.2014.02.001
2014 doi
-
[10]
Wallflower: Principles and Practice of Background Maintenance
Kentaro Toyama, John Krumm, Barry Brumitt, and Brian Meyers. “Wallflower: Principles and Practice of Background Maintenance”. In:IEEE Int. Conf. Comput. Vis. (ICCV). Kerkyra, Greece, Sept. 1999, pp. 255–261. URL: https://doi.org/10.1109/ICCV .1999.791228
1999 doi
-
[11]
Evaluation of background subtraction tech- niques for video surveillance
Sebastian Brutzer, Benjamin Hoferlin, and Gunther Hei- demann. “Evaluation of background subtraction tech- niques for video surveillance”. In:IEEE Int. Conf. Comput. Vis. Pattern Recognit. (CVPR). Providence, RI, USA: IEEE, June 2011, pp. 1937–1944.URL: https : //doi.org/10.11...
2011 doi
-
[12]
A Benchmark Dataset for Outdoor Foreground/Background Extraction
Antoine Vacavant, Thierry Chateau, Alexis Wilhelm, and Laurent Lequi `evre. “A Benchmark Dataset for Outdoor Foreground/Background Extraction”. In:Asian Conf. Comput. Vis. (ACCV). V ol. 7728. Lect. Notes 17 Comput. Sci. Springer Berl. Heidelb., Nov. 2012, pp. 291–300.URL: http...
2012
-
[13]
Foun- dations of the Theory of Performance-Based Rank- ing
S ´ebastien Pi ´erard, Ana ¨ıs Halin, Anthony Cioppa, Adrien Deli `ege, and Marc Van Droogenbroeck. “Foun- dations of the Theory of Performance-Based Rank- ing”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recog- nit. (CVPR). Nashville, TN, USA: IEEE, June 2025, pp. 14293–14302.URL...
2025
-
[14]
The Tile: A 2D Map of Ranking Scores for Two-Class Classification
S ´ebastien Pi ´erard, Ana ¨ıs Halin, Anthony Cioppa, Adrien Deli `ege, and Marc Van Droogenbroeck. “The Tile: A 2D Map of Ranking Scores for Two-Class Classification”. In:arXivabs/2412.04309 (2024). arXiv: 2412.04309.URL: https://doi.org/10.48550/arXiv.2412. 04309
-
[15]
CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. “CARLA: An Open Urban Driving Simulator”. In:Annu. Conf. Robot. Learn.V ol. 78. Proc. Mach. Learn. Res. Mountain View, CA, USA: ML Research Press, Nov. 2017, pp. 1–16. URL: https://proceedings...
2017
-
[16]
Real-Time Semantic Background Subtrac- tion
Anthony Cioppa, Marc Van Droogenbroeck, and Marc Braham. “Real-Time Semantic Background Subtrac- tion”. In:IEEE Int. Conf. Image Process. (ICIP). Abu Dhabi, United Arab. Emir.: IEEE, Oct. 2020, pp. 3214– 3218.URL: https://doi.org/10.1109/ICIP40778.2020. 9190838
2020 doi
-
[17]
BSUV-Net 2.0: Spatio-Temporal Data Augmen- tations for Video-AgnosticSupervised Background Sub- traction
M. Ozan Tezcan, Prakash Ishwar, and Janusz Kon- rad. “BSUV-Net 2.0: Spatio-Temporal Data Augmen- tations for Video-AgnosticSupervised Background Sub- traction”. In:IEEE Access9 (2021), pp. 53849–53860. URL: https://doi.org/10.1109/ACCESS.2021.3071163
2021 doi
-
[18]
Compar- ative study of background subtraction algorithms
Yannick Benezeth, Pierre-Marc Jodoin, Bruno. Emile, H´el`ene Laurent, and Christophe Rosenberger. “Compar- ative study of background subtraction algorithms”. In: J. Electron. Imaging19.3 (July 2010), pp. 1–12.URL: https://doi.org/10.1117/1.3456695
2010 doi
-
[19]
A Probabilistic In- terpretation of Precision, Recall and F-Score, with Im- plication for Evaluation
Cyril Goutte and Eric Gaussier. “A Probabilistic In- terpretation of Precision, Recall and F-Score, with Im- plication for Evaluation”. In:Advances in Information Retrieval (Proceedings of ECIR). V ol. 3408. Lect. Notes Comput. Sci. Springer, Mar. 2005, pp. 345–359.URL: https:...
2005 doi
-
[20]
Nouvelles recherches sur la distribution florale
Paul Jaccard. “Nouvelles recherches sur la distribution florale”. In:Bull. De La Soci ´et´e Vaudoise Des Sci. Nat. 44.163 (1908), pp. 223–270
1908
-
[21]
Cornelis Joost van Rijsbergen.Information Retrieval. second. London, Engl.: Butterworths, 1979
1979
-
[22]
A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives
Peter Christen, David J. Hand, and Nishadi Kirielle. “A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives”. In:ACM Comput. Surv. 56.3 (Oct. 2023), pp. 1–24.URL: https://doi.org/10. 1145/3606367
2023
-
[23]
PToPI: A Comprehensive Review, Analysis, and Knowledge Representation of Binary Classification Performance Measures/Metrics
G ¨urol Canbek, Tugba Taskaya Temizel, and Seref Sa- giroglu. “PToPI: A Comprehensive Review, Analysis, and Knowledge Representation of Binary Classification Performance Measures/Metrics”. In:SN Computer Sci- ence4.1 (Oct. 2022).URL: https://doi.org/10.1007/ s42979-022-01409-1
2022
-
[24]
Classification assessment methods
Alaa Tharwat. “Classification assessment methods”. In: Appl. Comput. Informatics17.1 (2021), pp. 168–192. URL: https://doi.org/10.1016/j.aci.2018.08.003
2021 doi
-
[25]
Evaluation: from precision, recall and F-measure to ROC, informedness, marked- ness and correlation
David M. W. Powers. “Evaluation: from precision, recall and F-measure to ROC, informedness, marked- ness and correlation”. In:arXivabs/2010.16061 (2020). arXiv: 2010 . 16061.URL: https : / / doi . org / 10 . 48550 / arXiv.2010.16061
2010 arXiv
-
[26]
Performance Measures for Binary Clas- sification
Daniel Berrar. “Performance Measures for Binary Clas- sification”. In:Encycl. Bioinform. Comput. Biology (2019), pp. 546–560.URL: https : / / doi . org / 10 . 1016 / B978-0-12-809633-8.20351-8
2019
-
[27]
A systematic analysis of performance measures for classification tasks
Marina Sokolova and Guy Lapalme. “A systematic analysis of performance measures for classification tasks”. In:Inf. Process. & Manag.45.4 (July 2009), pp. 427–437.URL: https://doi.org/10.1016/j.ipm.2009. 03.002
2009 doi
-
[28]
An Analysis of Performance Measures for Binary Classifiers
Charles Parker. “An Analysis of Performance Measures for Binary Classifiers”. In:IEEE Int. Conf. Data Min. Vancouver, Can.: IEEE, Dec. 2011, pp. 517–526.URL: https://doi.org/10.1109/ICDM.2011.21
2011 doi
-
[29]
An experimental comparison of performance measures for classification
C `esar Ferri, Jos ´e Hern´andez-Orallo, and Ramona Mod- roiu. “An experimental comparison of performance measures for classification”. In:Pattern Recognit. Lett. 30.1 (Jan. 2009), pp. 27–38.URL: https://doi.org/10. 1016/j.patrec.2008.08.010
2009
-
[30]
Beweis der invarianz desn- dimensionalen gebiets
Luitzen E. J. Brouwer. “Beweis der invarianz desn- dimensionalen gebiets”. In:Math. Ann.71.3 (Sept. 1911), pp. 305–313.URL: https : / / doi . org / 10 . 1007 / BF01456846
1911
-
[31]
Zur Invarianz desn- dimensionalen Gebiets
Luitzen E. J. Brouwer. “Zur Invarianz desn- dimensionalen Gebiets”. In:Math. Ann. 72.1 (Mar. 1912), pp. 55–56.URL: https : //doi.org/10.1007/BF01456889
1912 doi
-
[32]
Ad- vancing Precision, Recall, F-score, and Jaccard index: An approach for continuous, ratio-scale measurements
Katarzyna Krasnodebska, Wojciech Goch, Johannes H. Uhl, Judith A. Verstegen, and Martino Pesaresi. “Ad- vancing Precision, Recall, F-score, and Jaccard index: An approach for continuous, ratio-scale measurements”. In:Environ. Model. & Softw.193 (Sept. 2025), pp. 1–9. URL: http...
2025 doi
-
[33]
Frequentist Probability and Frequentist Statistics
Jerzy Neyman. “Frequentist Probability and Frequentist Statistics”. In:Synthese36.1 (1977), pp. 97–131.URL: https://www.jstor.org/stable/20115217
1977
-
[34]
The Relationship Between Precision-Recall and ROC Curves
Jesse Davis and Mark Goadrich. “The Relationship Between Precision-Recall and ROC Curves”. In:Int. Conf. Mach. Learn. (ICML). Pittsburgh, Pennsylvania: ML Res. Press, June 2006, pp. 233–240
2006
-
[35]
Sum- marizing the performances of a background subtraction algorithm measured on several videos
S ´ebastien Pi´erard and Marc Van Droogenbroeck. “Sum- marizing the performances of a background subtraction algorithm measured on several videos”. In:IEEE Int. Conf. Image Process. (ICIP). Abu Dhabi, United Arab. Emir., Oct. 2020, pp. 3234–3238.URL: https://doi.org/ 10.1109/I...
2020 doi
-
[36]
Unachievable Region in Precision-Recall Space 18 and Its Effect on Empirical Evaluation
Kendrick Boyd, Victor Costa, Jesse Davis, and David Page. “Unachievable Region in Precision-Recall Space 18 and Its Effect on Empirical Evaluation”. In:Int. Conf. Mach. Learn. (ICML). Edinburgh, UK, June 2012, pp. 639–646
2012
-
[37]
Multi-domain performance analysis with scores tailored to user preferences
S ´ebastien Pi ´erard, Adrien Deli `ege, and Marc Van Droogenbroeck. “Multi-domain performance analysis with scores tailored to user preferences”. In:arXiv abs/2512.08715 (2025). arXiv: 2512.08715.URL: https: //doi.org/10.48550/arXiv.2512.08715
2025 doi
-
[38]
Unpublished work, submitted to ESANN
S ´ebastien Pi ´erard, Adrien Deli `ege, and Marc Van Droogenbroeck.Multi-domain performance analysis with scores tailored to user preferences. Unpublished work, submitted to ESANN. 2025
2025
-
[39]
A Hitchhiker’s Guide to Understanding Performances of Two-Class Classifiers
Ana ¨ıs Halin, S ´ebastien Pi ´erard, Anthony Cioppa, and Marc Van Droogenbroeck. “A Hitchhiker’s Guide to Understanding Performances of Two-Class Classifiers”. In:arXivabs/2412.04377 (2024). arXiv: 2412.04377. URL: https://doi.org/10.48550/arXiv.2412.04377
-
[40]
A Methodology to Evaluate Strategies Predicting Rankings on Unseen Domains
S ´ebastien Pi ´erard, Adrien Deli `ege, Ana ¨ıs Halin, and Marc Van Droogenbroeck. “A Methodology to Evaluate Strategies Predicting Rankings on Unseen Domains”. In:IEEE Int. Conf. Multimedia Expo Work. (ICMEW), Work. Big Surveill. Data Anal. Process. (big-surv). Nantes, Franc...
2025 doi
-
[41]
Model Cards for Model Reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. “Model Cards for Model Reporting”. In:Proc. Conf. Fairness, Accountability, Transpar.Atlanta, GA, USA: ACM, Jan. 2019, pp. 220–...
2019
-
[42]
Sorbetto: a Python library for producing classification tiles with different flavors to visualize, analyze, and compare performances in two- class classification problems
S ´ebastien Pi´erard, Ana¨ıs Halin, Franc ¸ois Marelli, Simon Pernas, and J ´erˆome Pierre. “Sorbetto: a Python library for producing classification tiles with different flavors to visualize, analyze, and compare performances in two- class classification problems”. In:Zenodo(S...
2025 doi
-
[43]
Tornado predictions
John P. Finley. “Tornado predictions”. In:Am. Meteorol. J.1.3 (July 1884), pp. 85–88.URL: https://archive.org/ details/sim american-meteorological-journal 1884-07 1 3/page/84/mode/2up
-
[44]
Communications Through Limited Response Ques- tioning
Edward M. Bennett, Renee Alpert, and A. C. Goldstein. “Communications Through Limited Response Ques- tioning”. In:Public Opin. Q.18.3 (1954), pp. 303–308. URL: https://doi.org/10.1086/266520
1954 doi
-
[45]
Reliability of Content Analysis: The Case of Nominal Scale Coding
William A. Scott. “Reliability of Content Analysis: The Case of Nominal Scale Coding”. In:Public Opin. Q. 19.3 (1955), pp. 321–325.URL: https : / / doi . org / 10 . 1086/266577
1955
-
[46]
A Coefficient of Agreement for Nom- inal Scales
Jacob Cohen. “A Coefficient of Agreement for Nom- inal Scales”. In:Educ. Psychol. Meas.20.1 (Apr. 1960), pp. 37–46.URL: https : / / doi . org / 10 . 1177 / 001316446002000104
1960
-
[47]
A Fallacy in the Use of Skill Scores
Herbert S. Appleman. “A Fallacy in the Use of Skill Scores”. In:Bull. Am. Meteorol. Soc.41.2 (Feb. 1960), pp. 64–67.URL: https://doi.org/10.1175/1520- 0477- 41.2.64
1960 doi
-
[48]
Finley’s tornado predictions
Grove Karl Gilbert. “Finley’s tornado predictions”. In: Am. Meteorol. J.1.5 (1884), pp. 166–172
-
[49]
The numerical measure of the success of predictions
Charles S. Peirce. “The numerical measure of the success of predictions”. In:Science4.93 (Nov. 1884), pp. 453–454.URL: http://www.jstor.org/stable/1760565
-
[50]
On the association of attributes in statistics: with illustrations from the material of the childhood society, &c
George Udny Yule. “On the association of attributes in statistics: with illustrations from the material of the childhood society, &c.” In:Philosophical Transactions of the Royal Society of London. Series A, Contain- ing Papers of a Mathematical or Physical Character 194.252-26...
1900 doi
-
[51]
Berechnung Des Erfolges Und Der G ¨ute Der Windst ¨arkevorhersagen Im Sturmwarnungsdienst
Paul Heidke. “Berechnung Des Erfolges Und Der G ¨ute Der Windst ¨arkevorhersagen Im Sturmwarnungsdienst”. In:Geografiska Annaler8.4 (Dec. 1926), pp. 301–349. URL: https://doi.org/10.1080/20014422.1926.11881138
1926 doi
-
[52]
Rating weather forecasts
H. Helm Clayton. “Rating weather forecasts”. In:Bull. Am. Meteorol. Soc.15.12 (Dec. 1934), pp. 279–283. URL: https://doi.org/10.1175/1520-0477-15.12.279
1934 doi
-
[53]
The theory of signal detectability
Wesley W. Peterson, Theodore G. Birdsall, and W. C. Fox. “The theory of signal detectability”. In:Trans. IRE Prof. Group Inf. Theory4.4 (Sept. 1954), pp. 171–212. URL: https://doi.org/10.1109/TIT.1954.1057460
1954 doi
-
[54]
Measuring the accuracy of diagnostic systems
John A. Swets. “Measuring the accuracy of diagnostic systems”. In:Science240 (1988), pp. 1285–1293
1988
-
[55]
Area under the Precision-Recall Curve: Point Estimates and Confidence Intervals
Kendrick Boyd, Kevin Eng, and David Page. “Area under the Precision-Recall Curve: Point Estimates and Confidence Intervals”. In:Eur. Conf. Mach. Learn. Princ. Pr. Knowl. Discov. Databases (ECML/PKDD). V ol. 8190. Lect. Notes Comput. Sci. Prague, Czech Repub.: Springer, Sept. 2...
2013 doi
-
[56]
The Geometry of ROC Space: Under- standing Machine Learning Metrics through ROC Iso- metrics
Peter A. Flach. “The Geometry of ROC Space: Under- standing Machine Learning Metrics through ROC Iso- metrics”. In:Int. Conf. Mach. Learn. (ICML). Washing- ton, DC, USA: ML Res. Press, Aug. 2003, pp. 194–201
2003
-
[57]
A Novel Video Dataset for Change Detection Benchmarking
Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, and Prakash Ishwar. “A Novel Video Dataset for Change Detection Benchmarking”. In:IEEE Trans. Image Process.23.11 (Nov. 2014), pp. 4663–4679. URL: https://doi.org/10.1109/TIP.2014.2346013
2014 doi
-
[58]
An introduction to ROC analysis
Tom Fawcett. “An introduction to ROC analysis”. In: Pattern Recognit. Lett.27.8 (June 2006), pp. 861–874. URL: https://doi.org/10.1016/j.patrec.2005.10.010
2006 doi
-
[59]
What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is RarelyF 1
S ´ebastien Pi ´erard, Adrien Deli `ege, and Marc Van Droogenbroeck. “What Is the Optimal Ranking Score Between Precision and Recall? We Can Always Find It and It Is RarelyF 1”. In:IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). Denver, CO, USA: IEEE, June 2026
2026
-
[60]
A unifying view on dataset shift in classification
Jose G. Moreno-Torres, Troy Raeder, Roc ´ıo Alaiz- Rodr´ıguez, Nitesh V . Chawla, and Francisco Herrera. “A unifying view on dataset shift in classification”. In: Pattern Recognit.45.1 (Jan. 2012), pp. 521–530.URL: https://doi.org/10.1016/j.patcog.2011.06.019
2012 doi
-
[61]
Official Journal of the European Union, OJ L, 12.7.2024
European Parliament and Council of the European Union.Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying 19 down harmonised rules on artificial intelligence (Artifi- cial Intelligence Act). Official Journal of the European Union, OJ ...
2024
-
[62]
Detecting moving shadows: algorithms and evaluation
Andrea Prati, Ivana Mikic, Mohan M. Trivedi, and Rita Cucchiara. “Detecting moving shadows: algorithms and evaluation”. In:IEEE Trans. Pattern Anal. Mach. Intell. 25.7 (July 2003), pp. 918–923.URL: https://doi.org/10. 1109/TPAMI.2003.1206520
2003 arXiv
-
[63]
A Two-Stage Template Approach to Person Detection in Thermal Imagery
James W. Davis and Mark A. Keck. “A Two-Stage Template Approach to Person Detection in Thermal Imagery”. In:IEEE Workshops on Applications of Com- puter Vision (WACV/MOTION). V ol. 1. Breckenridge, CO, USA: IEEE, Jan. 2005, pp. 364–369.URL: https: //doi.org/10.1109/ACVMOT.2005.14
2005 doi
-
[64]
Web site https:// vcipl-okstate.org/pbvs/bench/
Roland Miezianko.IEEE OTCBVS WS Series Bench: Terravic Research Infrared Database. Web site https:// vcipl-okstate.org/pbvs/bench/. 2005.URL: https://vcipl- okstate.org/pbvs/bench/
2005
-
[65]
ETISEO, performance eval- uation for video surveillance systems
Anh-Tuan Nghiem, Franc ¸ois Bremond, Monique Thon- nat, and Val ´ery Valentin. “ETISEO, performance eval- uation for video surveillance systems”. In:IEEE Int. Conf. Adv. Video Signal Based Surveill. (AVSS). IEEE, Sept. 2007, pp. 476–481.URL: https://doi.org/10.1109/ A VSS.2007.4425357
2007
-
[66]
Modeling, Clustering, and Segmenting Video with Mixtures of Dy- namic Textures
Antoni B. Chan and Nuno Vasconcelos. “Modeling, Clustering, and Segmenting Video with Mixtures of Dy- namic Textures”. In:IEEE Trans. Pattern Anal. Mach. Intell.30.5 (May 2008), pp. 909–926.URL: https://doi. org/10.1109/TPAMI.2007.70738
2008 doi
-
[67]
Change Detection in Optical Aerial Images by a Multilayer Conditional Mixed Markov Model
Csaba Benedek and Tam ´as Szir´anyi. “Change Detection in Optical Aerial Images by a Multilayer Conditional Mixed Markov Model”. In:IEEE Trans. Geosci. Remote Sens.47.10 (Oct. 2009), pp. 3416–3430.URL: https : //doi.org/10.1109/TGRS.2009.2022633
2009 doi
-
[68]
A large-scale benchmark dataset for event recognition in surveillance video
Sangmin Oh, Anthony Hoogs, Amitha Perera, Naresh Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, J. K. Aggarwal, Hyungtae Lee, Larry Davis, Eran Swears, Xioyang Wang, Qiang Ji, Kishore Reddy, Mubarak Shah, Carl V ondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen,...
2011 doi
-
[69]
Real time moving vehicle detection and reconstruction for improving classifica- tion
Tao Wang and Zhigang Zhu. “Real time moving vehicle detection and reconstruction for improving classifica- tion”. In:IEEE Work. Appl. Comput. Vis. (WACV). Breckenridge, CO, USA: IEEE, Jan. 2012, pp. 497–502. URL: https://doi.org/10.1109/W ACV .2012.6163039
2012 doi
-
[70]
Background Subtraction Based on Color and Depth Using Active Sensors
Enrique Fernandez-Sanchez, Javier Diaz, and Eduardo Ros. “Background Subtraction Based on Color and Depth Using Active Sensors”. In:Sensors13.7 (July 2013), pp. 1–21.URL: https : / / doi . org / 10 . 3390 / s130708895
2013
-
[71]
Person Re-identification by Video Ranking
Taiqing Wang, Shaogang Gong, Xiatian Zhu, and Shengjin Wang. “Person Re-identification by Video Ranking”. In:Eur. Conf. Comput. Vis. (ECCV). V ol. 8692. Lect. Notes Comput. Sci. Springer Int. Publ., 2014, pp. 688–703.URL: https://doi.org/10.1007/978- 3-319-10593-2 45
2014 doi
-
[72]
Towards Benchmarking Scene Background Initialization
Lucia Maddalena and Alfredo Petrosino. “Towards Benchmarking Scene Background Initialization”. In: Int. Conf. Image Anal. Process. Work. (ICIAP Work.) V ol. 9281. Lect. Notes Comput. Sci. Springer Int. Publ., Sept. 2015, pp. 469–476.URL: https://doi.org/10.1007/ 978-3-319-23222-5 57
2015
-
[73]
PETS 2016: Dataset and Challenge
Luis Patino, Tom Cane, Alain Vallee, and James Ferryman. “PETS 2016: Dataset and Challenge”. In: IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Work. (CVPRW). Las Vegas, NV , USA: IEEE, June 2016, pp. 1240–1247.URL: https://doi.org/10.1109/CVPRW. 2016.157
2016 doi
-
[74]
Weighted Low-rank Decomposi- tion for Robust Grayscale-Thermal Foreground Detec- tion
Chenglong Li, Xiao Wang, Lei Zhang, Jin Tang, Hejun Wu, and Liang Lin. “Weighted Low-rank Decomposi- tion for Robust Grayscale-Thermal Foreground Detec- tion”. In:IEEE Trans. Circuits Syst. Video Technol.27.4 (Apr. 2017), pp. 725–738.URL: https://doi.org/10.1109/ TCSVT.2016.2556586
2017
-
[75]
The SYNTHIA Dataset: A Large Collection of Synthetic Images for Se- mantic Segmentation of Urban Scenes
German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio M. Lopez. “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Se- mantic Segmentation of Urban Scenes”. In:IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR). Las Vegas, NV , USA: IEEE, June 2...
2016 doi
-
[76]
Extensive Benchmark and Survey of Modeling Methods for Scene Background Initial- ization
Pierre-Marc Jodoin, Lucia Maddalena, Alfredo Pet- rosino, and Yi Wang. “Extensive Benchmark and Survey of Modeling Methods for Scene Background Initial- ization”. In:IEEE Trans. Image Process.26.11 (Nov. 2017), pp. 5244–5256.URL: https://doi.org/10.1109/ TIP.2017.2728181
2017
-
[77]
Comparative Evaluation of Background Subtraction Algorithms in Remote Scene Videos Cap- tured by MWIR Sensors
Guangle Yao, Tao Lei, Jiandan Zhong, Ping Jiang, and Wenwu Jia. “Comparative Evaluation of Background Subtraction Algorithms in Remote Scene Videos Cap- tured by MWIR Sensors”. In:Sensors17.9 (Aug. 2017), pp. 1–31.URL: https://doi.org/10.3390/s17091945
2017 doi
-
[78]
A Benchmarking Framework for Background Subtraction in RGBD Videos
Massimo Camplani, Lucia Maddalena, Gabriel Moy ´a Alcover, Alfredo Petrosino, and Luis Salgado. “A Benchmarking Framework for Background Subtraction in RGBD Videos”. In:Int. Conf. Image Anal. Process. (ICIAP). V ol. 10590. Lect. Notes Comput. Sci. Springer Int. Publ., 2017, pp...
2017
-
[79]
MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?
Matteo Fabbri, Guillem Braso, Gianluca Maugeri, Or- cun Cetintas, Riccardo Gasparini, Aljosa Osep, Si- mone Calderara, Laura Leal-Taixe, and Rita Cucchiara. “MOTSynth: How Can Synthetic Data Help Pedestrian Detection and Tracking?” In:IEEE/CVF Int. Conf. Comput. Vis. (ICCV). M...
2021
-
[80]
AGVS: A New Change Detection Dataset for Airport Ground Video Surveillance
Xiang Zhang, Chang Shu, Shuai Li, Celimuge Wu, and Zhi Liu. “AGVS: A New Change Detection Dataset for Airport Ground Video Surveillance”. In:IEEE Trans. Intell. Transp. Syst.23.11 (Nov. 2022), pp. 20588– 20600.URL: https://doi.org/10.1109/tits.2022.3184978
2022 doi
-
[81]
eMammal–citizen science camera trapping as a solu- tion for broad-scale, long-term monitoring of wildlife populations
Tavis Forrester, William J. McShea, R. W. Keys, Robert Costello, Megan Baker, and Arielle Parsons. “eMammal–citizen science camera trapping as a solu- tion for broad-scale, long-term monitoring of wildlife populations”. In:Annual Meeting of the Ecological So- ciety of America ...
2013
-
[82]
Recognition in Terra Incognita
Sara Beery, Grant Van Horn, and Pietro Perona. “Recognition in Terra Incognita”. In:Eur. Conf. Com- put. Vis. (ECCV). V ol. 11220. Lect. Notes Comput. Sci. Springer Int. Publ., 2018, pp. 472–489.URL: https:// doi.org/10.1007/978-3-030-01270-0 28
2018 doi
-
[83]
A Fusion Framework for Camouflaged Moving Foreground Detection in the Wavelet Domain
Shuai Li, Dinei Florencio, Wanqing Li, Yaqin Zhao, and Chris Cook. “A Fusion Framework for Camouflaged Moving Foreground Detection in the Wavelet Domain”. In:IEEE Trans. Image Process.27.8 (Aug. 2018), pp. 3918–3930.URL: https : / / doi . org / 10 . 1109 / TIP. 2018.2828329
2018
-
[84]
Foreground detection in camouflaged scenes
Shuai Li, Dinei Florencio, Yaqin Zhao, Chris Cook, and Wanqing Li. “Foreground detection in camouflaged scenes”. In:IEEE Int. Conf. Image Process. (ICIP). Beijing, China: IEEE, Sept. 2017, pp. 4247–4251.URL: https://doi.org/10.1109/ICIP.2017.8297083
2017 doi
-
[85]
Comparison of Background Subtraction Methods for a Multimedia Application
Fida El Baf, Thierry Bouwmans, and Bertrand Vachon. “Comparison of Background Subtraction Methods for a Multimedia Application”. In:Int. Work. Syst. Signals Image Process.Maribor, Slovenia: IEEE, June 2007, pp. 385–388.URL: https://doi.org/10.1109/IWSSIP. 2007.4381122
2007 doi
-
[86]
Fisher, Yun-Heh Chen-Burger, Daniela Giordano, Lynda Hardman, and Fang-Pang Lin, eds
Robert B. Fisher, Yun-Heh Chen-Burger, Daniela Giordano, Lynda Hardman, and Fang-Pang Lin, eds. Fish4Knowledge: Collecting and Analyzing Massive Coral Reef Fish Video Data. V ol. 104. Intell. Syst. Ref. Libr. Springer Int. Publ., 2016.URL: https://doi.org/10. 1007/978-3-319-30208-9
2016
-
[87]
Real-World Underwater Enhance- ment: Challenges, Benchmarks, and Solutions Under Natural Light
Risheng Liu, Xin Fan, Ming Zhu, Minjun Hou, and Zhongxuan Luo. “Real-World Underwater Enhance- ment: Challenges, Benchmarks, and Solutions Under Natural Light”. In:IEEE Trans. Circuits Syst. Video Technol.30.12 (Dec. 2020), pp. 4861–4875.URL: https: //doi.org/10.1109/TCSVT.201...
2020 doi
-
[88]
ARGOS-Venice Boat Classification
Domenico D. Bloisi, Luca Iocchi, Andrea Pennisi, and Luigi Tombolini. “ARGOS-Venice Boat Classification”. In:IEEE Int. Conf. Adv. Video Signal Based Surveill. (AVSS). Karlsruhe, Germany: IEEE, Aug. 2015, pp. 1–6. URL: https://doi.org/10.1109/A VSS.2015.7301727
2015 doi
-
[89]
Fast Image-Based Obstacle Detection From Unmanned Surface Vehicles
Matej Kristan, Vildana Suli ´c Kenk, Stanislav Kova ˇciˇc, and Janez Per ˇs. “Fast Image-Based Obstacle Detection From Unmanned Surface Vehicles”. In:IEEE Trans. Cybern.46.3 (Mar. 2016), pp. 641–654.URL: https : //doi.org/10.1109/TCYB.2015.2412251
2016 doi
-
[90]
Video Processing From Electro-Optical Sensors for Object Detection and Track- ing in a Maritime Environment: A Survey
Dilip K. Prasad, Deepu Rajan, Lily Rachmawati, Eshan Rajabally, and Chai Quek. “Video Processing From Electro-Optical Sensors for Object Detection and Track- ing in a Maritime Environment: A Survey”. In:IEEE Trans. Intell. Transp. Syst.18.8 (Aug. 2017), pp. 1993– 2016.URL: htt...
2017 doi
-
[91]
SeaShips: A Large-Scale Pre- cisely Annotated Dataset for Ship Detection
Zhenfeng Shao, Wenjing Wu, Zhongyuan Wang, Wan Du, and Chengyuan Li. “SeaShips: A Large-Scale Pre- cisely Annotated Dataset for Ship Detection”. In:IEEE Trans. Multimedia20.10 (Oct. 2018), pp. 2593–2604. URL: https://doi.org/10.1109/TMM.2018.2865686
2018 doi
-
[92]
Automatic Ship Classification from Optical Aerial Im- ages with Convolutional Neural Networks
Antonio-Javier Gallego, Antonio Pertusa, and Pablo Gil. “Automatic Ship Classification from Optical Aerial Im- ages with Convolutional Neural Networks”. In:Remote Sens.10.4 (Mar. 2018), pp. 1–20.URL: https://doi.org/ 10.3390/rs10040511
2018 doi
-
[93]
Real-Time Ship Segmentation in Maritime Surveillance Videos Using Automatically Annotated Synthetic Datasets
Miguel Ribeiro, Bruno Damas, and Alexandre Bernardino. “Real-Time Ship Segmentation in Maritime Surveillance Videos Using Automatically Annotated Synthetic Datasets”. In:Sensors22.21 (Oct. 2022), pp. 1–18.URL: https://doi.org/10.3390/s22218090
2022 doi
-
[94]
Mass- MIND: Massachusetts Maritime INfrared Dataset
Shailesh Nirgudkar, Michael DeFilippo, Michael Sacarny, Michael Benjamin, and Paul Robinette. “Mass- MIND: Massachusetts Maritime INfrared Dataset”. In: Int. J. Robot. Res.42.1-2 (Jan. 2023), pp. 21–32.URL: https://doi.org/10.1177/02783649231153020
2023 doi
-
[95]
Modality-Aware In- frared and Visible Image Fusion with Target-Aware Supervision
Tianyao Sun, Dawei Xiang, Tianqi Ding, Xiang Fang, Yijiashun Qi, and Zunduo Zhao. “Modality-Aware In- frared and Visible Image Fusion with Target-Aware Supervision”. In:Int. Conf. Comput. Vis. Data Min. (IC- CVDM). London, Engl.: IEEE, Sept. 2025, pp. 180–184. URL: https://doi...
2025 doi
-
[96]
It focuses on outdoor surveillance scenarios affected by weather, illumination changes, dynamic back- grounds, and shadows
provides real and synthetic video sequences for eval- uating background subtraction and foreground detection algorithms. It focuses on outdoor surveillance scenarios affected by weather, illumination changes, dynamic back- grounds, and shadows. •Audio-Visual Vehicle (A VV) (20...
2012
-
[97]
It is useful in urban surveillance because foreground objects may remain static for long periods, creating intermittent motion and background-initialization challenges
is designed for evaluating methods that estimate a clean background image from video sequences. It is useful in urban surveillance because foreground objects may remain static for long periods, creating intermittent motion and background-initialization challenges. •PETS (2016)...
2016
-
[98]
It includes challenges such as low contrast, video noise, dynamic background, camouflage, and varying foreground motion
provides infrared video sequences captured from re- mote scenes for evaluating background subtraction meth- ods. It includes challenges such as low contrast, video noise, dynamic background, camouflage, and varying foreground motion. •SBM-RGBD (2017):The SBM-RGBD dataset [78] ...
2017
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.