Pith. sign in

REVIEW 4 major objections 6 minor 47 references

MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that a unified, publicly checkable framework for feature selection evaluation can replace fragmented, proprietary benchmarking in Android malware detection, and that under identical conditions LASSO, RFE, and SigAPI are…

desk verdict A useful benchmarking artifact whose result matrix is broader than anything in the cited literature, but the central ranking depends on unvalidated reimplementations and the paper contains an internal contradiction about which methods were actually evaluated. read the letter →

arxiv 2507.10591 v1 pith:62CVTAWY submitted 2025-07-11 cs.LG cs.AIcs.CRcs.PF

classification cs.LGcs.AIcs.CRcs.PF
keywords featureselectionAndroidmalwaredetectionbenchmarkingreproducibilityclassimbalanceLASSOdomain-specificmethodsevaluationframework
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that feature selection research in Android malware detection has been held back by two fixable problems: evaluations rely on private datasets, and new methods are compared against only a handful of similar baselines. To fix this, the authors introduce MH-FSF, a modular, publicly available framework that reimplements 17 feature selection methods—11 classical and 6 domain-specific—and runs them through the same data preparation, model training, and evaluation pipeline on 10 public Android malware datasets. The comparative results show that no method wins everywhere: performance shifts between balanced and imbalanced data, and the most consistently reliable methods are LASSO, RFE, and SigAPI. If the framework is taken up, the field gains a common yardstick for new methods and a checkable record of how published techniques behave outside their original private datasets.

What carries the argument

The load-bearing mechanism is the MH-FSF pipeline, a four-stage architecture: data manipulation (NaN and duplicate removal, balancing, sampling), feature selection, classifier training and evaluation, and visualization. Each feature selection method is isolated as an independent module with a standard interface, so new techniques can be added without structural changes. The comparison is carried by this uniformity: all 17 methods feed the same reduced datasets to the same three classifiers (KNN, Random Forest, SVM) under stratified 5-fold cross-validation, and performance is aggregated over accuracy, precision, recall, F1, ROC-AUC, and MCC.

What would settle it

Re-run the six domain-specific methods on the datasets named in their original papers and compare the selected feature sets or reported F1 and recall values with those in the original publications; if the reproduced outputs differ materially, the MH-FSF comparison measures the reconstructions rather than the published methods.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a single reproducible evaluation platform can change what we believe about feature selection in this domain. Under identical conditions across 10 datasets and three classifiers, classical methods LASSO and RFE, together with the API-focused domain-specific method SigAPI, deliver the highest average F1 and recall with the least variability, while methods like PCA, ReliefF, and SigPID degrade markedly on imbalanced data. The paper further finds that several domain-specific methods, originally validated on single proprietary datasets, do not generalize: JOWMDroid and SigPID perform poorly outside their original feature types, suggesting that earlier isolated comparisons overstated their value. The framework claim is that MH-FSF itself—with 17 implementations, 10 public datasets, and all results released—provides the infrastructure to make such comparisons routine and reproducible.

Load-bearing premise

The paper's ranking stands on the assumption that its implementations of the six domain-specific methods faithfully reproduce the published methods, since it reports no comparison of the reproduced feature sets or scores against the originals.

Editorial extensions

If this is right

  • Researchers can now check any of the 17 methods against the same 10 public datasets, so new proposals need not rely on private data to claim an improvement.
  • The ranking implies that LASSO and RFE are the safe default choices for Android malware detection, with SigAPI the strongest domain-specific option on API-call datasets.
  • Because balancing changes which methods win, the paper's results imply that class imbalance should be reported and treated as a first-class experimental variable in future feature selection studies.
  • Methods validated only on one private dataset, such as JOWMDroid and SigPID, should be re-evaluated on diverse public data before being adopted.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: if the framework is applied to other domains it claims to support—network traffic, biomedical data, fraud detection—the same protocol could expose which classical methods are domain-robust, something the current Android-only evaluation cannot show.
  • The paper does not compare its reimplementations against the original methods' reported feature sets, so a fair reading is that its rankings measure the reconstructions unless such validation is added.
  • A testable extension is to treat dataset balancing as an intervention: the MCC heatmaps suggest balancing helps most methods, so a controlled study varying only the balancing rule could identify which selection methods benefit most.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents MH-FSF, a modular framework for feature selection evaluation, and uses it to compare 17 feature selection methods (11 classical and 6 domain-specific) on 10 publicly available Android malware datasets, using three classifiers and metrics including MCC. The authors report that LASSO, RFE, and SigAPI are the most reliable methods across balanced and imbalanced datasets, and they argue that class balancing is critical for obtaining consistent feature-selection performance. The framework and full results are said to be publicly available on GitHub.

Significance. If the claims hold, the paper would provide a valuable reproducible benchmark for feature selection in Android malware detection, with a public, extensible framework integrating both classical and domain-specific methods. The authors are explicit about using public datasets and default scikit-learn configurations, and they provide MCC heatmaps and per-method boxplots, which are welcome additions. However, the empirical conclusions currently rest on several load-bearing methodological details that are either unspecified or internally inconsistent, so the benchmark's value will be fully realized only after these points are resolved.

major comments (4)
  1. [Section III, Table V, Figures 2-4] The text states that ABC and LR were 'excluded from the current analysis' because of inferior performance, yet Table V and Figures 2-4 report results for both ABC and LR. This is a direct internal contradiction about which 17 methods were actually evaluated; the authors must reconcile the roster and remove or correct the exclusion statement.
  2. [Section V and Figure 1] The experimental protocol does not state whether feature selection is performed inside or outside the cross-validation loop. The pipeline in Figure 1 places 'Features Selection' before 'Model Training and Evaluation,' and Section V says the three classifiers are used to evaluate 'the datasets resulting from feature selection,' suggesting the feature subsets were derived from the full datasets and then evaluated with stratified 5-fold CV. If that is the case, the test folds have influenced feature selection, producing optimistically biased metrics and invalidating the relative ranking of methods. The paper must specify the exact placement of feature selection relative to CV; if it occurs before CV, the experiments should be rerun with feature selection inside each training fold.
  3. [Section V and Table IV] The paper's central conclusion about the benefits of class balancing cannot be verified because no balancing method is described. Section V only mentions stratified cross-validation and the use of default scikit-learn configurations; the procedure used to produce the 'Balanced Per Class' counts in Table IV (e.g., random undersampling of the majority class, SMOTE, or another resampling scheme) is never stated. This omission is a reproducibility blocker for all balanced-dataset results, which are a key component of the empirical claims.
  4. [Section III and Section VI] The fidelity of the six domain-specific reimplementations is supported only by an internal three-researcher review, with no quantitative comparison against the feature sets or performance numbers reported in the original papers (SemiDroid [4], RFG [7], JOWNDroid [8], MT [9], SigPID [10], SigAPI [6]). Because SigAPI is named in the headline result (Section VI), the ranking of domain-specific versus classical methods measures the authors' reconstructions unless external validation is provided. Include a comparison of reproduced outputs (e.g., selected features or classification performance on the original datasets) or explicitly limit the claims to the implementations as provided.
minor comments (6)
  1. [Figure 2 caption] The caption contains a typo: 'distribuition' should be 'distribution'.
  2. [Figure 1] The figure contains 'KroinoDroid R' while the text and tables use 'KronoDroid R'; please unify the spelling.
  3. [Section III] The sentence mentioning FSDroid as an additional framework method is confusing because FSDroid is not listed in Table II and is never defined; please explain or remove the reference.
  4. [Tables and text] The method name 'ANOV A' appears with an extra space throughout (e.g., Table II and Table V); this should be corrected to 'ANOVA'.
  5. [Section V and Table IV] The exact versions of the public datasets are not given; for example, KronoDroid has multiple releases, and the paper should identify the specific versions used, even if they are available in the repository.
  6. [Section VII] The phrase 'in presence of class imbalance' should be 'in the presence of class imbalance'; similar small grammatical fixes are needed in several places.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the ranking is an empirical measurement on public data and held-out folds, not a quantity forced by the framework's definitions.

full rationale

This paper contains no formal derivation chain. Its central outputs are (1) a modular software framework and (2) an empirical comparison of 17 feature-selection methods on 10 public Android malware datasets. The ranking of LASSO, RFE, and SigAPI as most reliable is obtained by running the methods, training KNN/RF/SVM classifiers under stratified 5-fold cross-validation, and aggregating F1, recall, and MCC scores. Nothing in Sections V-VI fits a parameter to the reported outcome, and no method's performance is defined in terms of the conclusion. The self-citations to references [31]-[33] document the framework's provenance, the reproduction workflow, and the earlier exclusion of FSDroid, but these are roster choices, not predictions, and the current numbers are independently measured on held-out data. The lack of external validation of the six domain-specific reimplementations is a legitimate reproducibility and fidelity concern, but it is not circularity: the comparison measures the authors' implementations, and the paper does not claim the result is forced by construction. The internal inconsistency about ABC and LR (Section III says they were excluded while Table V includes them) affects reliability and clarity but does not make the result equivalent to its inputs. Verdict: no significant circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central comparison rests on two unspecified choices (feature retention counts per method, and the class balancing scheme) and on four background assumptions: library correctness, dataset ground truth, fairness of default classifier hyperparameters, and the adequacy of internal review as validation of the six reimplemented domain-specific methods. No invented entities are introduced; MH-FSF is software, not a postulated theory.

free parameters (2)
  • Number of features retained by each selection method = unspecified; implementation defaults
    RFE, LASSO, PCA, ANOVA, and the domain-specific subset selectors each require a target feature count or threshold. The paper never states these settings, yet the reduced datasets that feed the classifiers are direct outputs of these choices.
  • Class balancing scheme = minority-class size per dataset (inferred from Table IV)
    The balanced per-class counts in Table IV equal the minority class size, implying undersampling to the minority count, but the paper never states the procedure (undersampling, oversampling, or synthetic), and balancing changes every ranking in Section VI.
assumptions (4)
  • standard math scikit-learn 1.5.2 implementations of classifiers and classical feature selection methods are correct for the intended purpose
    Section V states the scikit-learn version and default configurations; the entire pipeline treats library outputs as ground truth without independent verification.
  • domain assumption The 10 public datasets are correctly labeled and their feature extractions (permissions, API calls, intents, opcodes) faithfully represent malware behavior
    Section V (Table IV) builds all rankings on dataset ground truth; label noise or extraction errors would propagate through every method and classifier.
  • domain assumption Default scikit-learn hyperparameters for KNN, RF, and SVM provide a fair basis for comparing feature selection methods
    Section V: 'we used the default configuration of the scikit-learn library, version 1.5.2'. If defaults advantage or disadvantage certain feature subsets, the method rankings shift.
  • ad hoc to paper The internal three-researcher review process guarantees faithful reimplementation of the six domain-specific methods
    Section III: each reproduction 'undergo review and technical evaluation by at least three researchers'. This substitutes for external validation against the original papers' reported feature sets or results; if the reimplementations drift, the domain-specific versus classical comparison measures the authors' reconstructions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation." pith.science (2026). https://pith.science/paper/62CVTAWY

@misc{pith2026250710591,
  author       = {Pith},
  title        = {Pith review of: MH-FSF: A Unified Framework for Overcoming Benchmarking and Reproducibility Limitations in Feature Selection Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/62CVTAWY}},
  note         = {Machine review of arXiv:2507.10591}
}
read the original abstract

Feature selection is vital for building effective predictive models, as it reduces dimensionality and emphasizes key features. However, current research often suffers from limited benchmarking and reliance on proprietary datasets. This severely hinders reproducibility and can negatively impact overall performance. To address these limitations, we introduce the MH-FSF framework, a comprehensive, modular, and extensible platform designed to facilitate the reproduction and implementation of feature selection methods. Developed through collaborative research, MH-FSF provides implementations of 17 methods (11 classical, 6 domain-specific) and enables systematic evaluation on 10 publicly available Android malware datasets. Our results reveal performance variations across both balanced and imbalanced datasets, highlighting the critical need for data preprocessing and selection criteria that account for these asymmetries. We demonstrate the importance of a unified platform for comparing diverse feature selection techniques, fostering methodological consistency and rigor. By providing this framework, we aim to significantly broaden the existing literature and pave the way for new research directions in feature selection, particularly within the context of Android malware detection.

Figures

Figures reproduced from arXiv: 2507.10591 by the authors.

Figure 1
Figure 1. Overview of MH-FSF: The four main steps of the pipeline. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. , provides clear evidence of the varying effectiveness of feature selection methods across both balanced and com￾plete datasets. LASSO and RFE emerge as top-performing methods, con￾sistently achieving F1 scores and recall values above 0.9. Their boxplots exhibit low variability, with compact interquartile ranges and an absence of significant outliers. This visual consistency reinforces their reliability and robustne… view at source ↗
Figure 3
Figure 3. MCC: Complete Datasets X Feature Selection Methods. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: MCC: Balanced Datasets X Feature Selection Methods. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 46 canonical work pages

  1. [4]

    SemiDroid: A Behavioral Malware Detector Based on Unsupervised Machine Learning Techniques Using Feature Selection Approaches,

    A. Mahindru and A. L. Sangal, “SemiDroid: A Behavioral Malware Detector Based on Unsupervised Machine Learning Techniques Using Feature Selection Approaches,” International Journal of Machine Learn- ing and Cybernetics , vol. 12, no. 5, pp. 1369–1411, 2021

  2. [7]

    Automated Malware Detection in Mobile App Stores Based on Robust Feature Generation,

    M. Alazab, “Automated Malware Detection in Mobile App Stores Based on Robust Feature Generation,” Electronics, vol. 9, p. 435, 03 2020

  3. [8]

    JOWMDroid: Android Malware Detection Based on Feature Weighting with Joint Optimization of Weight-Papping and Classifier Parameters,

    L. Cai, Y . Li, and Z. Xiong, “JOWMDroid: Android Malware Detection Based on Feature Weighting with Joint Optimization of Weight-Papping and Classifier Parameters,” Computers & Security , vol. 100, p. 102086, 2021. [Online]. Available: https://www.sciencedirect. com/science/article/pii/S016740482030359X

  4. [9]

    A Multi-Tiered Feature Selection Model for Android Malware Detection Based on Feature Discrimination and Information Gain,

    P. Bhat and K. Dutta, “A Multi-Tiered Feature Selection Model for Android Malware Detection Based on Feature Discrimination and Information Gain,” Journal of King Saud University - Computer and Information Sciences , vol. 34, no. 10, Part B, pp. 9464–9477, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S1319157821003049

  5. [10]

    SigPID: Significant Per- mission Identification for Android Malware Detection,

    L. Sun, Z. Li, Q. Yan, W. Srisa-an, and Y . Pan, “SigPID: Significant Per- mission Identification for Android Malware Detection,” in MALWARE. IEEE Computer Society, 2016, pp. 1–8

  6. [6]

    Significant API Calls in Android Malware Detection (Using Feature Selection Techniques and Correlation Based Feature Elimination),

    A. H. Galib and M. Hossain, “Significant API Calls in Android Malware Detection (Using Feature Selection Techniques and Correlation Based Feature Elimination),” in International Conference on Software Engineering and Knowledge Engineering (SEKE) , 2020. [Online]. Available: https://api.semanticscholar.org/CorpusID:221193478

  7. [1]

    A Comprehensive Survey on Feature Selection in the Various Fields of Machine Learning,

    P. Dhal and C. Azad, “A Comprehensive Survey on Feature Selection in the Various Fields of Machine Learning,” Applied Intelligence, vol. 52, no. 4, pp. 4543–4581, 2022

  8. [2]

    Importance of Features Selection, Attributes Selection, Challenges and Future Directions for Medical Imaging Data: A Review,

    N. Naheed, M. Shaheen, S. A. Khan et al. , “Importance of Features Selection, Attributes Selection, Challenges and Future Directions for Medical Imaging Data: A Review,” CMES, vol. 125, no. 1, 2020

Show all 47 references
  1. [3]

    PermDroid: A Framework Developed Using Proposed Feature Selection Approach and Machine Learning Techniques for Android Malware Detection,

    A. Mahindru, H. Arora, A. Kumar, S. K. Gupta, S. Mahajan, S. Kadry, and J. Kim, “PermDroid: A Framework Developed Using Proposed Feature Selection Approach and Machine Learning Techniques for Android Malware Detection,” Scientific Reports, vol. 14, no. 1, p. 10724, 2024

  2. [5]

    A New Feature Selection Method Based on a Self-Variant Genetic Algorithm Applied to Android Malware Detection,

    L. Wang, Y . Gao, S. Gao, and X. Yong, “A New Feature Selection Method Based on a Self-Variant Genetic Algorithm Applied to Android Malware Detection,” Symmetry, vol. 13, no. 7, p. 1290, 2021

  3. [11]

    DroidRL: Feature Selection for Android Malware Detection with Reinforcement Learning,

    Y . Wu, M. Li, Q. Zeng, T. Yang, J. Wang, Z. Fang, and L. Cheng, “DroidRL: Feature Selection for Android Malware Detection with Reinforcement Learning,” Computers & Security , vol. 128, p. 103126, 2023

  4. [12]

    A Hybrid Feature Selection Approach-Based Android Malware Detection Framework Us- ing Machine Learning Techniques,

    S. K. Smmarwar, G. P. Gupta, and S. Kumar, “A Hybrid Feature Selection Approach-Based Android Malware Detection Framework Us- ing Machine Learning Techniques,” in Cyber Security, Privacy and Networking: Proceedings of ICSPN 2021 . Springer, 2022, pp. 347– 356

  5. [13]

    BFEDroid: A Feature Selection Technique to Detect Malware in Android Apps Using Machine Learning,

    C. Chimeleze, N. Jamil, R. Ismail, K.-Y . Lam, J. S. Teh, J. Samual, and C. Akachukwu Okeke, “BFEDroid: A Feature Selection Technique to Detect Malware in Android Apps Using Machine Learning,” Security and Communication Networks , vol. 2022, no. 1, p. 5339926, 2022

  6. [14]

    A Novel Android Malware Detection System: Adaption of Filter-based Feature Selection Methods,

    D. ¨O. S ¸ahin, O. E. Kural, S. Akleylek, and E. Kılıc ¸, “A Novel Android Malware Detection System: Adaption of Filter-based Feature Selection Methods,” Journal of Ambient Intelligence and Humanized Computing , pp. 1–15, 2023

  7. [15]

    Captur- ing the Behavior of Android Malware with MH-100K: A Novel and Multidimensional Dataset,

    H. Braganc ¸a, V . Rocha, E. Souto, D. Kreutz, and E. Feitosa, “Captur- ing the Behavior of Android Malware with MH-100K: A Novel and Multidimensional Dataset,” in XXIII SBSeg, 2023

  8. [16]

    Detecc ¸˜ao de Malwares Android: Datasets e Reprodutibilidade,

    T. Soares, G. Siqueira, L. Barcellos et al. , “Detecc ¸˜ao de Malwares Android: Datasets e Reprodutibilidade,” in Anais da XIX ERRC . Porto Alegre, RS, Brasil: SBC, 2021, pp. 43–48. [Online]. Available: https://sol.sbc.org.br/index.php/errc/article/view/18540

  9. [17]

    Debiasing Android Malware Datasets: How Can I Trust Your Results If Your Dataset Is Biased?

    T. C. Miranda, P.-F. Gimenez, J.-F. Lalande, V . V . T. Tong, and P. Wilke, “Debiasing Android Malware Datasets: How Can I Trust Your Results If Your Dataset Is Biased?” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 2182–2197, 2022

  10. [18]

    Effective and Efficient Android Malware Detection and Category Classification Using the Enhanced KronoDroid Dataset,

    M. Waheed and S. Qadir, “Effective and Efficient Android Malware Detection and Category Classification Using the Enhanced KronoDroid Dataset,” Security and Communication Networks , vol. 2024, no. 1, p. 7382302, 2024

  11. [19]

    A Modified ResNeXt for Android Malware Identification and Classification,

    M. A. Albahar, M. S. ElSayed, and A. Jurcut, “A Modified ResNeXt for Android Malware Identification and Classification,” Computational Intelligence and Neuroscience , vol. 2022, no. 1, p. 8634784, 2022

  12. [20]

    AndroOBFS: Time- tagged Obfuscated Android Malware Dataset with Family Information,

    S. Kumar, D. Mishra, B. Panda, and S. K. Shukla, “AndroOBFS: Time- tagged Obfuscated Android Malware Dataset with Family Information,” in Proceedings of the 19th International Conference on Mining Software Repositories, 2022, pp. 454–458

  13. [21]

    Mal- Radar: Demystifying Android Malware in the New Era,

    L. Wang, H. Wang, R. He, R. Tao, G. Meng, X. Luo, and X. Liu, “Mal- Radar: Demystifying Android Malware in the New Era,” Proceedings of the ACM on Measurement and Analysis of Computing Systems , vol. 6, no. 2, pp. 1–27, 2022

  14. [22]

    A Novel Android Malware Detection System: Adaption of Filter-Based Feature Selection Methods,

    D. ¨O. S ¸ahin, O. E. Kural, S. Akleylek, and E. Kılıc ¸, “A Novel Android Malware Detection System: Adaption of Filter-Based Feature Selection Methods,” JAIHC, pp. 1–15, 2023

  15. [23]

    Android Malware Classification Using Optimum Feature Selection and Ensemble Machine Learning,

    R. Islam, M. I. Sayed, S. Saha et al., “Android Malware Classification Using Optimum Feature Selection and Ensemble Machine Learning,” Internet of Things and Cyber-Physical Systems (IOTCPS) , vol. 3, pp. 100–111, 2023

  16. [24]

    Deepdroid: Feature Selection Approach to Detect Android Malware Using Deep Learning,

    A. Mahindru and A. Sangal, “Deepdroid: Feature Selection Approach to Detect Android Malware Using Deep Learning,” in IEEE ICSESS . IEEE, 2019, pp. 16–19

  17. [25]

    Fest: A Feature Extraction and Selection Tool for Android Malware Detection,

    K. Zhao, D. Zhang, X. Su, and W. Li, “Fest: A Feature Extraction and Selection Tool for Android Malware Detection,” in IEEE ISCC. IEEE, 2015, pp. 714–720

  18. [26]

    A Lightweight Android Malware Classifier Using Novel Feature Selection Methods,

    A. Salah, E. Shalabi, and W. Khedr, “A Lightweight Android Malware Classifier Using Novel Feature Selection Methods,” Symmetry, vol. 12, no. 5, p. 858, 2020

  19. [27]

    Android Malware Detection Using Genetic Algorithm Based Optimized Feature Selection and Machine Learning,

    A. Fatima, R. Maurya et al., “Android Malware Detection Using Genetic Algorithm Based Optimized Feature Selection and Machine Learning,” in TSP. IEEE, 2019, pp. 220–223

  20. [28]

    Malware Detection Using Deep Learning and Correlation-Based Feature Selection,

    E. S. Alomari, R. R. Nuiaa, Z. A. A. Alyasseri, H. J. Mohammed et al., “Malware Detection Using Deep Learning and Correlation-Based Feature Selection,” Symmetry, vol. 15, no. 1, p. 123, 2023

  21. [29]

    Opportunities and Challenges of Feature Selection Methods for High Dimensional Data: A Review,

    S. S. Subbiah and J. Chinnappan, “Opportunities and Challenges of Feature Selection Methods for High Dimensional Data: A Review,” Ing´enierie des Syst `emes d’Information, vol. 26, no. 1, 2021

  22. [30]

    A Survey on Intrusion Detection System: Feature Selection, Model, Performance Measures, Application Perspec- tive, Challenges, and Future Research Directions,

    A. Thakkar and R. Lohiya, “A Survey on Intrusion Detection System: Feature Selection, Model, Performance Measures, Application Perspec- tive, Challenges, and Future Research Directions,” Artificial Intelligence Review, vol. 55, no. 1, pp. 453–563, 2022

  23. [31]

    Uma An ´alise de M ´etodos de Selec ¸˜ao de Caracter ´ısticas Aplicados `a Detecc ¸˜ao de Malwares Android,

    T. Soares, D. Kreutz, V . Rocha, E. Costa, L. Le ˜ao, J. Pontes, J. Assolin, G. Rodrigues, and E. Feitosa, “Uma An ´alise de M ´etodos de Selec ¸˜ao de Caracter ´ısticas Aplicados `a Detecc ¸˜ao de Malwares Android,” in Anais do XXII SSBSeg , 2022, pp. 288–301. [Online]. Avail...

  24. [32]

    FS3E: Uma Ferramenta para Execuc ¸ ˜ao e Avaliac ¸˜ao de M ´etodos de Selec ¸ ˜ao de Caracter ´ısticas para Detecc ¸ ˜ao de Malwares Android,

    E. Costa, D. Kreutz, V . Rocha, L. Le ˜ao, S. Sab ´oia, N. Neves, and E. Feitosa, “FS3E: Uma Ferramenta para Execuc ¸ ˜ao e Avaliac ¸˜ao de M ´etodos de Selec ¸ ˜ao de Caracter ´ısticas para Detecc ¸ ˜ao de Malwares Android,” in Anais Estendidos do XXII SBSeg , 2022, pp. 151–1...

  25. [33]

    Avaliac ¸˜ao de M ´etodos de Selec ¸˜ao de Caracter ´ısticas de Amostras Android com a Ferramenta FS3E (v2),

    N. Neves, V . Rocha, D. Kreutz et al. , “Avaliac ¸˜ao de M ´etodos de Selec ¸˜ao de Caracter ´ısticas de Amostras Android com a Ferramenta FS3E (v2),” in Anais da XX ERRC , 2023, pp. 139–144. [Online]. Available: https://sol.sbc.org.br/index.php/errc/article/view/26019

  26. [34]

    MH-FSF: um Framework para Reproduc ¸˜ao, Experimentac ¸˜ao e Avaliac ¸˜ao de M ´etodos de Selec ¸˜ao de Caracter ´ısticas,

    V . Rocha, H. Braganc ¸a, D. Kreutz, and E. Feitosa, “MH-FSF: um Framework para Reproduc ¸˜ao, Experimentac ¸˜ao e Avaliac ¸˜ao de M ´etodos de Selec ¸˜ao de Caracter ´ısticas,” 2024, https://github.com/SBSegSF24/ MH-FSF

  27. [35]

    A Comprehen- sive Survey: Artificial Bee Colony (ABC) Algorithm and Applications,

    D. Karaboga, B. Gorkemli, C. Ozturk, and N. Karaboga, “A Comprehen- sive Survey: Artificial Bee Colony (ABC) Algorithm and Applications,” AI Review, vol. 42, pp. 21–57, 2014

  28. [36]

    Analysis of Variance (ANOV A),

    L. Sthle and S. Wold, “Analysis of Variance (ANOV A),” Chemometr Intell Lab., vol. 6, no. 4, pp. 259–272, 1989

  29. [37]

    Chi- Square Test,

    R. J. Tallarida, R. B. Murray, R. J. Tallarida, and R. B. Murray, “Chi- Square Test,” Manual of Pharmacologic Calculations with Computer Programs, pp. 140–142, 1987

  30. [38]

    Feature Selection Based on Information Gain,

    B. Azhagusundari, A. S. Thanamani et al., “Feature Selection Based on Information Gain,” IJITEE, vol. 2, no. 2, pp. 18–21, 2013

  31. [39]

    LASSO Regression,

    J. Ranstam and J. A. Cook, “LASSO Regression,” Journal of British Surgery, vol. 105, 2018

  32. [40]

    Mean-Absolute Deviation Model,

    H. Konno and T. Koshizuka, “Mean-Absolute Deviation Model,” IIE Transactions, vol. 37, no. 10, pp. 893–900, 2005

  33. [41]

    Principal Component Analysis (PCA),

    T. Kurita, “Principal Component Analysis (PCA),” Computer Vision: A Reference Guide, pp. 1–4, 2019

  34. [42]

    Pearson Correlation Coefficient,

    I. Cohen, Y . Huang, J. Chen, J. Benesty, J. Benesty, J. Chen, Y . Huang, and I. Cohen, “Pearson Correlation Coefficient,” Noise Reduction in Speech Processing, pp. 1–4, 2009

  35. [43]

    Theoretical and Empirical Analysis of ReliefF and RReliefF,

    M. Robnik- ˇSikonja and I. Kononenko, “Theoretical and Empirical Analysis of ReliefF and RReliefF,” Machine Learning, vol. 53, pp. 23– 69, 2003

  36. [44]

    Using Recursive Feature Elimination in Random Forest to Account for Correlated Variables in High Dimensional Data,

    B. F. Darst et al. , “Using Recursive Feature Elimination in Random Forest to Account for Correlated Variables in High Dimensional Data,” BMC Genetics, vol. 19, pp. 1–6, 2018

  37. [45]

    Scikit-learn: Machine Learning in Python,

    F. Pedregosa, G. Varoquaux, A. Gramfort, V . Michelet al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011

  38. [46]

    The MCC-F1 Curve: A Performance Evaluation Technique for Binary Classification,

    C. Cao, D. Chicco, and M. M. Hoffman, “The MCC-F1 Curve: A Performance Evaluation Technique for Binary Classification,” arXiv preprint arXiv:2006.11278, 2020

  39. [47]

    James, D

    G. James, D. Witten, T. Hastie, and R. Tibshirani, An Introduction to Statistical Learning: with Applications in R . Springer, 2013. [Online]. Available: https://faculty.marshall.usc.edu/gareth-james/ISL/

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.