Pith. sign in

REVIEW 2 cited by

Filter Methods for Feature Selection in Supervised Machine Learning Applications -- Review and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.12140 v1 pith:RFYMFM6U submitted 2021-11-23 cs.LG cs.DBstat.ML

classification cs.LGcs.DBstat.ML
keywords methodsfeaturefeaturesfilterapplicationsnumberperformanceselection
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The amount of data for machine learning (ML) applications is constantly growing. Not only the number of observations, especially the number of measured variables (features) increases with ongoing digitization. Selecting the most appropriate features for predictive modeling is an important lever for the success of ML applications in business and research. Feature selection methods (FSM) that are independent of a certain ML algorithm - so-called filter methods - have been numerously suggested, but little guidance for researchers and quantitative modelers exists to choose appropriate approaches for typical ML problems. This review synthesizes the substantial literature on feature selection benchmarking and evaluates the performance of 58 methods in the widely used R environment. For concrete guidance, we consider four typical dataset scenarios that are challenging for ML models (noisy, redundant, imbalanced data and cases with more features than observations). Drawing on the experience of earlier benchmarks, which have considered much fewer FSMs, we compare the performance of the methods according to four criteria (predictive performance, number of relevant features selected, stability of the feature sets and runtime). We found methods relying on the random forest approach, the double input symmetrical relevance filter (DISR) and the joint impurity filter (JIM) were well-performing candidate methods for the given dataset scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feature Shift Localization Network

    cs.LG 2025-06 conditional novelty 6.0 of 10

    FSL-Net localizes shifted features between two datasets using a network trained on 1,350 datasets, matching DataFix's F1 while being about 36x faster on average.

  2. ROOFS: RObust biOmarker Feature Selection

    stat.ML 2026-01 conditional novelty 5.0 of 10

    A benchmarking package and a PIONeeR case study conclude that a BH-adjusted p-value filter beats LASSO on stability, predictive AUC, and discovery rate for lung-cancer immunotherapy-resistance biomarkers.

Pith tools