REVIEW 5 cited by
An Approach to Multiple Comparison Benchmark Evaluations that is Stable Under Manipulation of the Comparate Set
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The measurement of progress using benchmarks evaluations is ubiquitous in computer science and machine learning. However, common approaches to analyzing and presenting the results of benchmark comparisons of multiple algorithms over multiple datasets, such as the critical difference diagram introduced by Dem\v{s}ar (2006), have important shortcomings and, we show, are open to both inadvertent and intentional manipulation. To address these issues, we propose a new approach to presenting the results of benchmark comparisons, the Multiple Comparison Matrix (MCM), that prioritizes pairwise comparisons and precludes the means of manipulating experimental results in existing approaches. MCM can be used to show the results of an all-pairs comparison, or to show the results of a comparison between one or more selected algorithms and the state of the art. MCM is implemented in Python and is publicly available.
Forward citations
Cited by 5 Pith papers
-
Soft-MSM: Differentiable Context-Aware Elastic Alignment for Time Series
Soft-MSM is a smooth, gradient-enabled version of the context-aware MSM distance for time series alignment that outperforms Soft-DTW alternatives in clustering and nearest-centroid classification.
-
Statistical comparisons of time-series feature sets on classification tasks
Across 124 time-series classification datasets, six open-source feature sets perform mostly equivalently, with tsfresh winning most often and simple quantile/FFT baselines competitive on several problems.
-
Scaling Time Series Classification via XAI-Driven Data Reduction
drXAI uses XAI attributions from a fast classifier to choose important channels/time points, achieving 80–90% data reduction with comparable classification accuracy.
-
A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning
Rehab-Pile, a standardized archive of 60 rehabilitation skeleton datasets with reproducible baselines, reports that the authors' lightweight LITEMV model wins on efficiency and usually on accuracy.
-
Time series classification with random convolution kernels: pooling operators and input representations matter
SelF-Rocket dynamically selects input representations and pooling operators within random convolution kernel methods for TSC and reports SOTA accuracy on UCR datasets.
Discussion (0). Sign in to comment.