REVIEW 11 cited by
Metrics for Multi-Class Classification: an Overview
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of a multi-class classifier. Those metrics turn out to be useful at different stage of the development process, e.g. comparing the performance of two different models or analysing the behaviour of the same model by tuning different parameters. In this white paper we review a list of the most promising multi-class metrics, we highlight their advantages and disadvantages and show their possible usages during the development of a classification model.
Forward citations
Cited by 11 Pith papers
-
TransformEEG: Towards Improving Model Generalizability in Deep Learning-based EEG Parkinson's Disease Detection
TransformEEG, a convolutional-transformer with a depthwise tokenizer, reports the highest median balanced accuracy (78.4-80.1%) and lowest variability across splits among eight EEG models for Parkinson's detection on ...
-
Rhythm Features for Speaker Identification
Using character-duration sequences from WhisperX, a transformer identifies speakers with balanced accuracy of 0.39 on LibriSpeech but only 0.03 on VoxCeleb1, and fusion with x-vectors does not improve accuracy.
-
Exploring LLM Capabilities in Extracting DCAT-Compatible Metadata for Data Cataloging
Large language models can generate DCAT-compatible metadata for data catalogs at quality close to human annotations, though the strongest evidence is for simple extraction tasks.
-
Comparing Credit Risk Estimates in the Gen-AI Era
Few-shot GPT-4o underperforms logistic regression and KNN on German Credit Data across all tested prompt and example-selection configurations.
-
Runtime Analysis of Evolutionary NAS for Multiclass Classification
On a hand-built multiclass benchmark, (1+1)-ENAS with one-bit or bit-wise mutation finds an optimal architecture in O(rM ln(rM)) expected generations, with lower bound Omega(rM ln M), so the two mutations have nearly ...
-
Rule-Based Modeling of Low-Dimensional Data with PCA and Binary Particle Swarm Optimization (BPSO) in ANFIS
ANFIS-PCA-BPSO applies PCA to normalized firing strengths and selects components with BPSO, reducing ANFIS rule count and training time at a small accuracy cost.
-
Towards Continuous-variable Quantum Neural Networks for Biomedical Imaging
A 4-qumode Gaussian CV-QNN classifies MedMNIST images with accuracy statistically indistinguishable from a 42-parameter classical linear model and a DV-QNN.
-
Proactive HIV Care: AI-Based Comorbidity Prediction from Routine EHR Data
Across six models and 2,200 HIV outpatients, including demographic features always improved multi-label comorbidity prediction, with XGBoost best at 45.8% macro F1; gender was recoverable from labs at 92.8%.
-
Ensemble BERT for Medication Event Classification on Electronic Health Records (EHRs)
Majority voting over multiple pretrained BERT models improves medication event classification on the n2c2 CMED dataset, but the paper lacks error bars and code.
-
Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models
Fine-tuned VGG16, VGG19 and ResNet-50v2 classify a public Kaggle chest X-ray set (COVID-19, normal, viral pneumonia) at 85–96% accuracy; the abstract claims tuberculosis and lung-cancer coverage the experiments never include.
-
Multidimensional classification of posts for online course discussion forum curation
Bayesian fusion of a generic LLM and a local classifier ties the best individual classifier on MOOC forum labels and lags fine-tuning, undermining the paper's headline claim.
Discussion (0). Continue with ORCID to comment.