REVIEW 19 cited by
Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Commonly used evaluation measures including Recall, Precision, F-Measure and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the statistic. Using these measures a system that performs worse in the objective sense of Informedness, can appear to perform better under any of these commonly used measures. We discuss several concepts and measures that reflect the probability that prediction is informed versus chance. Informedness and introduce Markedness as a dual measure for the probability that prediction is marked versus chance. Finally we demonstrate elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision, and outline the extension from the dichotomous case to the general multi-class case.
Forward citations
Cited by 19 Pith papers
-
PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data
Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.
-
Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance
Chest X-ray AI model rankings and image-quality metric rankings change substantially with the choice of evaluation reference, so benchmark scores are not neutral.
-
SWDL: Stratum-Wise Difference Learning with Deep Laplacian Pyramid for Semi-Supervised 3D Intracranial Hemorrhage Segmentation
SWDL-Net improves semi-supervised intracranial hemorrhage segmentation by learning from differences between a Laplacian pyramid upsampler and a convolutional upsampler, reaching 89.3% Dice with 2% labeled data.
-
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
A multi-agent vision-language framework that extracts event, time, and location from public event images, evaluated with a new soft metric on VLM-augmented datasets.
-
Exoplanet Transit Candidate Identification in TESS Full-Frame Images via a Transformer-Based Algorithm
A Transformer-based network detects transit-like dips in TESS FFI light curves without phase folding, yielding 214 new exoplanet candidates including single-transit and multi-planet systems.
-
Event Detection in Videos: A Framework for the Development of New Methods
A framework of tagged multi-environment datasets (including new FSD and SUC), probabilistic Tile-based ranking, and explicit application scenarios for fair video event detection.
-
From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data
Phenotype-first unsupervised subgroups yield comparable held-out policy utilities across clustering methods, with no statistically significant differences, while the individuals prioritized differ substantially.
-
From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement
LLM-scored offer predicates aggregated by a Logic Tensor Network classify procurement documents about as accurately as BERT or LLM baselines while exposing auditable predicate and rule truth values.
-
MVRS: The Multimodal Virtual Reality Stimuli-based Emotion Recognition Dataset
A new VR-based emotion dataset with synchronized eye tracking, body motion, EMG, and GSR from 13 participants, evaluated with classifiers but with questionable validation.
-
CS-Agent: LLM-based Community Search via Dual-agent Collaboration
CS-Agent, a Solver-Validator two-agent dialogue with a Decider selector, improves LLM community search on synthetic graphs, and GraphCS is a new benchmark for measuring it.
-
Physical Layer Authentication Based on Hierarchical Variational Auto-Encoder for Industrial Internet of Things
A hierarchical autoencoder-plus-variational-autoencoder scheme authenticates industrial IoT transmitters from channel impulse responses, claiming higher F1 than three baselines without attacker channel priors.
-
Benchmarking Unsupervised Strategies for Anomaly Detection in Multivariate Time Series
Across ten public datasets, a reconstruction-based inverted transformer with per-variate anomaly labelling achieves the best or tied best MCC on most datasets, but the comparison is weakened by test-set-based configur...
-
A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment
End-Net, a multiscale inception-based CNN, reaches 0.9761 accuracy on a balanced multi-class MRI dataset of Alzheimer, tumors, MS and controls and is deployed as a public web service.
-
A Novel Data Augmentation Strategy for Robust Deep Learning Classification of Biomedical Time-Series Data: Application to ECG and EEG Analysis
A proposed ResNet plus attention model with concatenated time-domain augmentations reports 99.96%, 99.78%, and 100% accuracy on UCI EEG, MIT-BIH ECG, and PTB ECG, but no ablation or split protocol supports the state-o...
-
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
A visual heatmap tool lets people rate vision-language model reliability in video by inspecting patterns of green and red cells, with user ratings tracking objective F1 scores when those exist.
-
FoundationalECGNet: A Lightweight Foundational Model for ECG-based Multitask Cardiac Analysis
A multi-architecture ECG classifier reports near-perfect scores on a small test set, but the evaluation is compromised by pre-split oversampling and inconsistent metric reporting.
-
A Novel Convolutional Neural Network-Based Framework for Complex Multiclass Brassica Seed Classification
A custom 23-layer CNN classifies ten Brassica seed types from microscope images with 93% test accuracy on a newly collected dataset.
-
A Framework for Multi-View Multiple Object Tracking using Single-View Multi-Object Trackers on Fish Data
A YOLOv8-ByteTrack pipeline plus stereo triangulation can produce 3D fish tracks for some underwater video pairs, but the claimed multi-view accuracy improvement is not demonstrated.
-
Absolute Evaluation Measures for Machine Learning: A Survey
A survey compiles bounded absolute evaluation metrics for classification, clustering, and ranking and proposes decision trees for metric selection, but several formulas are reproduced incorrectly.
Discussion (0). Continue with ORCID to comment.