Pith. sign in

REVIEW 19 cited by

Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.16061 v1 pith:ESXXDISC submitted 2020-10-11 cs.LG stat.MEstat.ML

classification cs.LGstat.MEstat.ML
keywords informednessmeasurescasechancemarkednessprecisionrecallused
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Commonly used evaluation measures including Recall, Precision, F-Measure and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the statistic. Using these measures a system that performs worse in the objective sense of Informedness, can appear to perform better under any of these commonly used measures. We discuss several concepts and measures that reflect the probability that prediction is informed versus chance. Informedness and introduce Markedness as a dual measure for the probability that prediction is marked versus chance. Finally we demonstrate elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision, and outline the extension from the dichotomous case to the general multi-class case.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,564 citations worldwide. Full citation record

  1. PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data

    cs.HC 2026-07 conditional novelty 6.5 of 10

    Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.

  2. Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

    eess.IV 2026-07 conditional novelty 6.0 of 10

    Chest X-ray AI model rankings and image-quality metric rankings change substantially with the choice of evaluation reference, so benchmark scores are not neutral.

  3. SWDL: Stratum-Wise Difference Learning with Deep Laplacian Pyramid for Semi-Supervised 3D Intracranial Hemorrhage Segmentation

    eess.IV 2025-06 conditional novelty 6.0 of 10

    SWDL-Net improves semi-supervised intracranial hemorrhage segmentation by learning from differences between a Laplacian pyramid upsampler and a convolutional upsampler, reaching 89.3% Dice with 2% labeled data.

  4. GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A multi-agent vision-language framework that extracts event, time, and location from public event images, evaluated with a new soft metric on VLM-augmented datasets.

  5. Exoplanet Transit Candidate Identification in TESS Full-Frame Images via a Transformer-Based Algorithm

    astro-ph.EP 2025-02 conditional novelty 6.0 of 10

    A Transformer-based network detects transit-like dips in TESS FFI light curves without phase folding, yielding 214 new exoplanet candidates including single-transit and multi-planet systems.

  6. Event Detection in Videos: A Framework for the Development of New Methods

    cs.CV 2026-07 conditional novelty 5.5 of 10

    A framework of tagged multi-environment datasets (including new FSD and SUC), probabilistic Tile-based ranking, and explicit application scenarios for fair video event detection.

  7. From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Phenotype-first unsupervised subgroups yield comparable held-out policy utilities across clustering methods, with no statistically significant differences, while the individuals prioritized differ substantially.

  8. From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    LLM-scored offer predicates aggregated by a Logic Tensor Network classify procurement documents about as accurately as BERT or LLM baselines while exposing auditable predicate and rule truth values.

  9. MVRS: The Multimodal Virtual Reality Stimuli-based Emotion Recognition Dataset

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A new VR-based emotion dataset with synchronized eye tracking, body motion, EMG, and GSR from 13 participants, evaluated with classifiers but with questionable validation.

  10. CS-Agent: LLM-based Community Search via Dual-agent Collaboration

    cs.SI 2025-08 conditional novelty 5.0 of 10

    CS-Agent, a Solver-Validator two-agent dialogue with a Decider selector, improves LLM community search on synthetic graphs, and GraphCS is a new benchmark for measuring it.

  11. Physical Layer Authentication Based on Hierarchical Variational Auto-Encoder for Industrial Internet of Things

    eess.SP 2025-08 conditional novelty 5.0 of 10

    A hierarchical autoencoder-plus-variational-autoencoder scheme authenticates industrial IoT transmitters from channel impulse responses, claiming higher F1 than three baselines without attacker channel priors.

  12. Benchmarking Unsupervised Strategies for Anomaly Detection in Multivariate Time Series

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Across ten public datasets, a reconstruction-based inverted transformer with per-variate anomaly labelling achieves the best or tied best MCC on most datasets, but the comparison is weakened by test-set-based configur...

  13. A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment

    cs.CV 2026-06 unverdicted novelty 4.5 of 10

    End-Net, a multiscale inception-based CNN, reaches 0.9761 accuracy on a balanced multi-class MRI dataset of Alzheimer, tumors, MS and controls and is deployed as a public web service.

  14. A Novel Data Augmentation Strategy for Robust Deep Learning Classification of Biomedical Time-Series Data: Application to ECG and EEG Analysis

    eess.SP 2025-07 reject novelty 4.0 of 10

    A proposed ResNet plus attention model with concatenated time-domain augmentations reports 99.96%, 99.78%, and 100% accuracy on UCI EEG, MIT-BIH ECG, and PTB ECG, but no ablation or split protocol supports the state-o...

  15. IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A visual heatmap tool lets people rate vision-language model reliability in video by inspecting patterns of green and red cells, with user ratings tracking objective F1 scores when those exist.

  16. FoundationalECGNet: A Lightweight Foundational Model for ECG-based Multitask Cardiac Analysis

    cs.LG 2025-09 reject novelty 3.0 of 10

    A multi-architecture ECG classifier reports near-perfect scores on a small test set, but the evaluation is compromised by pre-split oversampling and inconsistent metric reporting.

  17. A Novel Convolutional Neural Network-Based Framework for Complex Multiclass Brassica Seed Classification

    cs.CV 2025-05 reject novelty 3.0 of 10

    A custom 23-layer CNN classifies ten Brassica seed types from microscope images with 93% test accuracy on a newly collected dataset.

  18. A Framework for Multi-View Multiple Object Tracking using Single-View Multi-Object Trackers on Fish Data

    cs.CV 2025-05 reject novelty 3.0 of 10

    A YOLOv8-ByteTrack pipeline plus stereo triangulation can produce 3D fish tracks for some underwater video pairs, but the claimed multi-view accuracy improvement is not demonstrated.

  19. Absolute Evaluation Measures for Machine Learning: A Survey

    cs.LG 2025-07 unverdicted novelty 1.0 of 10

    A survey compiles bounded absolute evaluation metrics for classification, clustering, and ranking and proposes decision trees for metric selection, but several formulas are reproduced incorrectly.

Pith tools