Pith. sign in

REVIEW 6 cited by

Macro F1 and Macro F1

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03347 v3 pith:VJEL7DZZ submitted 2019-11-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords computationsmacrodifferentonlybinarycalculatecircumstancesclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The 'macro F1' metric is frequently used to evaluate binary, multi-class and multi-label classification problems. Yet, we find that there exist two different formulas to calculate this quantity. In this note, we show that only under rare circumstances the two computations can be considered equivalent. More specifically, one formula well 'rewards' classifiers which produce a skewed error type distribution. In fact, the difference in outcome of the two computations can be as high as 0.5. The two computations may not only diverge in their scalar result but can also lead to different classifier rankings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. Soprano voices in opera seria: a corpus-based inquiry into eighteenth-century vocal types

    stat.AP 2026-08 conditional novelty 6.0 of 10

    Composers of eighteenth-century opera seria encoded character gender in vocal writing but barely tailored notation to the singer's biological sex.

  2. Dependency Triad: A Metric to Quantify the Dependencies Between Attributes for Local Differential Privacy

    cs.CR 2026-08 conditional novelty 6.0 of 10

    The Dependency Triad summarizes pairwise attribute dependence with three parameters and delivers a constant-time upper-bound estimate of correlation-induced privacy leakage.

  3. Joint Modeling of Entities and Discourse Relations for Coherence Assessment

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Jointly modeling entities and discourse relations improves coherence assessment accuracy over text-only and single-feature models on GCDC, CoheSentia, and TOEFL.

  4. Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Meta-PerSER uses MAML-style meta-training with combined-set training, derivative annealing, and per-layer learning rates to personalize speech emotion recognition to unseen annotators from 32 labeled examples, outperf...

  5. Timestamp calibration for time-series single cell RNA-seq expression data

    q-bio.GN 2024-12 conditional novelty 5.0 of 10

    ScPace, a self-paced SVM that drops high-loss cells before retraining, improves timestamp annotation and supervised pseudotime analysis on simulated and real noisy time-series single-cell RNA-seq data.

  6. Financial Fine-tuning a Large Time Series Model

    q-fin.CP 2024-12 conditional novelty 4.0 of 10

    Fine-tuning TimesFM on log-transformed financial price data improves directional accuracy and mock-trading Sharpe ratios over the base model.

Pith tools