REVIEW 7 cited by
Macro F1 and Macro F1
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The 'macro F1' metric is frequently used to evaluate binary, multi-class and multi-label classification problems. Yet, we find that there exist two different formulas to calculate this quantity. In this note, we show that only under rare circumstances the two computations can be considered equivalent. More specifically, one formula well 'rewards' classifiers which produce a skewed error type distribution. In fact, the difference in outcome of the two computations can be as high as 0.5. The two computations may not only diverge in their scalar result but can also lead to different classifier rankings.
Forward citations
Cited by 7 Pith papers
-
Soprano voices in opera seria: a corpus-based inquiry into eighteenth-century vocal types
Composers of eighteenth-century opera seria encoded character gender in vocal writing but barely tailored notation to the singer's biological sex.
-
Dependency Triad: A Metric to Quantify the Dependencies Between Attributes for Local Differential Privacy
The Dependency Triad summarizes pairwise attribute dependence with three parameters and delivers a constant-time upper-bound estimate of correlation-induced privacy leakage.
-
Joint Modeling of Entities and Discourse Relations for Coherence Assessment
Jointly modeling entities and discourse relations improves coherence assessment accuracy over text-only and single-feature models on GCDC, CoheSentia, and TOEFL.
-
Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning
Meta-PerSER uses MAML-style meta-training with combined-set training, derivative annealing, and per-layer learning rates to personalize speech emotion recognition to unseen annotators from 32 labeled examples, outperf...
-
Gaze-Enhanced Multimodal Turn-Taking Prediction in Triadic Conversations
A gaze-enhanced multimodal model predicts turn-taking in triadic conversations better than voice activity alone, and multi-user gaze gives the largest gains.
-
Timestamp calibration for time-series single cell RNA-seq expression data
ScPace, a self-paced SVM that drops high-loss cells before retraining, improves timestamp annotation and supervised pseudotime analysis on simulated and real noisy time-series single-cell RNA-seq data.
-
Financial Fine-tuning a Large Time Series Model
Fine-tuning TimesFM on log-transformed financial price data improves directional accuracy and mock-trading Sharpe ratios over the base model.
Discussion (0). Continue with ORCID to comment.