Pith. sign in

REVIEW 1 cited by

vocadito: A dataset of solo vocals with $f_0$, note, and lyric annotations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.05580 v2 pith:45PGEEJT submitted 2021-10-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords noteannotationsvocaditoalgorithmsdatasetdifferentestimationsinging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

To compliment the existing set of datasets, we present a small dataset entitled vocadito, consisting of 40 short excerpts of monophonic singing, sung in 7 different languages by singers with varying of levels of training, and recorded on a variety of devices. We provide several types of annotations, including $f_0$, lyrics, and two different note annotations. All annotations were created by musicians. We provide an analysis of the differences between the two note annotations, and see that the agreement level is low, which has implications for evaluating vocal note estimation algorithms. We also analyze the relation between the $f_0$ and note annotations, and show that quantizing $f_0$ values in frequency does not provide a reasonable note estimate, reinforcing the difficulty of the note estimation task for singing voice. Finally, we provide baseline results from recent algorithms on vocadito for note and $f_0$ transcription. Vocadito is made freely available for public use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwiftF0: Fast and Accurate Monophonic Pitch Detection

    cs.SD 2025-08 conditional novelty 6.0 of 10

    SwiftF0 estimates monophonic pitch from a compact STFT-CNN, reporting better accuracy than CREPE under 10 dB noise at 42x lower CPU cost, alongside a new synthetic speech dataset and a six-component evaluation metric.

Pith tools