Pith. sign in

REVIEW 4 cited by

RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.15412 v2 pith:3UMXOKQB submitted 2023-06-27 cs.SD eess.AS

classification cs.SDeess.AS
keywords musicvocalrmvpepitchpolyphonicmodelpitchesrobust
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vocal pitch is an important high-level feature in music audio processing. However, extracting vocal pitch in polyphonic music is more challenging due to the presence of accompaniment. To eliminate the influence of the accompaniment, most previous methods adopt music source separation models to obtain clean vocals from polyphonic music before predicting vocal pitches. As a result, the performance of vocal pitch estimation is affected by the music source separation models. To address this issue and directly extract vocal pitches from polyphonic music, we propose a robust model named RMVPE. This model can extract effective hidden features and accurately predict vocal pitches from polyphonic music. The experimental results demonstrate the superiority of RMVPE in terms of raw pitch accuracy (RPA) and raw chroma accuracy (RCA). Additionally, experiments conducted with different types of noise show that RMVPE is robust across all signal-to-noise ratio (SNR) levels. The code of RMVPE is available at https://github.com/Dream-High/RMVPE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Adding SVS-style phoneme conditioning and a length regulator to melody-guided cover generation cuts phoneme error rate from 45.6% to 18.7% while holding melody metrics.

  2. STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

    cs.SD 2025-07 conditional novelty 6.0 of 10

    STARS unifies lyric alignment, note transcription, vocal technique detection, and global style prediction into one multi-level neural model that matches or beats several single-task baselines.

  3. Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

    cs.SD 2026-07 conditional novelty 5.0 of 10

    A unified hierarchical-LM-plus-flow-matching system reports top-tier full-song vocal generation, ranking 2–3 on an external blind leaderboard, but releases neither code nor evaluation data.

  4. RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding

    eess.AS 2025-06 conditional novelty 5.0 of 10

    RT-VC converts a speaker's voice to a new target voice in real time on a CPU with 61.4ms latency, matching the quality of the current SOTA StreamVC.

Pith tools