Pith. sign in

REVIEW 3 cited by

Neural Speech Extraction with Human Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.03041 v1 pith:MYNXKMM3 submitted 2025-08-05 cs.SD cs.LGeess.AS

Neural Speech Extraction with Human Feedback

classification cs.SD cs.LGeess.AS
keywords extractionhumanneuralrefinementspeechapproachdatasetsfeedback
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then improves the marked sections while preserving unmarked regions. Since large-scale datasets of human-marked errors are difficult to collect, we generate synthetic datasets using various automated masking functions and train models on each. Evaluations show that models trained with noise power-based masking (in dBFS) and probabilistic thresholding perform best, aligning with human annotations. In a study with 22 participants, users showed a preference for refined outputs over baseline TSE. Our findings demonstrate that human-in-the-loop refinement is a promising approach for improving the performance of neural speech extraction.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Unmixing The Crowd: Learning Persistent Speaker Representations from Mixture-Derived Multi-Speaker Embeddings

    eess.AS 2026-04 unverdicted novelty 7.0

    A neural model predicts a set of speaker embeddings from noisy mixtures to enable enrollment-free target speech extraction, outperforming baselines on LibriMix and generalizing to real recordings.

  2. Unmixing The Crowd: Learning Persistent Speaker Representations from Mixture-Derived Multi-Speaker Embeddings

    eess.AS 2026-04 conditional novelty 5.0

    A LoRA-generating hypernetwork unifies conventional pulmonary and opportunistic cardiac screening from non-contrast chest CT, matching single-task models while outperforming multi-task baselines on prospective multi-c...

  3. GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

    eess.AS 2025-12 conditional novelty 5.0

    A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.