Pith. sign in

Paper Citation Record · LEDGER

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

As of 20 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.19314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19314 v3

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:07.475291Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:20:49.132229Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact7
  • verified fuzzy55
  • unresolved17
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d86250cd-c963-4ac1-9256-ddbe38fad0ee · outbound

This paper cites The cocktail-party problem revisited: early pro- cessing and selection of multi-talker speech,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline The cocktail-party problem revisited: early pro- cessing and selection of multi-talker speech,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.316621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.154451Z digest=sha256:534fc8b80fd01cc866e629c328ca98a9637e8e6e1a5e6d19e8a3b671fe019256

Observation d24a1ce8-4e6c-4d93-9999-a8c2e4cf756a · outbound

This paper cites Neural target speech extraction: An overview,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Neural target speech extraction: An overview,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.167867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.158760Z digest=sha256:e1843b321e9ec1f334af30b2a417d431d01820acb9dd9a92caadb625920f7be2

Observation a5f4c12f-df00-4fbe-9813-b06d63e71704 · outbound

This paper cites Neural spatial filter: Target speaker speech separation assisted with directional information,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Neural spatial filter: Target speaker speech separation assisted with directional information,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:16.021303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.163056Z digest=sha256:02cf52a5fd2019c81292c2584da2ed7b20bba3aee17fbe4557a2556d5108d634

Observation 8b003eec-47d2-42bc-b634-cecfc738d80d · outbound

This paper cites Far-field location guided target speech extraction using end-to-end speech recognition objectives,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Far-field location guided target speech extraction using end-to-end speech recognition objectives,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.913298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.168216Z digest=sha256:b318c88be47eb8fcad8e7b14ec446c4712fc19b4f35b0527fc9caf755ffadabe

Observation 1eedf57c-0046-4639-a424-985e93779fbf · outbound

This paper cites Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.435237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.171996Z digest=sha256:07fc225b35ecb6640ab75d63c8b64eaf7067860fa2e71d3f3ab2729936bef706

Observation 27460b18-8024-4c28-a89d-4479dbe0a0ef · outbound

This paper cites Conceptbeam: Concept driven target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Conceptbeam: Concept driven target speech extraction,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.283829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.176146Z digest=sha256:41d2328e679a337ac6f9dae06a4dd45eed3a2752bf70b96f337d272310ad38f7

Observation 693bb2b9-cf52-4e47-8d02-945d341d136c · outbound

This paper cites V oicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline V oicefilter: Targeted voice separation by speaker-conditioned spectrogram masking,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.164815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.180844Z digest=sha256:a57fcf91aeee39432b06fd15f3abaca3e1c9402aefb0cdc99f345f58df72cadf

Observation 930acadf-25cf-4460-b1bf-70371d4bf6ca · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:15.002077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.184966Z digest=sha256:0c322987cb5b0afd1b43111d37c313bd308a8ff9ff822b88a007031ab18f934a

Observation 33604da3-1634-4163-ac55-ae8543afb007 · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.875258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.188682Z digest=sha256:69d2789e3c4c9101c8156efea52b3bb61fa5a338e1e642751840984815195e1c

Observation 460ed315-7fe9-41ff-b469-ddcbc96234dd · outbound

This paper cites Target confusion in end-to-end speaker extraction: Analysis and approaches,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target confusion in end-to-end speaker extraction: Analysis and approaches,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.758242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.192850Z digest=sha256:ad3aaeb0c3db5fbb3a2adac5e1c847f4473363aa0dac90f74617f2c3ca4ca9ee

Observation f22eb8bc-4f60-4f6c-803d-194c16e497aa · outbound

This paper cites Dpccn: Densely-connected pyramid complex convolutional network for robust speech separation 11 and extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dpccn: Densely-connected pyramid complex convolutional network for robust speech separation 11 and extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.633307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.196434Z digest=sha256:167972399d84e88236f1646463831af3ac7d5d64239c349645c5d57f5e6768f4

Observation a5f74896-a48a-44e2-a4d5-4a405cdc5bdc · outbound

This paper cites Improving target sound extraction with timestamp information,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving target sound extraction with timestamp information,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.516695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.200583Z digest=sha256:060e0e60844f13b79e8eff921967004a43ea227a2488c94bb405ca4e1d52c29a

Observation 610f6f0c-164e-427f-ad8d-55bce01922a0 · outbound

This paper cites WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.204106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.204106Z digest=sha256:dff036e2efdfe1f524ebe3810cacb5b8116bf703a42e6d2089e165cac4974e7d

Observation 611ca561-de03-4463-8427-7df5c5ba0e9c · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Spex: Multi-scale time domain speaker extraction network,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.357625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.208667Z digest=sha256:8771d52fd56ff5a0796609670ee3b7d6d5c78fb64c6f39265c978a7231961f60

Observation 3858436a-b8b9-43ec-9a1f-78e582e3d462 · outbound

This paper cites Spex+: A complete time domain speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Spex+: A complete time domain speaker extraction network,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.233683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.212662Z digest=sha256:9ec6719d71ad554ba3a5db14514e4ab5e1c847916ab97525085f6531010f04bd

Observation 85db493b-472e-4c86-895d-e4730dc4985c · outbound

This paper cites X-SEPFORMER: end-to- end speaker extraction network with explicit optimization on speaker confusion,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X-SEPFORMER: end-to- end speaker extraction network with explicit optimization on speaker confusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:14.115190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.216422Z digest=sha256:a700328d06117fa9b6faa31cd41e40bd89af331ce7727e4824b960dfe5869a2a

Observation 8da9cead-1b4d-4c0c-877a-787dddb08cab · outbound

This paper cites X-tf-gridnet: A time-frequency domain target speaker extraction network with adaptive speaker embedding fusion,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X-tf-gridnet: A time-frequency domain target speaker extraction network with adaptive speaker embedding fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.979985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.221377Z digest=sha256:fc7b2fc308cefb3e8b288c63002be729a6a1ae2dec00b742dc93c5ac551a54b1

Observation 1de9266a-258a-4493-b63b-e2782b33b22b · outbound

This paper cites USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.834953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.225317Z digest=sha256:5f30039aba4f4f17b069adec54215b6ed2d125e193046e27ee5220f201c76326

Observation ed61ae57-6b30-475d-b20a-a71d244102f2 · outbound

This paper cites Target speech extraction with conditional diffusion model,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with conditional diffusion model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.866038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.230095Z digest=sha256:3451a79cb21fae7e8b9c1f108a0077e300b5ff5ca1af7149042c2ef018374594

Observation 3f2ee616-3578-4865-9a29-f62fd92bdbb1 · outbound

This paper cites Noise-robust Speech Separation with Fast Generative Correction.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Noise-robust Speech Separation with Fast Generative Correction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.817493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.234297Z digest=sha256:70bc4ce02363b413569048222994d64f8255eca8365b5f567dd74502ad6dfbe5

Observation 9f531854-3c29-43d1-bf74-867007a691f8 · outbound

This paper cites Speech enhancement and dereverberation with diffusion-based generative models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speech enhancement and dereverberation with diffusion-based generative models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.683036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.239647Z digest=sha256:d2e86d2373ca0eac326b84db8775208ba77c4778a02895aeb5054be598462fbf

Observation 1c4af11a-89ad-4505-9ea5-a24441951168 · outbound

This paper cites Diffusion-based generative speech source separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion-based generative speech source separation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.243716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.243716Z digest=sha256:9e871dccc334e65abf2df755c3188ca9839bab8a9f3339b440ee26a13e0d7ed9

Observation 63742595-d3a9-4838-bd75-3a4f27419613 · outbound

This paper cites Generative pre-training for speech with flow matching,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generative pre-training for speech with flow matching,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.475291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.248153Z digest=sha256:4704039f411f80416c35bfeb4388456fd6e0d81beb88627ac1acb6330035a744

Observation ffb056b2-5e5a-4357-a714-17287e0219e7 · outbound

This paper cites Metis: A Foundation Speech Generation Model with Masked Generative Pre-training.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.252044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.252044Z digest=sha256:7f8e21bbb6bdb5b07d9c991e68d80a15777cc9fb506cf68e5f46cf89110fc273

Observation af0f0304-7d6e-4afa-91ce-5577d88b2cbd · outbound

This paper cites SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.256774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.256774Z digest=sha256:ea3ea0caed2b51f711cd6ec9c9910ea54e158f3e75e1fc77a6559a63238176e4

Observation bf62dc7e-f961-46a4-be01-5f8568be862f · outbound

This paper cites Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.261084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.261084Z digest=sha256:847f17f483a09d6282d9d41a8499c46e4e629b90ea8fe26264538384662013e2

Observation eadbdd15-2a50-446f-88e2-1527b08990b5 · outbound

This paper cites Attention is all you need,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.265276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.265276Z digest=sha256:9aae17d9e365e469299ac5b398cd4adebee2dd09447e983971bfc9dcd88d833a

Observation 17e636f9-cabc-4df8-b41f-2cbca8dbe101 · outbound

This paper cites Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition,

Reference 28

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:22:07.780630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.269521Z digest=sha256:04587e0a445cccd8ea49356f2beeefd7137871df2edb9cfe09cc61754fd7c1ec

Observation 58f778a2-4cd6-4848-b207-4291a63e5542 · outbound

This paper cites SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.710817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.274041Z digest=sha256:1dc3ce9ec5293ee57808c6566f81169ec16db802f9545b065e0e13f9a3fa4cd0

Observation ff992483-89bd-4c28-8710-1079063a1ac4 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.278378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.278378Z digest=sha256:383eb5a294786b2c778edd3699fccf84c4cc5c481e0b467ab799704e8dd0a68e

Observation 335f9c72-14f9-4d0b-9d3d-7ba3e47b1cba · outbound

This paper cites Target speech extraction with conditional diffusion model,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with conditional diffusion model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.348423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.282398Z digest=sha256:643eca188cd5ce0f8975a791176ae19c1ec418bcec0e8e9d05d1e1dc51425ed3

Observation e99c4cc7-f709-4a9a-a5f6-b9cb4bf63e49 · outbound

This paper cites Dpm-tse: A diffusion probabilistic model for target sound extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dpm-tse: A diffusion probabilistic model for target sound extraction,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:13.098187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.286482Z digest=sha256:64c3997c31693b7f5eb679135a0a2e5608f68c6802fd8ea8113fb1e174e757b0

Observation 73ea1afc-061b-4f60-a9f9-019697bca4fb · outbound

This paper cites Diffusion- based generative speech source separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion- based generative speech source separation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.750074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.290826Z digest=sha256:4e71a0d4de1a18be4c63b7dc21bbf5ac252bfc43b8fefea68bc8cd7819d03357

Observation 246e030c-ca12-4386-8f20-fdd375bd3a4e · outbound

This paper cites Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.683925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.295184Z digest=sha256:e83c23b3e45c9729244e790dea342c56bb7aa5c970c3beeebd528fd8ca6c0bba

Observation aa515a27-f2d8-4c0c-99c9-7d5b713dd5b7 · outbound

This paper cites Generation- based target speech extraction with speech discretization and vocoder,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Generation- based target speech extraction with speech discretization and vocoder,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.458042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.299420Z digest=sha256:73c49e376f181a98ebdbce14b8ced1eec00c9e6136bf6cff3ae4cdd239ae5fb5

Observation 25dbd4bf-1922-457e-abd1-dd3165775919 · outbound

This paper cites Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.667941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.303374Z digest=sha256:841274d0639dcaaa9930051335b217ff26def2b8e0534482c597f5ad106069af

Observation 72098b94-b6c8-4c80-a0a5-18864715c414 · outbound

This paper cites Diffusion-based signal refiner for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Diffusion-based signal refiner for speech separation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.307054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.307054Z digest=sha256:14030b2f1e0ab98f8ad93900cd147d9530a479b0994943f2d1a96ece7b84433a

Observation 4842be6c-412c-46b7-83ea-3ef0264395a8 · outbound

This paper cites Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:12.042351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.311336Z digest=sha256:8c4f4e2f55e3a5a71b1d309d0a199ec08e8d06219cf35f7d22d09571f2b77fdc

Observation 22f71e89-d785-415d-b410-649fa43d864b · outbound

This paper cites Ddtse: Discriminative diffusion model for target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Ddtse: Discriminative diffusion model for target speech extraction,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:11.826875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.315289Z digest=sha256:b3b15548bcabce3e6fcf660cb1315154a0729bd99007dce42ac3925a0ef20422

Observation 4c8713c5-9933-4c50-ba86-8f2066bbe735 · outbound

This paper cites Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:11.286750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.319275Z digest=sha256:41e46356cf219265cdea3388607dae95fbbe74bbc4e65e4789cad9de5ed3a4d5

Observation 52a51c3a-7be6-4adb-80a6-4269e55219eb · outbound

This paper cites X- vectors: Robust DNN embeddings for speaker recognition,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline X- vectors: Robust DNN embeddings for speaker recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.825851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.323186Z digest=sha256:5ece419844260df96608799ae364b65946992082965aede07a64c1346c243490

Observation 4f347263-8123-49eb-a228-bd5c259de03c · outbound

This paper cites Probing self-supervised learning models with target speech extraction,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Probing self-supervised learning models with target speech extraction,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.383416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.326526Z digest=sha256:84ff18fede12396020a31a9f75489e799527d57419cf09f236db8d42ca1b0c67

Observation 70347c62-4802-4526-b4b1-884eaeda3a8c · outbound

This paper cites Target speech extraction with pre-trained self-supervised learning models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speech extraction with pre-trained self-supervised learning models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.184130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.329756Z digest=sha256:2379e7413a0eb62fb532d263804de73e080beea43c15c73ff37ae71b798e2622

Observation a6131e0a-85aa-4405-8e71-6099872bfdb6 · outbound

This paper cites Smma-net: An audio clue-based target speaker extraction network with spectrogram matching and mutual attention,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Smma-net: An audio clue-based target speaker extraction network with spectrogram matching and mutual attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:10.039313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.333165Z digest=sha256:4f90485504b1606f62ab29ebf3db997b95f6775c67e20a4210c44412908a1cc3

Observation dc6868b9-d78f-42f0-848f-bb5f43732e07 · outbound

This paper cites Target speaker extraction by directly exploiting contextual information in the time-frequency domain,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speaker extraction by directly exploiting contextual information in the time-frequency domain,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.902393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.336782Z digest=sha256:fb3c1d12be4bac87d8555882154b2e2ae9177ee79ac0945db7029d5ef0b9ac74

Observation f0f6c734-18fe-4623-9839-4327c25098b4 · outbound

This paper cites Target speaker extraction with ultra-short reference speech by VE-VE framework,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Target speaker extraction with ultra-short reference speech by VE-VE framework,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.751317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.340177Z digest=sha256:634214c2c031414390264a3a4a0015ec5514d84c6900df8df49371afbd6e0b85

Observation cfcbd0d9-aa6b-4853-b2a3-7dd89ec2c885 · outbound

This paper cites Sef-net: Speaker embedding free target speaker extraction network,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Sef-net: Speaker embedding free target speaker extraction network,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.585458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.344503Z digest=sha256:a11fe394827ccba2533c33511148676af625f0d9712d69fa13cc48d147132110

Observation 6dc8bc8b-9d24-4d81-9c76-b40ddc55d071 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Common diffusion noise schedules and sample steps are flawed,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.460733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.347592Z digest=sha256:5d4d68cb2c8a6e9acf0d8b04249044f17c1118af99c076eb4c5c5150e3b37341

Observation 23f2da4c-8b3e-435f-ad39-abb56ef48d70 · outbound

This paper cites Progressive distillation for fast sampling of diffusion models,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Progressive distillation for fast sampling of diffusion models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.360479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.350993Z digest=sha256:239f98aeb8f30e1ca9e317961803b7bd267d23efad12a345dc4ff297387ae823

Observation f157c4e6-1f2e-4093-a974-7438ab5c6e9d · outbound

This paper cites High- fidelity audio compression with improved RVQGAN,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline High- fidelity audio compression with improved RVQGAN,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.211137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.354792Z digest=sha256:0b8aab995c85d2f84d988ca09c212aea8e07674f666850029798893d8dc363f1

Observation 010a6aa9-89f3-47cb-9fbb-8b2e41a8ab8f · outbound

This paper cites Stable Audio Open.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Stable Audio Open

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.358306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.358306Z digest=sha256:af712a05b4dd9636303f95e4b80f89d24f99e4f45e0712713ba1f34c9173b1d4

Observation 96e4134f-1562-4fc0-96fe-b16912434989 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.362796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.362796Z digest=sha256:a1b0c6e53dcd32949d12c3ae50679a5051727659feb59f603439f4dd94173cd0

Observation 15bc5cc7-3c30-4004-a002-2d6a7760ec0b · outbound

This paper cites Tf- gridnet: Integrating full- and sub-band modeling for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Tf- gridnet: Integrating full- and sub-band modeling for speech separation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:09.018402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.366834Z digest=sha256:3273202961f2f7edc84c41b307f1657cc35894d67fa97c492c37a9e1b25cec15

Observation e6a27d0f-83b0-47a4-b831-3b0e46ed4003 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SPMamba: State-space model is all you need in speech separation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.370792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.370792Z digest=sha256:40c925fbfa1e2cdf4238e63eec4d9f81e786f87b3f8cb45bb82f7ea729d566e9

Observation 11054119-9270-47de-9613-91804548ef4b · outbound

This paper cites Complex ratio masking for monaural speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Complex ratio masking for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.871580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.374906Z digest=sha256:e47af49ac50af1ad7a7acc16b369ff22fd5b6f6053126642bb8c2f50ec58cef1

Observation aadc30e0-17e9-4e6e-ba89-151dfc7f0c01 · outbound

This paper cites auraloss: Audio focused loss functions in pytorch,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline auraloss: Audio focused loss functions in pytorch,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.379255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.379255Z digest=sha256:8def2050a0e056fa0cdeee6fe81e7a22df19c80b84c5acc39447ed3663d552c9

Observation 54a7bc8f-5781-4c69-b7e1-ad5fe4ce80f4 · outbound

This paper cites High fidelity neural audio compression,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline High fidelity neural audio compression,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.674926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.383128Z digest=sha256:4acacc7f0716d877f34d4046d773c31780da2d20a585aa809cda756473da82a6

Observation de2e9680-5b1a-41a2-9c2f-441ef16b5157 · outbound

This paper cites Scalable diffusion models with transformers,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Scalable diffusion models with transformers,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.454405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.387236Z digest=sha256:bd818bc42c50245c0289e7743a691715ee1fe2b43f64be81e02a9d0543930a3a

Observation 31cfd854-1be9-4e3c-804e-60f026a25267 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.264101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.390920Z digest=sha256:bdd9b2dd9ab4a16240ca071c0bdd6e0391884c22ea52ce74444131c7e9a0cdc9

Observation a0bfd558-bae4-425f-91e7-49c5aed8eebf · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.395030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.395030Z digest=sha256:6b3c6474d11b48fb38d6b287c18f3e48dec677d9ef2ce7a73124314db70d118c

Observation 2f748628-177d-4a38-b1bd-f2f82efc59c0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Roformer: Enhanced transformer with rotary position embedding,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.143124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.399020Z digest=sha256:a9e35a83c8766d3177d23afaf78102d16443567945150d2398bdaadd44930cc8

Observation 0edd31f9-8f3b-4bb4-9518-44c1cbd8072d · outbound

This paper cites Single- channel multi-speaker separation using deep clustering,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Single- channel multi-speaker separation using deep clustering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.040444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.402914Z digest=sha256:8d2579f407e7d44f385aa53f9d114379b391dee67a64cad2ef297d7c7b401dad

Observation 75bdcf37-727d-4132-8a8d-0786387e5952 · outbound

This paper cites Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.028706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.406508Z digest=sha256:0b897cb7271f65ad0532629b29d9f1d9c23d357e4903cca1f80d68202c004520

Observation 4918245f-567b-4278-be56-986fac214ce5 · outbound

This paper cites Wham!: Extending speech separation to noisy environments,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wham!: Extending speech separation to noisy environments,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.017035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.410086Z digest=sha256:db98fe796059e4646b373a822516e8c4f526add5bb7d5932c56472cae44be30e

Observation 79627c97-0cab-4265-9437-c08bde14e8bb · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Librispeech: An ASR corpus based on public domain audio books,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:08.005422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.413724Z digest=sha256:c61bc3741085d7a9e34971ee33cc8193454901a6e203f025b3bdc3f59f6ba83b

Observation 73d7aa18-fb73-4828-97c6-8dd42e67a898 · outbound

This paper cites Improving speaker discrimination of target speech extraction with time-domain speakerbeam,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving speaker discrimination of target speech extraction with time-domain speakerbeam,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.993118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.417535Z digest=sha256:25898afda66bd470c09a6ed6db8f38cadba6fa1edff96093e1660e20f3d284f2

Observation 7deb9100-49d1-4a44-81eb-7d1dac4d1144 · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline MUSAN: A Music, Speech, and Noise Corpus

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.421318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.421318Z digest=sha256:4db3021fa69214cb9392805c9e3f164709ebb07517621fb88a077e99d6dda60e

Observation d5dd377b-9cfe-4767-b06e-9e8dc831bfbf · outbound

This paper cites Multichannel audio database in various acoustic environments,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Multichannel audio database in various acoustic environments,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.979644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.425128Z digest=sha256:fcb8146ad34722b521565f6b3d7acd12ed829be92ae2690d691eb2e718c06566

Observation 6be34bc1-4554-44a6-bbce-60ee64db0d1c · outbound

This paper cites The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline The fifth ’chime’ speech separation and recognition challenge: Dataset, task and baselines,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.966005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.429542Z digest=sha256:51e6c1629ba16e1431c3ffcb137b12cbc8ca8f59a82f2a126c6265d9d60d7f1a

Observation b89ca30e-01c2-46bf-b953-d0904c8f36e3 · outbound

This paper cites SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.433411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.433411Z digest=sha256:02ff2e5af0e3c4e5e5040bd64143b6c04d07d3da2aaae0ede66f9fd750743055

Observation 35abf0fb-854c-4bf8-ae44-0209e167cd0d · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.951031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.437013Z digest=sha256:4a21012dba6d5b9dbb027d6931f02c9b09057c0bd875ab28f56313746a7e195e

Observation e24fe973-e9ee-44e8-88a7-9561bd4b5938 · outbound

This paper cites An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.931359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.440926Z digest=sha256:fb3375bc244432f050ee1a8bedbc04ddf47e02039fed570fb3467c86a6256f32

Observation 9ca04559-6a14-4eb9-95b6-5a04d202f30a · outbound

This paper cites Dnsmos P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Dnsmos P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.911881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.445060Z digest=sha256:741a30f4fc1f9bc5df2fe2b99570c9877f775a8ad4ac94afde531049389fb74e

Observation a04353fe-156d-442a-ba94-d5b2718ed1be · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Robust speech recognition via large-scale weak supervision,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.888087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.453503Z digest=sha256:f15446fb1d22187f4aeba7513a8c2df24a66a42e4a646e33ac39a286ca5a3deb

Observation 7f382b3e-dce3-42ac-8325-5ba18d26efb8 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.457443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.457443Z digest=sha256:51486c57a12c9f768ab80821fe16a046a40263977c40ad239d9dbc992ada5b6f

Observation e5c1acbc-6245-400f-8509-f8751ee5b120 · outbound

This paper cites Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.533270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.460985Z digest=sha256:ec50804b5df3146e7d080ae3b9d567175c94b4b955b51d265d58c5640f813410

Observation 31812576-3bd2-4f26-836a-81dda307c00c · outbound

This paper cites Classifier-Free Diffusion Guidance.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Classifier-Free Diffusion Guidance

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.466008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.466008Z digest=sha256:cdc75896a7312c74cae186ef92e1eb1bf67dd9ef19cb455ccae52be45ad923e3

Observation 2da558ad-64b7-47c4-b910-1665ea66ae00 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Photorealistic text-to-image diffusion models with deep language understanding,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.869297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.470504Z digest=sha256:57ca5b273cd2203d65aaebec2dcbf4059c13daedec3f18114236d73318f60b81

Observation 6f599b9a-c3e8-4b82-9d3d-e8ebea2852e9 · outbound

This paper cites Available: https://openreview.net/forum?id=08Yk-n5l2Al.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Available: https://openreview.net/forum?id=08Yk-n5l2Al

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:07.857925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:22:07.475291Z digest=sha256:7d01e4682149a169adcc9eae579c062d62ca62bd7fd8fa1c751f65f2ece41904

Observation f9ec7a85-4eca-4bf2-b7a6-cc964b864693 · outbound

This paper cites an unresolved cited work.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Unresolved cited work

Reference 2022

Resolution
parse uncertain
no resolver link, observed 2026-08-07T14:22:07.449434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.449434Z digest=sha256:b19675f18e909745f975997ac809299954c40b20b6a632dc22d3c41502553aaa

Pith citing papers

Observation cc329c18-7386-4acd-8846-95a8f1672c13 · inbound

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model cites this paper.

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:49.132229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:49.132229Z digest=sha256:5aed234ba76221f3f01575d26dff99f79bbf4562c522785533d9182d637ea9bb

Observation 0293e581-3ff6-4913-8d5f-4960ebfa64ed · inbound

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios cites this paper.

Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:20.849476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:20.849476Z digest=sha256:78170b4365fc09b8701640d06baf1f2802b7f46c2e36e29402c2086a4c66fd67

Observation 29ebe771-cc57-4139-8478-42c13570334e · inbound

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition cites this paper.

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T17:09:15.281290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:09:15.281290Z digest=sha256:7823cd5abe5f57d5470b7dd6fd6b964ae92c45564a6afe947e69c4e52f1d8198