Pith. sign in

Paper Citation Record · LEDGER

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 7 inbound Pith citation observations for arXiv:2506.19398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19398 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.534728Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.362189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:08:08.743449Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69262a98-01cc-4375-8be0-ce2617879210 · outbound

This paper cites While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.131306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.356498Z digest=sha256:865688be7a1734e783ef8e1ffd241debf317df8fb1a4d29684af6d6eff66b5ce

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · outbound

This paper cites ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:87a67b46b1edb98d5c67737dc0b39bd57442df818437b26ad67ec71afae39f4d

Observation ffbe1fee-e561-42d0-b535-2900b9f5f465 · outbound

This paper cites Training strategies 3.1.1.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Training strategies 3.1.1

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:11:25.112661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.366582Z digest=sha256:1f21174c28efc30ff37679505ae5f772da17e3758df4ff0e736aa1aab93c5a04

Observation 72b8ed4b-36ef-4577-8a86-85b840ff6436 · outbound

This paper cites Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.095263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.371021Z digest=sha256:0c2896b63e9f5de1bb25b1e1e54a2624ee4728eb9871c92ab64bec14a9a4027c

Observation 641f26e5-0e0d-4eb6-aae4-80b3166b30ec · outbound

This paper cites Deep learning for audio signal processing,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Deep learning for audio signal processing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.078020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.376338Z digest=sha256:248d7a84cbd6f7d7332ab08fffc87f631e013df54830c09795a0bbbb499f35c5

Observation 84304d9c-cd5d-423d-9c93-b1b942667db1 · outbound

This paper cites Mamba in Speech: Towards an Alternative to Self-Attention.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Mamba in Speech: Towards an Alternative to Self-Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.380674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.380674Z digest=sha256:73ae97190d8d3155beeeda361bad3ecd4f1ab38093720c69fa35ff97c049970e

Observation 6082d4d3-b631-4c69-ad1b-0d83a76af01a · outbound

This paper cites DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.063117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.385499Z digest=sha256:d3e0e455280d3d7791d742983ca55f9f3a8e6cc75b8d2dc3161eda9ff447786a

Observation 2a35cea7-d23a-4042-93f2-bdc8a40e17dd · outbound

This paper cites SpeechBrain: A general-purpose speech toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A general-purpose speech toolkit,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.047106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.389448Z digest=sha256:7efe7a295591c11085dcfef3ac9d2c4b3d21cbd5805d8ba0663834fc2d968f81

Observation 651b4301-b4e2-413f-8fbd-95a5ae5587c8 · outbound

This paper cites AudioSR: Versatile audio super-resolution at scale,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment AudioSR: Versatile audio super-resolution at scale,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.978245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.416110Z digest=sha256:92cd70fd5e87a57e6d9ae75bf72c75459955bbd4543c66eb814441f1b1426ff9

Observation 0a8da161-6948-4863-9ea9-090d76f982a9 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ESPnet: End-to-end speech processing toolkit,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.032797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.399073Z digest=sha256:9ed60e4cd4539be7644944c1c8dc7d6a791878ce5d81538a4c2f3416890418fd

Observation d4baf6c7-b9da-4399-9bab-d6526be4a8df · outbound

This paper cites Summary on the multimodal information-based speech processing 2023 challenge,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Summary on the multimodal information-based speech processing 2023 challenge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.020116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.403507Z digest=sha256:22c32d22726599e21c5747a81755b8938f845ddfe9e1b1bbb1ca9e39f07d747e

Observation 51bb2732-9407-4c47-a76c-6aebd1a7ea79 · outbound

This paper cites As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.007293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.407313Z digest=sha256:c26e3ccb0ee2354dea3aec4d3af343a407ee911e492682435cdd6cda4e949ec2

Observation 95a7cd92-8c58-4079-bed2-62ae14bd37ca · outbound

This paper cites DeepFilterNet: Perceptually motivated real-time speech en- hancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepFilterNet: Perceptually motivated real-time speech en- hancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.993239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.411596Z digest=sha256:4fde0594ca4641d67d6127bd9888a0d0431a796e91f4c5211aeca861acd8c86a

Observation cc184939-26b3-4313-8d7f-57dec752ac13 · outbound

This paper cites Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.915062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.435224Z digest=sha256:852318e22146c89e6c892180ed4248cee16af41146cde997a72bffe83e4c2b44

Observation b444fbaa-fe93-4bfe-8501-0343628453ef · outbound

This paper cites FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.964648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.420128Z digest=sha256:e60cb029cbd175ae461fa1018e0d9fd01ce1e46786bfda35283b62b90679b8c2

Observation 60e16eda-03ee-4f93-af38-77890af77535 · outbound

This paper cites FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.950815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.424200Z digest=sha256:fd1939f1559ee07a74b68184fe96a852efc74a83d7a83b3d191e253c46df0ab7

Observation f9a92ae9-a8d1-46e3-9c7c-c7226a044498 · outbound

This paper cites MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.937368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.427496Z digest=sha256:2db495100c7ae06aca92d145855688d11bc0d753dd4b8ff273ea90b792ea3c4b

Observation 3f76367b-dfd3-42af-8818-5c8d59f4bbf4 · outbound

This paper cites MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.926695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.431574Z digest=sha256:e171a0640150af92732ca2fd1a7d6edd702b9b0c9a3eff89f57613282e8349bb

Observation 5658879f-5b6d-49b8-a6d9-67f991f513c3 · outbound

This paper cites NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:11:24.586002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.452262Z digest=sha256:d1e89707d82320cd21cb0c36ec8e1386186f277ee3bdeea4608b4496f816bb22

Observation de92d52c-4bf6-4860-8822-ba96854e3fe5 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.902975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.438524Z digest=sha256:52db07c27df18a9892d6aeaf5ed9e465a27f0bc46c2571741bf2d7a66e7798a0

Observation aabc8ee1-2c97-4e64-aa86-a8da807dc826 · outbound

This paper cites Scenario-aware audio-visual TF- Gridnet for target speech extraction,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Scenario-aware audio-visual TF- Gridnet for target speech extraction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.891068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.441735Z digest=sha256:8ff184296fa1792f7fe438245d5dc8b517bd2d6c328a2659e0e0c3d2fbc012b8

Observation d8c239ae-1c8d-4c32-859d-5764628568ce · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Speaker extraction with co-speech gestures cue,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.878761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.445326Z digest=sha256:aa525b4ae634c7f29d7eab41aedc027694160bb27576b481609917f7e2b9402b

Observation bac71848-182d-4296-86e5-a0f824999631 · outbound

This paper cites SpEx+: A complete time domain speaker extraction network,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpEx+: A complete time domain speaker extraction network,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.866309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.448614Z digest=sha256:a457046d4a2038603f04a5e8f64982efc52522f2bd36f7a6189fa8f5c4443d74

Observation ec631b81-503b-41f1-8a50-b5ab29dfa231 · outbound

This paper cites DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.800145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.472350Z digest=sha256:5c02857dda785974ca939aa10ccff88602f2d53b3a553de630ee686c6a18fc86

Observation 56a28c20-a421-4a64-9fef-8351e699a4d6 · outbound

This paper cites The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.852653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.456214Z digest=sha256:991da3a2db75681e93d2e939810e84cc3e9c94f36115697324a5568769027451

Observation 3f12de00-69f9-4ec6-9155-361df8278134 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.840151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.459983Z digest=sha256:37624361d805883f909adda51d2049af0fbc076119cace37edfb39e2a38e3f75

Observation 659d3a06-5962-4ebf-b2e5-03c71bfa0038 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Audio set: An ontology and human-labeled dataset for audio events,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.825770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.463692Z digest=sha256:1c93523807f4553cb637687d5dbc4537595aeef6ab9c09aec9aece1e5d2cece8

Observation 95fdaf25-0bc1-40ec-a534-9a7e5ab7d920 · outbound

This paper cites DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.812415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.468069Z digest=sha256:d89d7e621e73635d37c34fd380b3e1c2772e91a3716bd508dee025472c58812b

Observation 84ac29bb-be00-47d6-9960-e826474b4b6e · outbound

This paper cites Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.743338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.492570Z digest=sha256:6ab8605307f45d4699ed1069cb2f653d4cb4d1df0761d305c12c087d04f67f38

Observation ae34a824-1169-40b3-8f35-5e3e3c074c7b · outbound

This paper cites A mask free neural network for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment A mask free neural network for monaural speech enhancement,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.788296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.476740Z digest=sha256:204450ac04a2d72f69d97d3e8d54e4203313ed3e3150350ec1b9b9c7f8c30e1a

Observation 33a9f56d-411a-468a-9246-0760cb623b87 · outbound

This paper cites TridentSE: Guiding speech enhancement with 32 global tokens,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TridentSE: Guiding speech enhancement with 32 global tokens,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.777289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.480847Z digest=sha256:b3c1955c8b354b3076f314d2efd0f4731b6f21b9e313a2be59bbe3051e0ce23f

Observation 5ab46bd6-5543-4693-94c5-762f92a30bea · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text- to-speech,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment LibriTTS: A corpus derived from librispeech for text- to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.766236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.484737Z digest=sha256:53c950e309380bc1a7ad792fa1fde835eb28377184d735afee3f156115868715

Observation 37ab05c7-7c56-42e3-8b87-94489e6e4ff0 · outbound

This paper cites Explor- ing strategies for training deep neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Explor- ing strategies for training deep neural networks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.754403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.488634Z digest=sha256:bf715e29974c9eba6e5f4626a03b22237c9d45afee8dfdd99fbc52a54b28269a

Observation 1a55352d-409e-441e-8e99-b07c8b85bb82 · outbound

This paper cites An efficient encoder-decoder archi- tecture with top-down attention for speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment An efficient encoder-decoder archi- tecture with top-down attention for speech separation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.683161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.511237Z digest=sha256:e4c71e71c8d83394824182251122df00be83e207a8f8e3210394079b3c00bb84

Observation a3042b7b-1c2e-4e49-9fcb-38ad66742e36 · outbound

This paper cites CMGAN: Conformer-based metric gan for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CMGAN: Conformer-based metric gan for speech enhancement,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.730747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.496146Z digest=sha256:15b0225ffa52ba3f7f0c36f0df75fdcb9ef763c0a7c8061d360ec526981ff98d

Observation d3c53e10-c4f2-4b40-953d-c1a1122527eb · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.719254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.499581Z digest=sha256:43c2140fffe6891033332aaf6c49e3426eb2858d98af9714c2d9b5c2dc65b2d0

Observation 9f24950c-a1cd-4283-9cd1-8ab001d8e2c4 · outbound

This paper cites Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.706871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.503833Z digest=sha256:ac316e910d3ef05140165927be24ea691139fc3f9008165932e35120b181f4f6

Observation d376de3c-6851-490f-a8d7-be8a5d857958 · outbound

This paper cites Attention is all you need in speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Attention is all you need in speech separation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.695833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.507563Z digest=sha256:d08a65553a88b2eeeef9da6a32d7c27cfb12190d16f77a80ff1ae81081330676

Observation 22a3c68e-0b4a-41ee-a4f2-14383c8ead30 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Selective listening by synchronizing speech with lips,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.640308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.530987Z digest=sha256:3e9f5a1d6503e5eb7a14152742a8f795cc9e934afadaf59397b08c2aa038996e

Observation 07b9b447-c93c-4a44-9e37-c8d9f9885845 · outbound

This paper cites TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.515143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.515143Z digest=sha256:6653843bf8fc4ee44bc4b7d2054a3e95afaede01f0c7f72a943c27a7571a8207

Observation 100526fc-2a2d-40f2-a4e1-ed08c42b5971 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SPMamba: State-space model is all you need in speech separation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.519069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.519069Z digest=sha256:4d061cf74cf67b9b4d3ff27a4174a6d49acd5b01f837a5a1d7a72d4f92be02be

Observation a9606353-6c7f-47a4-a8e3-39ca7e2241b3 · outbound

This paper cites Time domain audio visual speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Time domain audio visual speech separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.663849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.523257Z digest=sha256:8a7d92400a5719cc0640018bc90a60307ad450002baae6bdbd3171c2a0354af5

Observation f5854460-642f-4afb-9029-0d21a5d977fe · outbound

This paper cites MuSE: Multi-modal target speaker extraction with visual cues,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MuSE: Multi-modal target speaker extraction with visual cues,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.651853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.526741Z digest=sha256:1fc3f28516677b0daf1e6ba56aeba1b8f25ebb0a6875d56de2b944e0dd53c314

Observation ea871c47-8135-4ac8-a3a6-cdb06e9d4e2a · outbound

This paper cites USEV: Universal speaker extraction with visual cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment USEV: Universal speaker extraction with visual cue,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.629442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T23:11:24.534728Z digest=sha256:e1eeb2da31feaf2844039da61e6a0b3e7d3bb9fcf404fc4b6aa38ececa639cda

Observation 3ae2ef09-51e3-4364-9d36-43eabf7fb5d6 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A General-Purpose Speech Toolkit

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.394051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.394051Z digest=sha256:e64362c5aa4a4339b9768410c57043044dc975c77212ec44e0fcc6a09412ef38

Pith citing papers

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · inbound

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment cites this paper.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:87a67b46b1edb98d5c67737dc0b39bd57442df818437b26ad67ec71afae39f4d

Observation 02546d07-92d4-4255-a078-45e13b8b6517 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.604200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.604200Z digest=sha256:3c0a0c5cf99600324616efb767f1d3613b2f85bcaf401eb7f24aec693618602b

Observation 54563704-4c79-430a-bd33-33beede6abb6 · inbound

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels cites this paper.

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:01.909103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:16:53.603971Z digest=sha256:62ebde55cec74e5eb851b040169a4eb64d78a7dfd6c32ec9fba64956d47b4214

Observation deafe58a-e056-440e-8781-a9f84efcd48a · inbound

Hierarchical Codec Diffusion for Video-to-Speech Generation cites this paper.

Hierarchical Codec Diffusion for Video-to-Speech Generation ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.283645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:12:01.260833Z digest=sha256:62f55a2a79577c40ee39e1d57e5bf4b9da866e20a7ceb2cfcd9a96c9e1172403

Observation c8e55c73-53a0-4cd6-8931-5862900542ea · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.459132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:a1988e74eb4d654878ec85d2b8c81ae85a4600670c422846f75bd5b9c89ee774

Observation 6c4486fa-1b94-4290-bd04-2b0d6369226b · inbound

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions cites this paper.

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:08:08.745045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:27:00.881610Z digest=sha256:f00bfb4a74d7cf932ab70228affc0a40d10d6514709b98af282c035cb3a207d3

Observation f08e5579-001a-4a99-96e0-c7ceb8cf0d7d · inbound

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis cites this paper.

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T06:37:05.257817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:37:05.257817Z digest=sha256:b5ea5a339bc6d93bd3a6676c055be1f41f2520a839029744091913e167567a59