Pith. sign in

Paper Citation Record · LEDGER

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 7 inbound Pith citation observations for arXiv:2506.19398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19398 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.534728Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:24.362189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:08:08.743449Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69262a98-01cc-4375-8be0-ce2617879210 · outbound

This paper cites While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment While crucial for these applications, ac- curately processing speech is challenged by the often degraded quality of real-world audio

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.131306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.356498Z digest=sha256:cf0118745fb68eb74c50ff1aa0fa1bbc5b1bd18fcaa6d4ca6484c5305ea4753a

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · outbound

This paper cites ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:cf2efbdaf598672e5975ec3129b07b9ab1b48397c2549dabad75d828cc0967d6

Observation ffbe1fee-e561-42d0-b535-2900b9f5f465 · outbound

This paper cites Training strategies 3.1.1.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Training strategies 3.1.1

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T23:11:25.112661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.366582Z digest=sha256:e0f1d685e24163118056a5135feeba43ec2057a90755b00e8b6d98448275aef3

Observation 72b8ed4b-36ef-4577-8a86-85b840ff6436 · outbound

This paper cites Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Beyond the presented evaluations, ClearerV oice- Studio is available for live demos on HuggingFace and Mod- elScope, enabling users to experiment with real-world record- ings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.095263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.371021Z digest=sha256:067b8a48ac42a2eb984bcd368953798f9661fb4757299f9ad84144378c26a29e

Observation 641f26e5-0e0d-4eb6-aae4-80b3166b30ec · outbound

This paper cites Deep learning for audio signal processing,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Deep learning for audio signal processing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.078020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.376338Z digest=sha256:71e8cdf59820baa1e9b16348ba1a2e9ab8455c287dd72c574e41b38e0fad5ae1

Observation 84304d9c-cd5d-423d-9c93-b1b942667db1 · outbound

This paper cites Mamba in Speech: Towards an Alternative to Self-Attention.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Mamba in Speech: Towards an Alternative to Self-Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.380674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.380674Z digest=sha256:11d69a0b4d85ff3206a8044118f5ec5f188be455e3e333409709bfe751c05ac6

Observation 6082d4d3-b631-4c69-ad1b-0d83a76af01a · outbound

This paper cites DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.063117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.385499Z digest=sha256:1ec96eefa75756b3c3dce812425b30238d9bde27d6e3811a7b1e511750bfef14

Observation 2a35cea7-d23a-4042-93f2-bdc8a40e17dd · outbound

This paper cites SpeechBrain: A general-purpose speech toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A general-purpose speech toolkit,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.047106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.389448Z digest=sha256:df2e713adf34d58b3372bc19e7c2aa8969a941ca58aafaabbdc4c2248ad92636

Observation 651b4301-b4e2-413f-8fbd-95a5ae5587c8 · outbound

This paper cites AudioSR: Versatile audio super-resolution at scale,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment AudioSR: Versatile audio super-resolution at scale,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.978245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.416110Z digest=sha256:0a76099e6ba6c91409df5d1af261e631a7ada893dc6dcb551874248d5ece18ae

Observation 0a8da161-6948-4863-9ea9-090d76f982a9 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ESPnet: End-to-end speech processing toolkit,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.032797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.399073Z digest=sha256:1396de5e989fc59ee67afc8edd68252433191fbd24c82debe1ea644e02c8121c

Observation d4baf6c7-b9da-4399-9bab-d6526be4a8df · outbound

This paper cites Summary on the multimodal information-based speech processing 2023 challenge,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Summary on the multimodal information-based speech processing 2023 challenge,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.020116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.403507Z digest=sha256:cbacde02b92aea19b6c2b67d593913a1b4793a7cb6e2e528e933486ca9483b66

Observation 51bb2732-9407-4c47-a76c-6aebd1a7ea79 · outbound

This paper cites As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment As- teroid: the PyTorch-based audio source separation toolkit for re- searchers,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:25.007293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.407313Z digest=sha256:aa315304e3a9dc79470f43e2a9fadf817132b5fcf42ad98fbd75747894f8a560

Observation 95a7cd92-8c58-4079-bed2-62ae14bd37ca · outbound

This paper cites DeepFilterNet: Perceptually motivated real-time speech en- hancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DeepFilterNet: Perceptually motivated real-time speech en- hancement,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.993239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.411596Z digest=sha256:51be7c256a445bef0445e45224e476abf6eaabea8c3370dace4b1b2227775526

Observation cc184939-26b3-4313-8d7f-57dec752ac13 · outbound

This paper cites Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Hifi-SR: A unified generative transformer-convolutional adversarial net- work for high-fidelity speech super-resolution,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.915062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.435224Z digest=sha256:d260c1ef7d1b81cd24b0021689cd7e5a5cdd1192bfa021d51bf59b086413fe4f

Observation b444fbaa-fe93-4bfe-8501-0343628453ef · outbound

This paper cites FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FlowA VSE: Ef- ficient audio-visual speech enhancement with conditional flow matching,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.964648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.420128Z digest=sha256:30c45c088306b41e7d8055a20d99db8c4d46d90fe85516694e3700414502de96

Observation 60e16eda-03ee-4f93-af38-77890af77535 · outbound

This paper cites FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.950815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.424200Z digest=sha256:ef633120972444328546325a1ed99985711bec4eadd4cb897eaaa1cf05c0ae98

Observation f9a92ae9-a8d1-46e3-9c7c-c7226a044498 · outbound

This paper cites MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer2: Combin- ing transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.937368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.427496Z digest=sha256:253587e7057a964abf79ad4070bebbd51367a592cd191e5a1d1f88e5617ffb47

Observation 3f76367b-dfd3-42af-8818-5c8d59f4bbf4 · outbound

This paper cites MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MossFormer: Pushing the performance limit of monaural speech separation using gated single-head trans- former with convolution-augmented joint self-attentions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.926695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.431574Z digest=sha256:91525c7e914754e9f0fff43cc8b31507ff15781a45d7bb195037942bf1624b48

Observation 5658879f-5b6d-49b8-a6d9-67f991f513c3 · outbound

This paper cites NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:11:24.586002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.452262Z digest=sha256:fc6ad94d7c9fd19b866d07a8b6d40602544bc43f8a6997bbe0bb451b9ddd0c6c

Observation de92d52c-4bf6-4860-8822-ba96854e3fe5 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.902975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.438524Z digest=sha256:6fb2d2515e4f32f5399597f852f0c2b15e25106ed5909063edfada5d6ab6f081

Observation aabc8ee1-2c97-4e64-aa86-a8da807dc826 · outbound

This paper cites Scenario-aware audio-visual TF- Gridnet for target speech extraction,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Scenario-aware audio-visual TF- Gridnet for target speech extraction,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.891068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.441735Z digest=sha256:55da914ebe8451a7678b9dfcc136ec8616765413f7935724e9042453f0d8a8bd

Observation d8c239ae-1c8d-4c32-859d-5764628568ce · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Speaker extraction with co-speech gestures cue,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.878761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.445326Z digest=sha256:3574685a823da083437e0dec1821eeedf5e44706d8041259a8c990157f778d05

Observation bac71848-182d-4296-86e5-a0f824999631 · outbound

This paper cites SpEx+: A complete time domain speaker extraction network,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpEx+: A complete time domain speaker extraction network,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.866309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.448614Z digest=sha256:fdc805a68c05fa90ba149e74a778bd28fef3ff4fa381c3962eaf0a790751467b

Observation ec631b81-503b-41f1-8a50-b5ab29dfa231 · outbound

This paper cites DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DCCRN+: Channel-wise subband dccrn with snr estimation for speech enhancement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.800145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.472350Z digest=sha256:57d0a5c0bad6c576ee6984745d793b0bcb1d38efd1ab936ee6b676e6086abf94

Observation 56a28c20-a421-4a64-9fef-8351e699a4d6 · outbound

This paper cites The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.852653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.456214Z digest=sha256:7710d405af3b02275c6e444c4a39d95ff081d01fb6380c345c278debe340c3eb

Observation 3f12de00-69f9-4ec6-9155-361df8278134 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.840151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.459983Z digest=sha256:151e633ebe34a6fc76866276ae3073c42fdf3e1420ecc87b8acd3b8cbeeaa1ea

Observation 659d3a06-5962-4ebf-b2e5-03c71bfa0038 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Audio set: An ontology and human-labeled dataset for audio events,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.825770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.463692Z digest=sha256:879a66af201638144eab4049b85e8fbfa5bda30edc15211c475a9db9e243333b

Observation 95fdaf25-0bc1-40ec-a534-9a7e5ab7d920 · outbound

This paper cites DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment DEMAND: a collection of multi-channel recordings of acoustic noise in diverse environments,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.812415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.468069Z digest=sha256:6a0d42b67ea7bd80495b18c82ac21e0eaf24bc871888b59db6997965d997477d

Observation 84ac29bb-be00-47d6-9960-e826474b4b6e · outbound

This paper cites Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Phase- sensitive and recognition-boosted speech separation using deep recurrent neural networks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.743338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.492570Z digest=sha256:4eea7f035a01233da61798b2b56bc2a70cacf77ac00639e76d1ab1e4e24f01ba

Observation ae34a824-1169-40b3-8f35-5e3e3c074c7b · outbound

This paper cites A mask free neural network for monaural speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment A mask free neural network for monaural speech enhancement,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.788296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.476740Z digest=sha256:21bc26459aa53fb785f9a93729a16db8a63de775d690dd46beffd15aebd78a62

Observation 33a9f56d-411a-468a-9246-0760cb623b87 · outbound

This paper cites TridentSE: Guiding speech enhancement with 32 global tokens,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TridentSE: Guiding speech enhancement with 32 global tokens,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.777289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.480847Z digest=sha256:47a05a986eb14135904b2154114093edda1edad0606cdb876637c3dad88ca743

Observation 5ab46bd6-5543-4693-94c5-762f92a30bea · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text- to-speech,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment LibriTTS: A corpus derived from librispeech for text- to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.766236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.484737Z digest=sha256:61111255134de2647a57b7d3b5abb34fe9a72260ced01de75048bf1ed570faec

Observation 37ab05c7-7c56-42e3-8b87-94489e6e4ff0 · outbound

This paper cites Explor- ing strategies for training deep neural networks,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Explor- ing strategies for training deep neural networks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.754403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.488634Z digest=sha256:8e4a0556d4c5f186845343f02db6d3d7efc12f32c840720c63450c00c6a0848e

Observation 1a55352d-409e-441e-8e99-b07c8b85bb82 · outbound

This paper cites An efficient encoder-decoder archi- tecture with top-down attention for speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment An efficient encoder-decoder archi- tecture with top-down attention for speech separation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.683161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.511237Z digest=sha256:01d9e2b2ab4c3161b1d7b20b11aaea42029154c88b4c3eea89ee8a7532d5cc65

Observation a3042b7b-1c2e-4e49-9fcb-38ad66742e36 · outbound

This paper cites CMGAN: Conformer-based metric gan for speech enhancement,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment CMGAN: Conformer-based metric gan for speech enhancement,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.730747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.496146Z digest=sha256:c6c196d6a15ade55e25db5ee539fe3770d138ae5927ff54650fa263e799c9e7c

Observation d3c53e10-c4f2-4b40-953d-c1a1122527eb · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.719254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.499581Z digest=sha256:3824ea43e51b6072e9d32ad3f70875971ef321fcfd7074f934978baac8deedb3

Observation 9f24950c-a1cd-4283-9cd1-8ab001d8e2c4 · outbound

This paper cites Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.706871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.503833Z digest=sha256:b5c3e5424b835cd657e3a006d28cd0467af8efce62c1f73efc4f25531f102889

Observation d376de3c-6851-490f-a8d7-be8a5d857958 · outbound

This paper cites Attention is all you need in speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Attention is all you need in speech separation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.695833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.507563Z digest=sha256:949e57953a563f2802293781cc32c6ef038a1c92d21982dd2345208004f87ca3

Observation 22a3c68e-0b4a-41ee-a4f2-14383c8ead30 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Selective listening by synchronizing speech with lips,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.640308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.530987Z digest=sha256:3702759e61046e89d0e059b8204b02325f8f483f646829188ea39b9a315d154b

Observation 07b9b447-c93c-4a44-9e37-c8d9f9885845 · outbound

This paper cites TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.515143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.515143Z digest=sha256:b4aa317c19c50fe8ef1d0d21bcf687cdddff67298e333923bd7f2733bc858586

Observation 100526fc-2a2d-40f2-a4e1-ed08c42b5971 · outbound

This paper cites SPMamba: State-space model is all you need in speech separation.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SPMamba: State-space model is all you need in speech separation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.519069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.519069Z digest=sha256:c7e5b0ed11e09081105dda54e2c43b73dddf5c9cf86b2ce050c410d84e972738

Observation a9606353-6c7f-47a4-a8e3-39ca7e2241b3 · outbound

This paper cites Time domain audio visual speech separation,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment Time domain audio visual speech separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.663849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.523257Z digest=sha256:8d529e03dc4c024225654d3b10da3ffabf4994c990d867abda2f1d512ac3f47a

Observation f5854460-642f-4afb-9029-0d21a5d977fe · outbound

This paper cites MuSE: Multi-modal target speaker extraction with visual cues,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment MuSE: Multi-modal target speaker extraction with visual cues,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.651853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.526741Z digest=sha256:705df45a107a677beb9bb16bcd4c46a26217eaacd2e515d4c6877456d819ea97

Observation ea871c47-8135-4ac8-a3a6-cdb06e9d4e2a · outbound

This paper cites USEV: Universal speaker extraction with visual cue,.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment USEV: Universal speaker extraction with visual cue,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:11:24.629442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T23:11:24.534728Z digest=sha256:29c5c4663a1d6f05dbb8fad900f43d84890017713681501993acb90d1c606bc0

Observation 3ae2ef09-51e3-4364-9d36-43eabf7fb5d6 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment SpeechBrain: A General-Purpose Speech Toolkit

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.394051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.394051Z digest=sha256:f6c0deada777bdd6574dcb1e1d2f6509abc9446eef3352f4428be10b57a06584

Pith citing papers

Observation 714fd3b6-e2d1-45c8-b151-2d020951eb2c · inbound

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment cites this paper.

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:24.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:24.362189Z digest=sha256:cf2efbdaf598672e5975ec3129b07b9ab1b48397c2549dabad75d828cc0967d6

Observation 02546d07-92d4-4255-a078-45e13b8b6517 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.604200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.604200Z digest=sha256:3c9b9163dc34c3170fa50cb2ca4aceaf927230a3378d04ee217cbfb58057255c

Observation 54563704-4c79-430a-bd33-33beede6abb6 · inbound

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels cites this paper.

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:01.909103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:16:53.603971Z digest=sha256:921b1d84b2119cde3c260236946f05a8ab5b48fe5bb33a7ee1b18aa77fab177f

Observation deafe58a-e056-440e-8781-a9f84efcd48a · inbound

Hierarchical Codec Diffusion for Video-to-Speech Generation cites this paper.

Hierarchical Codec Diffusion for Video-to-Speech Generation ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:12:26.283645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:12:01.260833Z digest=sha256:5431f5005db54e99cad6efa3d7a8d70e0c6b9d04d75c6d82a36f8998cb58a5b0

Observation c8e55c73-53a0-4cd6-8931-5862900542ea · inbound

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities cites this paper.

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:17:57.459132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T23:17:08.124240Z digest=sha256:59ee4c1eb6017591309bd07c4625a949d930d3ca69da1cfd7829167d89443d9c

Observation 6c4486fa-1b94-4290-bd04-2b0d6369226b · inbound

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions cites this paper.

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:08:08.745045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:27:00.881610Z digest=sha256:42c534a4d540a9fa1c210f361e85e393a4bb35a9ad150d36350d79745bc20fd7

Observation f08e5579-001a-4a99-96e0-c7ceb8cf0d7d · inbound

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis cites this paper.

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T06:37:05.257817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:37:05.257817Z digest=sha256:41f9c6284f5e53d4c066c89bea6275322113cb9d90fea3293cfce8b3e150b197