Pith. sign in

Paper Citation Record · LEDGER

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

As of 11 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2501.14477.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14477 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:10:46.729901Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:37:08.860549Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:22:07.662926Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e64515a8-7a55-4d01-a669-1be15a21e8d1 · outbound

This paper cites Supervised speech separation based on deep learning: An overview,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Supervised speech separation based on deep learning: An overview,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.582092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.582092Z digest=sha256:f72ebbec0fdca75d416b803aba7c0b5d743540885dc28a9badb3a7004bbc20a1

Observation 3e8c4950-15d6-425b-a325-e7a2318b43bb · outbound

This paper cites SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.195307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.623637Z digest=sha256:5fb5213617bd21153d9631deeabbd384ff3976a685ac6c0d17553caa7e696da3

Observation fcfe31bd-1c2b-4892-b1df-dae02062b0af · outbound

This paper cites Improving speaker discrimination of target speech extraction with time-domain speakerbeam,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Improving speaker discrimination of target speech extraction with time-domain speakerbeam,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.185465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.627506Z digest=sha256:6c26e6267c2a11c92c0ebdf399e7ce3268417be7129613171e663ff37709d12a

Observation 93bdc0e6-9656-4e04-ae97-cd9bbbb1a8f9 · outbound

This paper cites SpEx: Multi-scale time domain speaker extraction network,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR SpEx: Multi-scale time domain speaker extraction network,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.631159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.631159Z digest=sha256:c53a62a2a6ace45a97106bb9c940af03551703420e9fbc99a09b1104a49b674b

Observation d4c717d8-3fbb-445e-8cac-eea030e9a485 · outbound

This paper cites Target speech extraction with conditional diffusion model,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Target speech extraction with conditional diffusion model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.167219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.635131Z digest=sha256:57b7b2e7f018af1edb7adab21620d13005d32d39219301abca4dc08dbf23ece7

Observation a8da5e04-21df-499d-88c1-a9a7da46c977 · outbound

This paper cites Generation- based target speech extraction with speech discretization and vocoder,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Generation- based target speech extraction with speech discretization and vocoder,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.156907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.638354Z digest=sha256:815219be90f399ae687c19f4161708f074814dd8a8d306729995887ea4f34132

Observation 72fe91f6-2256-4f34-859b-83c10e4fe051 · outbound

This paper cites TSELM: Target Speaker Extraction using Discrete Tokens and Language Models.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR TSELM: Target Speaker Extraction using Discrete Tokens and Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.642096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.642096Z digest=sha256:fc949faa8c8f78b89fb1f8bb9ecb47a9887ac63355c4708acff8f1b0bcd80edf

Observation d37f81d3-ac7a-4cde-8dce-ef0c2d00def9 · outbound

This paper cites X-Sepformer: End-to-end speaker extraction network with explicit optimization on speaker confusion,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR X-Sepformer: End-to-end speaker extraction network with explicit optimization on speaker confusion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.147120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.645463Z digest=sha256:c4b21d88940110067cf1d486f41fbfb02d29277850b34a937f81ab1ea3041d7e

Observation d282659c-dd21-4b2d-8c79-85235707f1db · outbound

This paper cites Personalized speech enhancement combining band-split rnn and speaker attentive module,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Personalized speech enhancement combining band-split rnn and speaker attentive module,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.121244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.648974Z digest=sha256:eaa70bbd19aa19b481d64f3b99d738b3976e45583f80f628162b10432e7d8ee5

Observation 13178b6d-e481-4a2e-9227-e84482ef5244 · outbound

This paper cites Multi-Level Speaker Representation for Target Speaker Extraction.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Multi-Level Speaker Representation for Target Speaker Extraction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.651894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.651894Z digest=sha256:1cc984ba3c596d45cb514d269c7a3a7f43c43870014b0fb5a88a60b2160bb9de

Observation fc4661f7-00c1-45cf-b940-995b05860785 · outbound

This paper cites Hierarchical speaker representation for target speaker extraction,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Hierarchical speaker representation for target speaker extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.101057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.655479Z digest=sha256:ad1e761e92307688071f9455b3b0c9fc8276450238bf5ef4e02272bf42ae8c22

Observation f9ef6128-fec9-402f-b9c4-c32700df0e34 · outbound

This paper cites Extending Whisper with prompt tuning to target-speaker ASR,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Extending Whisper with prompt tuning to target-speaker ASR,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.086726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.658867Z digest=sha256:6f170a2fe2ca3eb2ff75eb66947829f896757e9da9b4054c9b2b7552565aaf81

Observation ed4a4fa8-2d2c-479b-b71b-863bd58dde34 · outbound

This paper cites Empowering Whisper as a joint multi-talker and target-talker speech recognition system,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Empowering Whisper as a joint multi-talker and target-talker speech recognition system,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.076020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.661519Z digest=sha256:0c0ac433ba6d888c7249529d93e8a4a0ceb44dfd5dcca261e6ef79708e107480

Observation e342ca0a-6445-478a-bb8e-179f59ca1356 · outbound

This paper cites SQ-Whisper: Speaker-querying based Whisper model for target-speaker ASR,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR SQ-Whisper: Speaker-querying based Whisper model for target-speaker ASR,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.064932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.664308Z digest=sha256:4221b417874a1193a49569f20d710dbcc8cbaecbfb2fe770c177847bae79ef06

Observation 17239752-4b59-4f8a-a1cf-0f039e3fa314 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Robust speech recognition via large-scale weak super- vision,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.052192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.667861Z digest=sha256:d14a9a049588cac29290a55832a7d65fd4b474c254b582cd0ed084d678beecb7

Observation e9c76bbf-e0b0-45ce-906c-83d722439235 · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR LoRA: Low-rank adaptation of large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.037943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.670704Z digest=sha256:815463ef92130b01825471d5a2dc8ed5877bde1f4187a39303a42f595e4ea6b3

Observation 96b622a0-1d44-4d2e-b856-4e7e6a118a3d · outbound

This paper cites Flow matching for generative modeling,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Flow matching for generative modeling,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.673387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.673387Z digest=sha256:e25c3748a02dbd274f69e76fe98eb0bb781df166f3906d4567b2de547ae10619

Observation e4c93146-3020-42a2-9c87-b2d2bfd1be34 · outbound

This paper cites Matcha- TTS: A fast TTS architecture with conditional flow matching,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Matcha- TTS: A fast TTS architecture with conditional flow matching,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:47.008939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.675876Z digest=sha256:d5c3562a3dd0d00178d50c314f9d64917625f25aa4b85d5ce98819350b5945a2

Observation 6fa45df7-e2a2-428e-ae49-fbae5c7f5d18 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.678571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.678571Z digest=sha256:596afde331c1f1afaf7d5a0e9be1422a5748e96cfb015a5dbc0431b123b46128

Observation cd89a5ae-8e9b-4f10-b316-0089f7cd367d · outbound

This paper cites LibriMix: An open-source dataset for generalizable speech separation,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR LibriMix: An open-source dataset for generalizable speech separation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.986025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.681445Z digest=sha256:43a834460440a7013eae0b1920bb40a2db778e33f88487a07a6ff013c2609025

Observation e22286fa-806d-41bd-b8aa-2d4b69a430bd · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Deep clustering: Discriminative embeddings for segmentation and separation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.970794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.684511Z digest=sha256:55215ebda20de757eeea32b44e181f2f578a2c753edd52b3aba7465979294030

Observation 25e09a36-2e1e-4397-a44f-14f7a1b8b9b7 · outbound

This paper cites Attention is all you need,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Attention is all you need,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.687909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.687909Z digest=sha256:b0fa6ca0d5920e204fc781c8d0e25707c2ba8a4381ba9e0c66dc1e5d711456c8

Observation d860a433-1951-4857-ab67-e1a204d7d755 · outbound

This paper cites Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:10:46.790932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.691658Z digest=sha256:0d4b3fff8df781c3d693c563680156e24bedc4f71403a41c7d4de4e12da0bf4b

Observation 1f28baf5-cf7d-4094-997d-1e7b7545b1e9 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.695561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.695561Z digest=sha256:1266156f2ed6128a79b8c6424da48c2aab2fe2c885f4fa8c535f8554d9ffb9e2

Observation e0eb6be9-90a9-41bf-a34e-fc2692bc919e · outbound

This paper cites HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.937486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.699254Z digest=sha256:08e3d0a68d97b0c4dbd32b97864de88d532576ec498b63505b4165396e14714d

Observation 23f16fff-4312-4732-a716-753ebda57677 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR LibriSpeech: An ASR corpus based on public domain audio books,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.702516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.702516Z digest=sha256:ac382648d43792e8efa3b11cba3b73b8aae54615c3f58287651ec2c6167518b5

Observation 082f1ea6-a0ad-41f0-b487-b0394f311f39 · outbound

This paper cites Adapting self- supervised models to multi-talker speech recognition using speaker embeddings,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Adapting self- supervised models to multi-talker speech recognition using speaker embeddings,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.903210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.705345Z digest=sha256:a6d2d5c1f7884a8a8d20b46c6c8dbf68be376449e4ee11261b92779802701f77

Observation ea68ea39-dc31-43a3-b84c-dd07cc4a4ceb · outbound

This paper cites Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.889979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.708854Z digest=sha256:256aa55429cf5cc5d6f1b08d35292b6ea66eef7b389900ff87dd60be510dc315

Observation b8afd31f-848a-4962-82aa-76c6d4097dc5 · outbound

This paper cites DNSMOS P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR DNSMOS P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.879866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.712181Z digest=sha256:599151cf0803e6ecfc11f2bbb9c2ed8e1a043b60a6b193c6e8fba4689ae8942f

Observation 1c594bd9-7a28-49c4-9b0c-fd684e0db7b3 · outbound

This paper cites SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.715255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.715255Z digest=sha256:0ce9ecb69a6de71e9df559e9fbf99e56b8de4566c9756f7ad5b93b85f5e24b12

Observation 5720d9b5-9734-4c39-b064-5348efef89f8 · outbound

This paper cites CAM++: A fast and efficient network for speaker verification using context-aware masking,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR CAM++: A fast and efficient network for speaker verification using context-aware masking,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.869963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.718372Z digest=sha256:7f8862cf9e0d15e7811e4f17a64806ec35a8894a44ace61847c71b1f7e6cf1e5

Observation d1452fe5-15e0-49d6-9c63-f139780b8ebf · outbound

This paper cites SELM: Speech enhancement using discrete tokens and language mod- els,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR SELM: Speech enhancement using discrete tokens and language mod- els,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.859584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.721074Z digest=sha256:2df6c7ab93f811b47ee0cfe52e8a9a22abdc4cba799ba4697dd00911a3571b87

Observation ce9906db-1e5f-4ad5-8f8a-f117e3369987 · outbound

This paper cites Decoupled weight decay regularization,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Decoupled weight decay regularization,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T15:10:46.726366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:10:46.726366Z digest=sha256:093c6d042c0fb92a2f8f89c2cf7d48a4bdd55eaa503f336ca4fffb930b62fc57

Observation b19be8a5-f6a0-4511-a199-e0cdff379eb3 · outbound

This paper cites Music source separation with band-split rope transformer,.

Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR Music source separation with band-split rope transformer,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:10:46.839711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:10:46.729901Z digest=sha256:852e47d1615c6ce741d0c83a94099bf719de1421b1224bc0dc2fd2b176cc611e

Pith citing papers

Observation df8b61db-204e-4fd2-b94b-825f583d3405 · inbound

FlowTSE: Target Speaker Extraction with Flow Matching cites this paper.

FlowTSE: Target Speaker Extraction with Flow Matching Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:08.860549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:37:08.860549Z digest=sha256:ed448c4563042ac79bcf4ace5ccf6dc3c4d86270d0a9291a2bc3b9fe2b35435e

Observation 25dbd4bf-1922-457e-abd1-dd3165775919 · inbound

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline cites this paper.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:22:07.667941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:22:07.303374Z digest=sha256:e3caff1e5c36d06b7007bb9b0dc5647e3c37823a719567cd39cdbe453195b921