Pith. sign in

Paper Citation Record · LEDGER

Beyond Speaker Identity: Text Guided Target Speech Extraction

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2501.09169.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09169 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:14:00.754527Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5effb835-1de0-438e-89e6-e1e2f390e16f · outbound

This paper cites V oiceFilter: Targeted V oice Separation by Speaker-Conditioned Spectrogram Masking,.

Beyond Speaker Identity: Text Guided Target Speech Extraction V oiceFilter: Targeted V oice Separation by Speaker-Conditioned Spectrogram Masking,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.355088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.599522Z digest=sha256:a847612b7aaa25e86cbe3ee7798e0b26303158b10f8f15d10db782189270c3b4

Observation 943e1499-b812-41fa-bb79-310e48e11aae · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.605275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.605275Z digest=sha256:b384dc1b1421dcdf6e45dd2f18b29058ec4462ec60be3414196aba77fa01908c

Observation 7b4987b4-d832-4ac2-aae7-a985417ee1c3 · outbound

This paper cites Spex+: A complete time domain speaker extraction network,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Spex+: A complete time domain speaker extraction network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.326227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.610679Z digest=sha256:f6eff9f6d58dde03c81cb090b20efdcd45fc9ae76287feee56543577f55245b1

Observation f0fbe5aa-2ab5-461b-9476-76768b5506bd · outbound

This paper cites Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.308630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.615829Z digest=sha256:6731f654614689605530088a853ca55ebffb6fa541cc9f25b564ebf08db8f137

Observation 76c6c715-309f-4dc5-9c9f-3b9eb7f7f082 · outbound

This paper cites Multimodal SpeakerBeam: Single channel target speech extraction with audio-visual speaker clues.

Beyond Speaker Identity: Text Guided Target Speech Extraction Multimodal SpeakerBeam: Single channel target speech extraction with audio-visual speaker clues

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.269058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.621232Z digest=sha256:c3d7533910cb0a26d2eb82d555567e7529f6ae6269e90dfbc48329a23868ba83

Observation 0d3fb154-6984-4e9f-8660-e71b78b8e072 · outbound

This paper cites Text-driven separation of arbitrary sounds,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Text-driven separation of arbitrary sounds,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.241447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.625979Z digest=sha256:2c574f92cf23e8cd67c9aceb7612e39846cd1931cb1c6a883a7fa114c6a39766

Observation d5f7358f-6ef5-4371-a0e1-12e6f3660e6b · outbound

This paper cites CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,.

Beyond Speaker Identity: Text Guided Target Speech Extraction CLIPSep: Learning text-queried sound separation with noisy unlabeled videos,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.213502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.631961Z digest=sha256:75ece312be177a9fc66fd76f38004120cb196b2bb645f9ad994dfcdc84a51fb6

Observation 790022a0-24b9-470c-a242-b5139b912fd3 · outbound

This paper cites Separate Anything You Describe.

Beyond Speaker Identity: Text Guided Target Speech Extraction Separate Anything You Describe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.636550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.636550Z digest=sha256:1308aecd443ffe1cc69e8ca1eac11d3d77e1de72753f6a7c16e1973c93e50f68

Observation d799943c-7b47-4695-a756-f868eecb56de · outbound

This paper cites Target sound extraction with variable cross-modality clues,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Target sound extraction with variable cross-modality clues,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.193632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.641820Z digest=sha256:8d4b67766b6a9e227f1ed1294cf553dd0dce8da41e2055321702e4caa7952281

Observation c8714b0d-7845-4bf9-979a-8fafcb4697ea · outbound

This paper cites CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction.

Beyond Speaker Identity: Text Guided Target Speech Extraction CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.646171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.646171Z digest=sha256:f11c613a9369efe3eb74f0b871e1d901050aaea5d69228c359c97b6793bf3180

Observation 4577da83-0503-427c-b818-f8e6f8c9f220 · outbound

This paper cites Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction.

Beyond Speaker Identity: Text Guided Target Speech Extraction Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.651546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.651546Z digest=sha256:161fe4739819c7b594f6a9ee25cdc11accdb4e49546139f4b421fddfbbfc2259

Observation 529930c3-ceed-44b9-968d-e6124fc551b3 · outbound

This paper cites Target Speech Diarization with Multimodal Prompts.

Beyond Speaker Identity: Text Guided Target Speech Extraction Target Speech Diarization with Multimodal Prompts

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:14:00.859643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.657124Z digest=sha256:7db10f242f470c4d8823ef6913c2efdc3231858d04eb7d27f10247d6a1ab8304

Observation 41becd24-0ace-429c-ba4f-8606c5eb6231 · outbound

This paper cites Attention is all you need in speech separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Attention is all you need in speech separation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.165097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.662412Z digest=sha256:a82d6186611e867cb401f32b193f0e90937f0e2995bfe4c34e371313293a09be

Observation bdb63dc3-bd02-477b-beed-6ff8d0783781 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Beyond Speaker Identity: Text Guided Target Speech Extraction SpeechBrain: A General-Purpose Speech Toolkit

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.667684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.667684Z digest=sha256:ec3022493184613c0400226b33172f44213e74f2310d34ce1fd76f8b62a97759

Observation f3eed92a-141a-444f-925f-a158c500a3a3 · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.147416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.675548Z digest=sha256:981da928e5a0aebf658bbf289110aaf7ef8b10993a107b134eb332088c824106

Observation 17418e5a-10de-456d-8c14-58a147cb9a53 · outbound

This paper cites Optimization of speaker extraction neural network with magnitude and temporal spectrum approximation loss,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Optimization of speaker extraction neural network with magnitude and temporal spectrum approximation loss,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.124753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.681693Z digest=sha256:a21d16d92ae1d99bea9f6362f7b3f53afd41df705078efb8315f8aa640cdc8aa

Observation 2335e1c5-3911-4214-a303-5b7847321c5f · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Beyond Speaker Identity: Text Guided Target Speech Extraction LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.691235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.691235Z digest=sha256:551c4f17f1fa87a8044c6743a81cc5ccd03804825a08868aed4cc5761feb2828

Observation 80c66340-71a3-460d-93f6-0d24113b728a · outbound

This paper cites Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech sepa- ration,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech sepa- ration,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.103312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.703745Z digest=sha256:e58454093342f7fa20f50ba5d18931c9cc60fad9eb4e9177ecccab051e7164c2

Observation 4bb5b753-248c-4144-a981-612c9e84d0d8 · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.060804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.709583Z digest=sha256:90906ea9f6164213d8cb5b2f91da61c709029a6f7f2e2782448b72c6804bb461

Observation 540a7d0f-2926-478c-821e-9abe2cbc5e3f · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Beyond Speaker Identity: Text Guided Target Speech Extraction X-vectors: Robust dnn embeddings for speaker recognition,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.715163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.715163Z digest=sha256:39a9980963d8193452e61e0f83077a9585bc5a07d1b7929a525a4a830d19a850

Observation 06270b51-d33d-43fc-8a96-84f584383de9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Beyond Speaker Identity: Text Guided Target Speech Extraction BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:00.721189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:00.721189Z digest=sha256:fec924f71f24ec0544e7a5513de430925bf3d80cc33097c9bd062ce5b046e4f1

Observation 8491beed-4259-4cc3-a58f-5bb734d19bae · outbound

This paper cites Semi- supervised time domain target speaker extraction with attention,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Semi- supervised time domain target speaker extraction with attention,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.032498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.726748Z digest=sha256:f34b49c15ac40284acfc7c7cb32632befc2a8d954bfde75bfac85f0af413d9f4

Observation bb48f634-17e7-43d2-ae8f-024f92d5d5b3 · outbound

This paper cites Performance measure- ment in blind audio source separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Performance measure- ment in blind audio source separation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:01.011895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.731405Z digest=sha256:323dcd8fe285ce1a732fe016b3a3dfde1d779607fca2b64b484b5eb412e9a628

Observation 1ef555f1-f253-4f8b-a894-1524efce4221 · outbound

This paper cites Wavesplit: End-to-end speech sep- aration by speaker clustering,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Wavesplit: End-to-end speech sep- aration by speaker clustering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.985082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.737069Z digest=sha256:19ddfed8e74b2f4715c4c05023688577c835fa8fad05c2fede2231b8a1322aba

Observation 44f33c04-f025-45f8-bbd4-e5dbf6544a93 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.968610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.742812Z digest=sha256:b209a49bc4d980295bf4ffbe0d5dff017130d66189a8af3e7e327de3e34ca769

Observation c4e21c3a-685c-4476-af3d-6a1e468f98e7 · outbound

This paper cites Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.952605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.749916Z digest=sha256:3ed8dd1bfbc0586795dd666c5df5fc30228aa3afb41cd523a175652412533c3a

Observation 287ae483-999f-4ee8-ae62-19339118296c · outbound

This paper cites Self-supervised disentangled representation learning for robust target speech extraction,.

Beyond Speaker Identity: Text Guided Target Speech Extraction Self-supervised disentangled representation learning for robust target speech extraction,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:14:00.939537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:14:00.754527Z digest=sha256:13e283bb4fdc953a949df33423c2427ae0b81f738679d8231c1e0086a9edbd84

Pith citing papers

No inbound Pith citation observations are available.