Pith. sign in

Paper Citation Record · LEDGER

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data

As of 21 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2412.12512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12512 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:05:25.940672Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:58:02.481574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T14:24:45.592114Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19b250f4-bbc9-49f6-8971-e1667a6e15f7 · outbound

This paper cites Neural target speech extraction: An overview,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Neural target speech extraction: An overview,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.751756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.702675Z digest=sha256:12549e124cc9b856762e01179367b5b0aaf88bbaa038f8deb446bebc2c52103f

Observation be20176c-dc09-4550-a0b7-0238337c6f35 · outbound

This paper cites Speech separation with pretrained frontend to minimize domain mismatch,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Speech separation with pretrained frontend to minimize domain mismatch,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.738211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.707581Z digest=sha256:2d9f4dc088989691ebcf20458dde7cd3736372abbdcc24222fee2456d54230ff

Observation 3961b256-2393-4ed4-9444-dc071ccb6b22 · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Spex: Multi-scale time domain speaker extraction network,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.725395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.711986Z digest=sha256:f728371325fd7f2be4e719dd62c92d3b7f5bd437bdc3880d656a56f4c96569a2

Observation f491bc91-3d51-4ea3-bbaa-75fec09f2183 · outbound

This paper cites Spex+: A complete time domain speaker extraction network,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Spex+: A complete time domain speaker extraction network,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.712160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.716431Z digest=sha256:637e5d13c38383fa77f93c38e979f331acfbdf3046f4d176d8ed4b6a49b75822

Observation 09d2fa82-48e0-443b-934c-266b32585615 · outbound

This paper cites Target speaker verification with se- lective auditory attention for single and multi-talker speech,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Target speaker verification with se- lective auditory attention for single and multi-talker speech,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.698216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.720684Z digest=sha256:d90f5ef910329e58cb1992253fb634f0fef4359cee598e55ea857170eb14e52b

Observation be92ded4-772c-4814-8355-db5d3b23a9e6 · outbound

This paper cites V oxCeleb2: Deep speaker recognition,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data V oxCeleb2: Deep speaker recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.684196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.725194Z digest=sha256:6f3b77cdfb22fc53cb920b311dcb21d65adab90f718327e01561669ccacc1b13

Observation 5402311e-22b6-4b1c-a716-8762526b6352 · outbound

This paper cites LibriTTS: A corpus derived from librispeech for text-to- speech,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data LibriTTS: A corpus derived from librispeech for text-to- speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.669640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.730017Z digest=sha256:46ed549d0b7d68b27861e47a329a7b5286cab556539181be19fdeced3975eb49

Observation efc50805-be03-43fb-91f4-e11d6ea6c955 · outbound

This paper cites LibriSpeech: an asr corpus based on public domain audio books,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data LibriSpeech: an asr corpus based on public domain audio books,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.655365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.733998Z digest=sha256:fdb18808a8f11107de583af08a8a245c2168205b5d6d496a80af3311586f8e41

Observation e86aba8e-439f-454c-9f28-f12d5591fe2d · outbound

This paper cites Machine Learning for Synthetic Data Generation: A Review.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Machine Learning for Synthetic Data Generation: A Review

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.737993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.737993Z digest=sha256:055fb438af6717b78b664232a548b0a207d4b5a76e8235f8335484fe004a157e

Observation 8985dcd2-17ca-41a7-a31e-d7633fc9338e · outbound

This paper cites an unresolved cited work.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:05:26.640832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.742359Z digest=sha256:155386b8a853c04f563a17c685f2b4c9927de93155a21d897800732b377beb9c

Observation f699387c-056d-46b3-9eca-131c26961bd9 · outbound

This paper cites What makes good synthetic training data for learning dis- parity and optical flow estimation?.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data What makes good synthetic training data for learning dis- parity and optical flow estimation?

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.626515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.746171Z digest=sha256:52bff51c5c6f011e31b6a001f91a3344a05ac7b8519c78edde5680af1c532ed8

Observation 39b4c628-81e1-48bd-8bbb-7d4820bc7724 · outbound

This paper cites Deep learning-enabled medical computer vision,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Deep learning-enabled medical computer vision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.607128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.750681Z digest=sha256:04b66d6273d9080cb4870ecdceadb228462d524c1c7d977988206b834920ab44

Observation cf96305f-aa52-49fa-9d75-1d3cc51da5f3 · outbound

This paper cites Data augmentation for low- resource neural machine translation,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Data augmentation for low- resource neural machine translation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.593384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.755219Z digest=sha256:f5ed032518fd1a714b308fd760aaf134f53e7f543d9b88adf3d86bb73b1b7d41

Observation 42e86da4-6763-4692-8cee-6ba6485242e4 · outbound

This paper cites A survey of data augmentation approaches for nlp,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data A survey of data augmentation approaches for nlp,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.579956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.759531Z digest=sha256:80da774d8e24af87dc2271ab4a62dedfa9b17f5ef6ec5421b6d12753885d4bf0

Observation 1012f1fd-2c6e-4cb2-8a2b-4440191e488d · outbound

This paper cites A Survey on Data Synthesis and Augmentation for Large Language Models.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data A Survey on Data Synthesis and Augmentation for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.763909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.763909Z digest=sha256:aa99bc73cf2527fb8f5cf709e8acd5a5161cc960b5b620f52b33373ac761f205

Observation c71a945b-80a6-4d9d-9a53-c52991619646 · outbound

This paper cites SYNT++: Utilizing imperfect synthetic data to improve speech recognition,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data SYNT++: Utilizing imperfect synthetic data to improve speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.566353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.768611Z digest=sha256:a448f140edfc58964c99c3a4907a5b9708f0d9fa995063da362fab95ec882653

Observation a8b35c9a-e69a-4c40-bde6-33bf290bf0c3 · outbound

This paper cites SynthASR: Unlocking synthetic data for speech recogni- tion,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data SynthASR: Unlocking synthetic data for speech recogni- tion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.551974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.773119Z digest=sha256:6bd6a31531bab50dab0bce65c4272031a30a8588e48eed71f4e621dfdaf8de64

Observation f20d2363-edaa-44b1-9dae-93e634e7a4b3 · outbound

This paper cites Effective data augmentation methods for neural text-to-speech systems,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Effective data augmentation methods for neural text-to-speech systems,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.536732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.777524Z digest=sha256:2212bc50e211ebcd1ad9c3c0c3d3daa30dfe24df5fb3ebc430e260c124e651ef

Observation 4f623136-4b83-4772-98ec-083d16a362ef · outbound

This paper cites TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.520665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.781952Z digest=sha256:f5092caeac710e4d57e9167315777c56bc3fc91c6b94a742ab04732330490b94

Observation 3e0fbc96-5f59-4500-a78d-89731dd42161 · outbound

This paper cites Speaker augmentation for low resource speech recognition,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Speaker augmentation for low resource speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.505299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.786348Z digest=sha256:46af8761fa81d5265ea5b75b1a942cfc077f78ea821b806cbf71df3a8087bb0f

Observation 871c20cc-f190-43e9-b436-ab8a5947669c · outbound

This paper cites Overcoming data scarcity in speaker identification: Dataset augmentation with synthetic mfccs via character-level rnn,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Overcoming data scarcity in speaker identification: Dataset augmentation with synthetic mfccs via character-level rnn,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.489895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.790558Z digest=sha256:a87f3e7cda7ad7193351f8ce8c5b939f18af5c64b0bab507d745617c8a7c2de8

Observation eac30b13-f307-4f99-b3c5-6717e31c8146 · outbound

This paper cites Target speaker extraction with curriculum learning,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Target speaker extraction with curriculum learning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.474379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.795645Z digest=sha256:80f069b3596de60486afe85c8b783387af34156a64200cbd50496231d07a368b

Observation c5f08fdc-bb07-4df0-8766-8e36006b7d81 · outbound

This paper cites Curriculum learning: A survey,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Curriculum learning: A survey,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.800072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.800072Z digest=sha256:54b331e3fcac744cac69de87d999f153ce77e9fd40838d4faef468b85d0eed13

Observation 1617ea84-6a51-4321-97ab-468889d7cd89 · outbound

This paper cites Improving curriculum learning for target speaker extraction with synthetic speakers.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Improving curriculum learning for target speaker extraction with synthetic speakers

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:05:26.063213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.803945Z digest=sha256:acba9c654c4992fdf11ddd3a81cc9e5233aac65c154e789573b153812fa99e80

Observation 9316f2b1-420c-4572-aab6-6dd903560241 · outbound

This paper cites SALT: Distinguish- able speaker anonymization through latent space transformation,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data SALT: Distinguish- able speaker anonymization through latent space transformation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.450059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.808186Z digest=sha256:bc33d565adf849b49f606f0fd873a9739f4d984b7f2fc0383f5b8927de66fb09

Observation 804a8539-7d82-43dc-8fa8-3bfdbb3591e6 · outbound

This paper cites SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.435098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.812182Z digest=sha256:7af42dc4ee8448148b09e7fd9b680a2f8ec310facfbb34cbc2239fe2660cd5b4

Observation f6bf5473-da86-4929-a7b5-b5e6f429f65f · outbound

This paper cites V oiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data V oiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.420690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.816184Z digest=sha256:2cb7b045be2cc9974c497718e78ac5bc465ef031b6363b2152150a55c4e1c455

Observation d2ea3ce5-c068-443c-980b-9d92cda77b16 · outbound

This paper cites Deep neural networks for small footprint text-dependent speaker verification,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Deep neural networks for small footprint text-dependent speaker verification,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.405758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.820412Z digest=sha256:de11ae6a19ada7521bf31c1d35686874c2289fc32e46d9ef82f0d22dd525ac9b

Observation a3b5fd32-94fd-4060-9e11-bb5a4d452b45 · outbound

This paper cites Conformer: Convolution- JOURNAL OF LATEX CLASS FILES, VOL. XX, NO. X, XXX 2022 12 augmented transformer for speech recognition,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Conformer: Convolution- JOURNAL OF LATEX CLASS FILES, VOL. XX, NO. X, XXX 2022 12 augmented transformer for speech recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.390600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.824385Z digest=sha256:29320d49063976c8af33fba978e53044c962edc54be10faa66e94f2e9a2abd87

Observation 99968acd-29bc-4dae-9a59-acc260fb9d00 · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.375718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.828421Z digest=sha256:be94318d81e94de98dfa311e7d1227c782ba7e152006b0095a6c20030e76a3a3

Observation 19c705e8-7ddb-458f-91f3-89305faf7973 · outbound

This paper cites Complex ratio masking for monaural speech separation,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Complex ratio masking for monaural speech separation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.361821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.832434Z digest=sha256:b2baa9064fad7e3a5ba057c66c34d389b69cb6046307633413f3be249c17bb61

Observation 288dd581-62b1-4ac1-8f38-77c3b136e5bd · outbound

This paper cites CSR-I (WSJ0) Complete,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data CSR-I (WSJ0) Complete,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.348431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.836392Z digest=sha256:0a343f5dd7fea88fa0faf3127ad3cb01cff40fc263660ad1b16394c13a3de76c

Observation da9cba9c-7934-4220-bc05-bfe2a49c4a1f · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.840587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.840587Z digest=sha256:0efc0439185b7747cb2107f37fb358567cf6f5813208157a10af9c8e7a3f8c7e

Observation 2635c539-2ee4-45c7-aa1f-f127e5be13ad · outbound

This paper cites Data augmentation for speech separation,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Data augmentation for speech separation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.333658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.844693Z digest=sha256:e12ed9a26db10277bbe4bfc7e2591851e89a47f4f1b607c1e65e1aefd6b8d97a

Observation d2e71f56-a5fe-4421-9d43-0ffb6f9d0b29 · outbound

This paper cites Employing real training data for deep noise suppression,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Employing real training data for deep noise suppression,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.318647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.848730Z digest=sha256:a669273738a05a0caa3dc0b3a03b91878a30858fe338b0f9497ef2980d87f9bf

Observation 527d4c46-f7f3-4386-ac5c-57653f5451c3 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)—a new method for speech quality assessment of telephone networks and codecs,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Perceptual evaluation of speech quality (pesq)—a new method for speech quality assessment of telephone networks and codecs,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.304912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.852847Z digest=sha256:9118015539fc3e659e84a7f8f8ddc84b953a4d1e1b650160a0d3cd62da8bd4c7

Observation 854c88b4-ee8a-4a76-a158-6de42a5078fc · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.289956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.856726Z digest=sha256:60fd4823f6bedf3b8ab49e90e5fcf2f7c27f49833634202066046df660daf1b6

Observation 5e94b42c-f0ea-454e-979c-0521f4e787d6 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.865952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.865952Z digest=sha256:221e815b8a8e000dc821fae5b8790b5bdf2f6e54d8d7d2437e45616b334054e2

Observation 2ebe90da-d97e-4d12-a128-be2ba4309274 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.870697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.870697Z digest=sha256:e6c63460b29884e25a1de303fd04d61866d2a5ae5fb845bbfcaaa4bb30d00adc

Observation cf41df21-2d85-4d03-ad71-8eb0d50cda16 · outbound

This paper cites Synvox2: Towards a privacy-friendly V oxCeleb2 dataset,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Synvox2: Towards a privacy-friendly V oxCeleb2 dataset,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.259664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.875729Z digest=sha256:d5867fa0774b24d733c500da8f21fac679ab7c76e54c643db89f9490fa89891d

Observation 52217345-bdb1-499b-955b-a8c234d5b174 · outbound

This paper cites WavLM: Large-scale self-supervised pre- training for full stack speech processing,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data WavLM: Large-scale self-supervised pre- training for full stack speech processing,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.880173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.880173Z digest=sha256:eab6213354f25bd3ab0676aefc2329df43b2faa5a36c0d28d0977151db1c0aa9

Observation 096296db-06f7-4ded-94e9-73174d971e7e · outbound

This paper cites Nearest neighbor pattern classification,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Nearest neighbor pattern classification,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.884477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.884477Z digest=sha256:3bf44a504f0247839c451d490cdf43b5d8f0abc43523394ab219ff0d9982f1e0

Observation cdbb77f7-6cb8-436f-bb4e-a134358a7efd · outbound

This paper cites HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.226267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.889434Z digest=sha256:a6de87cc70de4e6ff82530b315205d04e354fe101a3412c7498bdadc1290d031

Observation e892e776-5dd9-47b3-8235-e48426ad68c9 · outbound

This paper cites Speaker anonymization using orthogonal householder neural network,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Speaker anonymization using orthogonal householder neural network,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.211189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.894042Z digest=sha256:d76745fbfd070b51271671b2f1f850d7b2ca038d95a6426c0fe2aac905c7959c

Observation 4c8eb7fe-090e-4898-b47f-7fde47838429 · outbound

This paper cites Yet another algorithm for pitch tracking (Y AAPT),.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Yet another algorithm for pitch tracking (Y AAPT),

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.196025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.898539Z digest=sha256:65882af65adb3c61e05d7ba45637394c3507c9ec6be06edc1cc4b872f4a6ed58

Observation 7e4ed4c6-8c3e-483b-b8f3-2b0f1bad3473 · outbound

This paper cites HuBERT: self-supervised speech representation learning by masked prediction of hidden units,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data HuBERT: self-supervised speech representation learning by masked prediction of hidden units,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.181280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.902947Z digest=sha256:47e99cf75682d58d87e588022cc8b351c71e536d414ca8f79d85fd2e86ee88e4

Observation d164a870-fcd7-4d8e-9aeb-6b2f8413bd49 · outbound

This paper cites Attentive statistics pooling for deep speaker embedding,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Attentive statistics pooling for deep speaker embedding,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.907289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.907289Z digest=sha256:bac9743613822ebc22676742c1a89864c831bafb1fd5ad894d0526c014373bfe

Observation 0bfad4d8-744c-4803-b2a0-e44cd4113727 · outbound

This paper cites Additive margin softmax for face verification,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Additive margin softmax for face verification,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.911910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.911910Z digest=sha256:59908935c9cccae8a290b40a413ea0b16009771030c9cc00b72d38d9eb5b3ad3

Observation 8224650f-5820-4953-82e6-ab950dfccf9e · outbound

This paper cites Open-Source Conversational AI with SpeechBrain 1.0.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Open-Source Conversational AI with SpeechBrain 1.0

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.916256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.916256Z digest=sha256:fbc9a5c7ec9d942a7e24f48d268933c5cde4d739d1fe6c24b99e9f1a9c9221e4

Observation c75973e2-3d35-4d1b-b005-75f4b6b5d6a0 · outbound

This paper cites V oxCeleb: a large- scale speaker identification dataset,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data V oxCeleb: a large- scale speaker identification dataset,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.148483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.921052Z digest=sha256:e7a265ccff0f6e3b271774f236350875abc8eeaeb647e41b66417649a896cbbf

Observation fb754191-333a-4408-9b30-a31a132ed685 · outbound

This paper cites CN-Celeb: multi-genre speaker recognition,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data CN-Celeb: multi-genre speaker recognition,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.134549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.926135Z digest=sha256:b73f3a8b5284348e5e92961fbcf34d0f54e63755266d23941a06c4438ab54419

Observation 44c5f1b1-47d2-47bc-bed1-0978a1be1f0d · outbound

This paper cites TorchMetrics - measuring reproducibility in pytorch,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data TorchMetrics - measuring reproducibility in pytorch,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.120482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.931173Z digest=sha256:e6aebf4509c15fc3027decb235617103ada1ee19fbfec7bf1064d687a3051b38

Observation 31b2afeb-3331-40cb-992a-ee4dd3bd7a85 · outbound

This paper cites ICASSP 2021 deep noise suppression challenge,.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data ICASSP 2021 deep noise suppression challenge,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.106950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.936181Z digest=sha256:297374a945e0f96c249c094d18d858686597c3c6e62e407501d59b1575e06440

Observation 57f5eb54-bea9-4e99-a9ad-6897ee97b23f · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.940672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.940672Z digest=sha256:f381c6cb452ba0b2e5e195c1940849af7bff7dc7c02de7f2d9fe2d28315e7cf1

Observation b0f1d672-b665-4b7c-bfdb-d58aeb88c18f · outbound

This paper cites 22 605–22 623.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data 22 605–22 623

Reference 235

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:05:26.273992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T14:05:25.861418Z digest=sha256:17d46b8a54945a4aec4b02e9b38e3b071a138db2c48587fcb8ef4037d3bc8faa

Pith citing papers

Observation ff96ee8b-2756-4927-82cc-6de689893135 · inbound

Interpolating Speaker Identities in Embedding Space for Data Expansion cites this paper.

Interpolating Speaker Identities in Embedding Space for Data Expansion Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:58:02.481574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:58:02.481574Z digest=sha256:10093d0d0a78de717cb0deac83ffaada4e72d499b45d3f2ea2d967e1c2dfca09

Observation 8503687a-c68d-471e-8c66-def7496542be · inbound

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models cites this paper.

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.594078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T05:26:46.541637Z digest=sha256:5f02a5db787e8d3226fde04303301df6cd441ad495d5e7c97f8f5b13430d66be