Pith. sign in

Paper Citation Record · LEDGER

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:1908.04737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04737 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:38:44.464676Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:38:43.983210Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T13:38:44.647594Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact2
  • verified fuzzy31
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 651a7a66-fbd6-481f-9e72-b12532fa685c · outbound

This paper cites Overlapped speech – well known in a more general context as the cocktail party problem – remains, however, to be a largely unsolved problem.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Overlapped speech – well known in a more general context as the cocktail party problem – remains, however, to be a largely unsolved problem

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:46.256787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.969214Z digest=sha256:0c54d0daea79ee807e964dc2a36870bc9eeee584a47a0811bfd931ca6f06ff1a

Observation 1679087c-a2fe-45f6-899e-da05fe52473a · outbound

This paper cites End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T13:38:44.652679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.983210Z digest=sha256:33d2bd586e2fea27cb48d3606d6b94821f4ef2e9db65ea8527c46ee7c8843415

Observation eb5e50e7-9421-40c9-ae83-1019dba0aa1b · outbound

This paper cites Datasets We evaluate our models on the widely used mixed speech datasets wsj0-2mix and wsj0-3mix [9, 10].

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Datasets We evaluate our models on the widely used mixed speech datasets wsj0-2mix and wsj0-3mix [9, 10]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:46.162959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.987681Z digest=sha256:3fd08fd6bd50a9cc00ef321d606513e8c22efb0de7aa67c887ffcb7d7e43d4ba

Observation c3b48b80-dcf2-4563-bc6d-ed4d7d4e1351 · outbound

This paper cites Speaker embeddings inclusion strategies The first set of experiments aims to determine the best strat- egy for inclusion of speaker embeddings in the model’s in- put.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Speaker embeddings inclusion strategies The first set of experiments aims to determine the best strat- egy for inclusion of speaker embeddings in the model’s in- put

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-14T13:38:46.098133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.021457Z digest=sha256:c9dc6f182f06252329e5194b1d92f9e75db40ef68862afb12198f4225a92dda6

Observation ce66aa74-d728-418d-80f4-049bc957192a · outbound

This paper cites an unresolved cited work.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:38:46.008014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.065655Z digest=sha256:c36bd5bc7d3c66fd61df37e90a42d7fe800b8441637b7344458e7daf0b7b780a

Observation 3fa5607e-9c79-466f-acb3-474e4d332e38 · outbound

This paper cites Single-channel speech sepa- ration using sparse non-negative matrix factorization,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Single-channel speech sepa- ration using sparse non-negative matrix factorization,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.757461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.135876Z digest=sha256:57c2f1470561fb0b893a4293eece7235b63fd1b601400b654580d0ab06320a6b

Observation 1d7798b8-b707-41d0-b106-b4bb392411ff · outbound

This paper cites Similarly to speech recognition, speech separation methods have also made major progress with the help of deep learning.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Similarly to speech recognition, speech separation methods have also made major progress with the help of deep learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:46.243718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.973977Z digest=sha256:4e8dcc4dc818952a5fe4054a51428cc2b61caea45c077d3da1a3c391b70e6cc1

Observation 80b02dcd-5bfe-4388-800c-ba63f749f3ff · outbound

This paper cites Learning spectral clustering, with application to speech separation,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Learning spectral clustering, with application to speech separation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.144009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.144009Z digest=sha256:94eb1eaacf46bcb9b13f31f5576cc9138adcc6b811722465cb8b182aa9e49a94

Observation 90517952-21c9-4dda-89e1-2432d8ff8d9b · outbound

This paper cites Deep Neural Networks for Acoustic Modeling in Speech Recognition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Deep Neural Networks for Acoustic Modeling in Speech Recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.956377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.102873Z digest=sha256:fbc21e021d3c0beb67f29114fc7a3a77732726c2a917a0e57878e2a04e1d74e4

Observation 50045803-b5ae-4dfc-b450-03dd4ea5bcd4 · outbound

This paper cites Context-Dependent Pre-trained Deep Neural Networks for Large V ocabulary Speech Recognition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Context-Dependent Pre-trained Deep Neural Networks for Large V ocabulary Speech Recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.942351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.117866Z digest=sha256:f5e71eec0e011c416c9c5337ca35c70d113a9967cd8ba1c219529e1a0c86a5b2

Observation 32128274-1284-4060-bd96-ba02c6d14182 · outbound

This paper cites The Microsoft 2016 Conversational Speech Recognition System,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning The Microsoft 2016 Conversational Speech Recognition System,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.926674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.122624Z digest=sha256:c10f0265856b371880f4ec2b86f034c7a92c101dfef5f3c78ba03d9324119fc7

Observation 14693f19-1e94-4cbf-ac15-eb8d6047ba35 · outbound

This paper cites We evaluate our proposed framework on overlapped speech datasets with two and three overlapped speakers, within and across set- tings.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning We evaluate our proposed framework on overlapped speech datasets with two and three overlapped speakers, within and across set- tings

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:46.178694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.979371Z digest=sha256:19695840467901c38b0ca23254d3076e7ea14eb5fac622e610bfd4674a003d96

Observation 2cfcf485-9698-4ef0-bb1a-83ddd39297de · outbound

This paper cites Purely Sequence-Trained Neural Networks for ASR Based on Lattice-Free MMI,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Purely Sequence-Trained Neural Networks for ASR Based on Lattice-Free MMI,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.822765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.127021Z digest=sha256:5cc2dd6123a0d91cc28a0093d3a57798211ffe1e90c17b354b7272bea57d7ca1

Observation 02cf7ebe-f578-4965-bddf-7b003df9831d · outbound

This paper cites Wang and G.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Wang and G

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.131178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.131178Z digest=sha256:627754ae90e047b88d3f0a15a77dbd6859cf2d6f0fa7e4c90674c97ecebdea00

Observation 343adcf4-5d1d-4057-bde8-52943c1d8180 · outbound

This paper cites Super-human multi-talker speech recognition: A graphical mod- eling approach,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Super-human multi-talker speech recognition: A graphical mod- eling approach,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.723549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.140284Z digest=sha256:41a58500a0c2a2970e5d9c446a1e9744f2b871cb76ba2401995253813795da7c

Observation 76f58950-e03b-4a51-93e3-d9256fa023d7 · outbound

This paper cites Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.243730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.243730Z digest=sha256:907b60e7df2ed9cdf1b21f1f820e42272ff1a8aa0e3fa458db6e749bd6559ec4

Observation 9291abe4-8011-4a7f-90f0-7bb8606d21b2 · outbound

This paper cites Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.630215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.147759Z digest=sha256:8dd267444b051bca6dcf9221c660acda25465a78ecd3a1e81b0f8667d06bf94d

Observation 48258b02-6caa-4026-8f76-1cd99ad347d6 · outbound

This paper cites Single-Channel Multi-Speaker Separation Using Deep Cluster- ing,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Single-Channel Multi-Speaker Separation Using Deep Cluster- ing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.613430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.151655Z digest=sha256:48e7cd80029d51bef944c5e621437488e9ee359edc43dc994d259d6111e781eb

Observation 3e709503-00f4-4fb0-9cae-dad1cbb09baf · outbound

This paper cites Alternative Objective Functions for Deep Clustering,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Alternative Objective Functions for Deep Clustering,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.547771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.155502Z digest=sha256:fcb59ae3053cdecfdec28abd3f93cb203198651478ec5d90db0a8fdb1a604ac3

Observation 45f9cf6b-e8a3-4aa0-9f84-112abc83ea01 · outbound

This paper cites VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.171139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.171139Z digest=sha256:c2ca9f350189ce11689463a987ecf92b4b2c7846a8873a925a46dec30515714f

Observation ddd76945-b2a4-4759-84cb-d106895b4453 · outbound

This paper cites Deep Speech: Scaling up end-to-end speech recognition.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Deep Speech: Scaling up end-to-end speech recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.200937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.200937Z digest=sha256:b00b2c3dec7db417b320133a79219e817ea7f5195120994f5a2f280ccddb5c92

Observation b5289f2c-c191-4c08-90fd-3884db4a0288 · outbound

This paper cites EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.531701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.234231Z digest=sha256:918798878434b08eb623c15f6621b7f1118629ef30e122be23fd54d4b8daf53e

Observation f2eb5c52-38b1-4e8f-a77d-244942965096 · outbound

This paper cites End-to-end attention-based large vocabulary speech recog- nition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning End-to-end attention-based large vocabulary speech recog- nition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.239254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.239254Z digest=sha256:a7f0efa8762553a433200239fdc7420768ba2717afa0a45ed9e9608f6092818b

Observation 3d55b2c1-1420-40b0-b4cd-1aec4d8a367f · outbound

This paper cites Transfer learning from speaker verification to multispeaker text-to-speech synthesis,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Transfer learning from speaker verification to multispeaker text-to-speech synthesis,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.278173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.278173Z digest=sha256:b571bdc7a6d66eea7903eb9effc43f9d5d0ba60222c8ce8f71ab4dd444b4da04

Observation f08052f1-9b32-4763-90af-b65645c2e978 · outbound

This paper cites Hybrid CTC/attention architecture for end-to-end speech recog- nition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Hybrid CTC/attention architecture for end-to-end speech recog- nition,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.247685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.247685Z digest=sha256:e68cea5753e87354891338bafe71f904ab1b5727901893fd51d34624a8f12d0d

Observation aff3f211-a2ca-4d01-b38e-4a522225d479 · outbound

This paper cites End-to-end multi-speaker speech recognition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning End-to-end multi-speaker speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.433805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.252011Z digest=sha256:4ca35a53ec5d1f6c9210830014fffba63a7290379754319f4ae91716ce736ae1

Observation 51ee53ba-e313-4c8a-8171-0bd8a61dfbbf · outbound

This paper cites A Purely End-to-End System for Multi-speaker Speech Recog- nition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning A Purely End-to-End System for Multi-speaker Speech Recog- nition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.374155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.257226Z digest=sha256:ed63ce2b7736a0c0ffafa18823e07cb0e76a5d8c0860e742edb812ca43f34c89

Observation 32da70e1-b1b3-47f0-ba1d-4a034ef7e7c4 · outbound

This paper cites Progressive joint mod- eling in unsupervised single-channel overlapped speech recogni- tion,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Progressive joint mod- eling in unsupervised single-channel overlapped speech recogni- tion,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.285680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.261153Z digest=sha256:8fb79ed4205def3b8c515e71b1c559493922d81c17013d0bf8de8279bd910142

Observation 85ad0927-995c-4d25-bc29-cdba43df20e1 · outbound

This paper cites Speaker-Aware Neural Network Based Beam- former for Speaker Extraction in Speech Mixtures,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Speaker-Aware Neural Network Based Beam- former for Speaker Extraction in Speech Mixtures,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.227035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.265107Z digest=sha256:118b4489ade8ecdb9959aa7f7db3e76be3a89e9a076f1c8d0b1be8d407ce334a

Observation 4a422157-4728-4c9a-b4cc-3d7ecd6c267c · outbound

This paper cites Deep Extractor Network for Target Speaker Recovery from Sin- gle Channel Speech Mixtures,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Deep Extractor Network for Target Speaker Recovery from Sin- gle Channel Speech Mixtures,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.172901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.269202Z digest=sha256:05de84e8ee1cee336930e36868efa28ba254fd67b919eb9258c40e9a4f58edc0

Observation ef0aa0f8-d6d3-458d-b841-7b7f035c8ad7 · outbound

This paper cites Speaker diarization with LSTM,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Speaker diarization with LSTM,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.110169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.273591Z digest=sha256:79186bc4509cb5cd0d35fcc3ee59f5d14bb579a72262de0223435bea0a8e0a04

Observation 17d809fd-63a3-4a3d-81a1-231379a0e2e3 · outbound

This paper cites VoxCeleb: A Large- Scale Speaker Identification Dataset,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning VoxCeleb: A Large- Scale Speaker Identification Dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.818701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.440189Z digest=sha256:d6b8f80a1040ccc7cc5b57c6fea837faf80d9480446fd2c3a3002b35602724c7

Observation d59d8051-2bae-443b-bdba-b7a353cdfe5a · outbound

This paper cites Multilingual acoustic models using dis- tributed deep neural networks,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Multilingual acoustic models using dis- tributed deep neural networks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.040173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.333123Z digest=sha256:e0dd91c528d1ae61d41bbb28457908d8466ce4d5d5019f00c6dabca317056bad

Observation 667993ef-e747-441b-a793-78a5a8f6be98 · outbound

This paper cites The input features of x-vector extractor are 30-dimensional MFCCs without cepstral truncation with a frame length of 25 ms and shift of 10 ms.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning The input features of x-vector extractor are 30-dimensional MFCCs without cepstral truncation with a frame length of 25 ms and shift of 10 ms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:46.144120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.992657Z digest=sha256:29340924b2660d1f82a826adce923ca64f34979b71c1e286ab98e0bb63c53fc0

Observation 43a8af3a-23de-49a6-972e-5ef8595165f3 · outbound

This paper cites Investigation of transfer learning for ASR using LF-MMI trained neural networks,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Investigation of transfer learning for ASR using LF-MMI trained neural networks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:45.026396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.371895Z digest=sha256:8000909ac66bbe7e47c8614f75958481f64b2349f67cc8c43e69b03bdf60fa93

Observation 89d25c21-6373-4c68-bf68-cc5bee4a5c46 · outbound

This paper cites Lib- rispeech: an ASR corpus based on public domain audio books,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Lib- rispeech: an ASR corpus based on public domain audio books,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.926033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.388157Z digest=sha256:50e96f15ed0911cde21751483c31c1c910ae5fbc6c65998f5b4fe4b0f1b5d1c3

Observation 589ec6d4-28d2-47a6-a474-f06a72dd5d66 · outbound

This paper cites ESP- net: End-to-End Speech Processing Toolkit,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning ESP- net: End-to-End Speech Processing Toolkit,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.874794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.393052Z digest=sha256:37569455222c75c2c67eb1d2c46a32f0ddcc2ea2e21910ae8eadbea9e680e048

Observation dda78843-61c7-4ccc-912f-d631e6d2a375 · outbound

This paper cites The Kaldi speech recognition toolkit,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning The Kaldi speech recognition toolkit,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.864517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.398596Z digest=sha256:2a92db8f64843802d0bd8791400fd8f0f4f2e1a5303d86ec862e4cd0445b5e9b

Observation 866d7ca0-0a36-4bb8-b333-808f1f91cec1 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning ADADELTA: An Adaptive Learning Rate Method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:44.402934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:38:44.402934Z digest=sha256:52f779ac357baeae59be11a2b9c4968ebe1f27ed6f7f256fc3b4c1598b300579

Observation 7eebc545-9de0-4bcf-ad0f-c7f33c2f7108 · outbound

This paper cites End-to-end Speech Recognition with Word-based RNN Language Models.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning End-to-end Speech Recognition with Word-based RNN Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:38:44.584400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.408490Z digest=sha256:a544b5dfdd665ebaa199eb4e6f043699e1868d0549f196ea1febb30ec2ceccb1

Observation 55fecded-c0c2-407f-a1b9-08a4ebf75a24 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning VoxCeleb2: Deep Speaker Recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.731660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.446192Z digest=sha256:2d9fa3724249b5c566cfb2817aa846a82f8495a96b49f2b36731c7e4d7c084c4

Observation 17b0b9ba-0761-475d-a354-e0d94cd488ab · outbound

This paper cites X-vectors: Robust DNN embeddings for speaker recogni- tion,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning X-vectors: Robust DNN embeddings for speaker recogni- tion,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.718804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.450882Z digest=sha256:9964511e6ad20f3e0c9d8e13479f82aa1e2c6572cba20cbb3ada97183b06f3df

Observation 29c5e611-c86a-4eae-be53-ae1808fa4eb0 · outbound

This paper cites The Speak- ers in the Wild (SITW) Speaker Recognition Database,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning The Speak- ers in the Wild (SITW) Speaker Recognition Database,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.676529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.455556Z digest=sha256:9a19598f5cb6250129d74beac3d1326f688cbe4a1805c9c5d73ec8f54ce4fabe

Observation 78a550f5-856e-40f6-9652-85e183c6b9e8 · outbound

This paper cites Single-channel multi-talker speech recognition with permutation invariant training,.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning Single-channel multi-talker speech recognition with permutation invariant training,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:38:44.665588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.459737Z digest=sha256:459878143365a7bcb5640acfe216341a25cc269e0f779a7518957cb17b69d0fc

Observation 76f8073d-ad2f-45ae-aac8-28317a70026e · outbound

This paper cites End-to-End Monaural Multi-speaker ASR System without Pretraining.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning End-to-End Monaural Multi-speaker ASR System without Pretraining

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:38:44.564274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:44.464676Z digest=sha256:f4ec14d015ec685f9a8b6694db3b2e171cf90305f454b855414d96906ca00c9e

Pith citing papers

Observation 1679087c-a2fe-45f6-899e-da05fe52473a · inbound

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning cites this paper.

End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T13:38:44.652679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-14T13:38:43.983210Z digest=sha256:33d2bd586e2fea27cb48d3606d6b94821f4ef2e9db65ea8527c46ee7c8843415