Pith. sign in

Paper Citation Record · LEDGER

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

As of 17 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2505.10975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10975 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:03:18.851116Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:46.748745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:58:03.064307Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact7
  • verified fuzzy1
  • unresolved33
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc2a4700-80b1-4cb0-93fa-5a3973f849a2 · outbound

This paper cites an unresolved cited work.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Unresolved cited work

Reference 2

Resolution
verified exact
doi, observed 2026-08-15T21:03:19.357349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:17.927997Z digest=sha256:1b839de3b0ee1f7054ebc12db7c1f64ab462ff967adb4607a93bf2d09f57d43a

Observation 87b5533e-fdc2-4835-be0f-1e716c748a59 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:17.992758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:17.992758Z digest=sha256:f69a460d1e27ec488d6a733617bfbcf5de31b72f47ebce44e16d9e31c0dae0ce

Observation 014b6ad6-e672-4968-8ba0-a7f9c1d65004 · outbound

This paper cites Continuous speech separation: dataset and analysis.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Continuous speech separation: dataset and analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:17.996199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:17.996199Z digest=sha256:bc0b2d516d175aa5a236c96d00b2a73a15931c3fa599e381c2dfd0750b68eb7c

Observation 80573499-0b52-463e-9876-63d12850ee2c · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:17.999940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:17.999940Z digest=sha256:3e40568c0b11d8450029f7c196e5589c4d90e8b01e115e063ec52b9a6e17a91c

Observation e0acadb9-521d-4ae5-ad88-dff1dd8f2fad · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.003997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.003997Z digest=sha256:3222da00650800cd652092574dabc1248091ab4464ae7395afaff80f77e64377

Observation 43644128-e3fd-42f2-83bc-2ddc70e9d148 · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.019220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.019220Z digest=sha256:1d0279efc2aa13b1c6c5240ad3b124ba400393415c02c09aca5a92254881efdd

Observation 9d7766db-b8b1-40dc-a37f-678cbeb20bed · outbound

This paper cites Computer Speech & Language 99 (2026) 101925 16 X.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Computer Speech & Language 99 (2026) 101925 16 X

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:03:20.839177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.134894Z digest=sha256:d020e135b35ce3072a89f5afbe4380ec1f8d10e93aa80c4b3c8f79b87264ba35

Observation 512e7606-df37-4d9f-a56c-bfb8a9125933 · outbound

This paper cites Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:03:19.263997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.142286Z digest=sha256:70ca892d7be2d2baf6bf3969e86c873e34cb282ca79160884292e07eacaf2d1c

Observation 4d685cac-b9d6-4e3a-85a8-51969f0a7926 · outbound

This paper cites Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:03:19.205362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.157660Z digest=sha256:c699d2abd1d29e65f5925f4b32b3296d18f853c711c4e6cb7962cae34208eceb

Observation 0838d9e9-e9d3-495b-8bb7-63442144cc61 · outbound

This paper cites Sparks of Large Audio Models: A Survey and Outlook.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Sparks of Large Audio Models: A Survey and Outlook

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.204759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.204759Z digest=sha256:25924cd46c5342da830e7eda78a388c73bb7c3e6b2187350a2d6cd16ec57ad0a

Observation b3b85933-601a-4297-8be2-124f1c694388 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.274712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.274712Z digest=sha256:a3ca87f1b1efcd771cf93f04a85bf824dd2a9cd6a66dfc8189353b23655da2b0

Observation 9c7dd01a-c8c4-4cae-ac02-e4472df349f3 · outbound

This paper cites In: ISCSLP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ISCSLP

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.311922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.311922Z digest=sha256:0db2e6d9fe07034573ad1e0b7ec074b2fe96a38dcb77a8dec68a7e95824caa98

Observation 265aac97-bd2d-4663-89ef-00fa60e374ea · outbound

This paper cites Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.316167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.316167Z digest=sha256:d1190a30ca11d249b034a247b7ba75ecfed00d065ed8423741d652bd8fbdc6d7

Observation a40a6551-a0f1-452a-a58c-5dba264cd584 · outbound

This paper cites 1838–1842.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio 1838–1842

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:03:18.327670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.327670Z digest=sha256:670374dab48c70b90b1455e86d16a8d20551967cf1d745c074ad0cfc1e8a4270

Observation 267c89a6-9dc7-4fd4-8ab2-251205ebb4ec · outbound

This paper cites Multimedia Tools Appl.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Multimedia Tools Appl

Reference 29

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:03:18.330539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.330539Z digest=sha256:6e9a493ced6576893da0f9c4e19d2e145b38c3a7b648e169ddd3517c319658a8

Observation 26ef89c8-db3a-4e53-862b-556e6a59fc91 · outbound

This paper cites Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.334013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.334013Z digest=sha256:efa0b210d16e742881a1f3fb2ed001f4fe1db9ac472a75604e965c2de54052b3

Observation 9c3e268b-e0d4-4878-a0d0-76b88a2e12b2 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.337916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.337916Z digest=sha256:fb51dd0d9b633982083287d722b8fc2f1406f10db80b1453eddb0f123201d7f2

Observation ffca1ec4-d02b-49ac-986c-ec99ae55d9b0 · outbound

This paper cites Alignment-Free Training for Transducer-based Multi-Talker ASR.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Alignment-Free Training for Transducer-based Multi-Talker ASR

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.340664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.340664Z digest=sha256:cc79fdb58dff7431040af34bf3994d53fdb30c0e1537b6722973220ca3b5353d

Observation 954e3bf7-3205-46e4-ad4b-b94af12df57a · outbound

This paper cites IEEE/ACM Trans.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio IEEE/ACM Trans

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.347670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.347670Z digest=sha256:fa3d959173dbf2c35665ea345050aa5e5cf7d4d4c3da82dca6c4f6accdc194d0

Observation cbd1a595-8add-4be5-bf26-775e5ff3ed5d · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio SpeechBrain: A General-Purpose Speech Toolkit

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.354196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.354196Z digest=sha256:45afce0ba5dd8e8c3eb52416213d10c35d15b2f5408a979bb577cee55ac8dd13

Observation d326b034-c01e-4348-a35d-c75e45851e84 · outbound

This paper cites Cascaded encoders for fine-tuning ASR models on overlapped speech.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Cascaded encoders for fine-tuning ASR models on overlapped speech

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:03:19.893147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.445989Z digest=sha256:1e8eec09b57fd26f80f0a64df83323f8c87c6576b13e41261123fc80e17a2cb7

Observation 85ebe427-1cab-46c5-977e-64fd68065f0b · outbound

This paper cites In: Proceedings of the Annual Meeting of the Assoc.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: Proceedings of the Annual Meeting of the Assoc

Reference 38

Resolution
verified exact
doi, observed 2026-08-15T21:03:18.998390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.549627Z digest=sha256:05f8bd790eb36da065934029d7d9bc84e1da38f7469989135e0af62efb4af7f9

Observation 3dcddab2-d1d2-4c40-b86b-8f9caa4d57d7 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:03:18.634500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.634500Z digest=sha256:01a51eabf50f3b4528bac8d6e481d2f605173e580ee764ca4f0da83a9baf738a

Observation 6c63f12b-4bf5-4957-9ad3-fa3e0ea7f5dc · outbound

This paper cites In: Interspeech.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: Interspeech

Reference 40

Resolution
verified exact
doi, observed 2026-08-15T21:03:18.986783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.637927Z digest=sha256:89ad2e2258f26a3d052724be250895134ce5fa37a4c1653b2a244058ff80b54e

Observation 1006d463-c766-47fd-891a-170a5bab1c34 · outbound

This paper cites Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.641177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.641177Z digest=sha256:54b61e049a502e082a3f4c323d380dc70517191bb995c1429d5fe05c3e382f08

Observation a0e49438-780b-4709-b8c6-48f82b74b6f3 · outbound

This paper cites Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.644641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.644641Z digest=sha256:56f3099ec3637cf629bcbc0dd5c7e9b3f84637a31231e7a2ff1575a5ed4fa065

Observation 7080dc02-17cb-4def-a954-e8273cd65e67 · outbound

This paper cites Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.648277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.648277Z digest=sha256:7cb36c5464524f8b95c11d8b5680233e93dea464e2d2bcf175e377e0122ec70d

Observation 2dc3d023-0840-4723-800d-2a9743f423ba · outbound

This paper cites Attention Is All You Need.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Attention Is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.652393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.652393Z digest=sha256:0aa51b3887185b335d5beaba319256b8678a8b7074f041580a21503877f7543e

Observation e68ba598-a4bc-440f-baf6-d95f5d7ec4e2 · outbound

This paper cites MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.656284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.656284Z digest=sha256:f74a15e9fab9b52b1596e6d50fcbaa1b378326505c21ccd788f96bea55079726

Observation c29aab31-1673-409b-9eec-01d2830e110e · outbound

This paper cites http://dx.doi.org/10.1109/SLT61566.2024.10832215.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio http://dx.doi.org/10.1109/SLT61566.2024.10832215

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.660875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.660875Z digest=sha256:9da18a4f995b44a58069675dcb7d44899a82fb6d7fdbe9fdaedc1ff26b43e4ad

Observation 82b913ae-81a6-47f3-a4b0-5db9691a1aa5 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.664293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.664293Z digest=sha256:4a139d5c30a7ed3596ef3e732b038887f1eaaaabd4a772a43338d81a5dbc672a

Observation 8b3f3ee2-5725-4406-865d-bc8b7705d8a0 · outbound

This paper cites Watanabe, S., Mandel, M.I., Barker, J., Vincent, E.,.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Watanabe, S., Mandel, M.I., Barker, J., Vincent, E.,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.667875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.667875Z digest=sha256:c395698c5ec4b08d943ac12a043441e7656def6bfe977a8b022c76acd1817796

Observation 3b86dc09-227b-4df0-a340-b871f9790bc9 · outbound

This paper cites CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.670824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.670824Z digest=sha256:e08417b2c9b449526792eb695d62e22573bc633667af7ea49927070c89aa1496

Observation 18eb5e2d-f94d-4048-b943-360f17f87371 · outbound

This paper cites 3021–3025.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio 3021–3025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.674604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.674604Z digest=sha256:9e41dc5a36ef51e1b87f9b4db4cf1fb363fcb1aacdd9c112dbe9ffbc4ba11fd4

Observation 2ade5721-5335-403b-8532-e10988ac9fe5 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.678121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.678121Z digest=sha256:fac735983fedfecef78d0f2f8e165060343c95ad7c8bf5db0048871dbc15ab05

Observation 4d4621ae-c651-45da-88b3-eebc141cc684 · outbound

This paper cites Zhang, W., Chang, X., Qian, Y., Watanabe, S.,.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Zhang, W., Chang, X., Qian, Y., Watanabe, S.,

Reference 52

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:03:20.827342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.721030Z digest=sha256:329f9de144c9d7253159aacf963a250631d8f78656ab1dc71f9ae0f2e1dab3b2

Observation 3ba6ee1b-f069-4025-9157-ccf4e8706791 · outbound

This paper cites IEEE/ACM Trans.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio IEEE/ACM Trans

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.816960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.816960Z digest=sha256:9ea6c97507f4a2d6ede597306b3b507df6e9118c15ccd0013c092ffb3831f987

Observation f8992d65-b335-44ff-8a88-5c1a996a103e · outbound

This paper cites IEEE Signal Process.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio IEEE Signal Process

Reference 54

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T21:03:19.563918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.851116Z digest=sha256:e02ae621b0b0cf59c59fea04cd9b6d59f15ffb716162a45d5de43c118cd90206

Observation 7dbacb38-7a71-4b55-b03b-b5733748ed6e · outbound

This paper cites an unresolved cited work.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Unresolved cited work

Reference 2010

Resolution
verified exact
doi, observed 2026-08-15T21:03:19.344412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.015535Z digest=sha256:143b16424e3e6d2b79ae8634232f8e1d1bd1b9e928e13b77a759690492d4eb13

Observation 74a1059e-80c3-4b29-8bc1-e367fe2ce770 · outbound

This paper cites IEEE Trans.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio IEEE Trans

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:17.913425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:17.913425Z digest=sha256:094e384ddbd16bd067108898b4255f58de6cc35936c8bf427acb0a1627ba51dc

Observation 5ff84733-7d6a-4f4e-be2c-3a254c86a367 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.011711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.011711Z digest=sha256:31409ae96bf3d55da4e525334abc1ffd7a347a88e814b2ca7656a9f14ef0c77e

Observation 1a6f3b07-ca1e-4114-a627-b60ec6a5a394 · outbound

This paper cites TasNet: time-domain audio separation network for real-time, single-channel speech separation.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio TasNet: time-domain audio separation network for real-time, single-channel speech separation

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:03:20.131691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.320162Z digest=sha256:684884a9fa49231d98631ce9b3cf6d59ee2add9e96eb600c5ad306310a8d35dd

Observation 0c9acae6-ad2b-4163-94f0-21f3adcf759c · outbound

This paper cites an unresolved cited work.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Unresolved cited work

Reference 2018

Resolution
verified exact
doi, observed 2026-08-15T21:03:19.090909Z

Source-reported events for the cited work

correction dated 2019-04-19. Source: crossref record 10.1631/fitee.19e0001->10.1631/fitee.1700814:correction, observed 2026-07-11T03:15:24.066474+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-08-15T21:03:18.350651Z digest=sha256:5a13201877e54cefd08491eb1a4bdbf26490fedf769fffaca65708b53566b7a5

Observation 11b1066b-0906-47cd-be23-3a99fd60c9f6 · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 2019

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:03:17.989518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:17.989518Z digest=sha256:4f6e518dd11296b0942eb6957ed6d962bd2df3cd82892eddbed17a20c06b6558

Observation 37d47b15-71d4-4360-b6cd-5b8e1970ac7e · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:17.950575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:17.950575Z digest=sha256:b3d817fc23e8f993340937799834e42b4541fef981fe06518dc34c74b0cc76ee

Observation 024f2e49-b789-4641-b9b3-848a3f1d696e · outbound

This paper cites In: ICASSP.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio In: ICASSP

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.131176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.131176Z digest=sha256:63308a61141095ce2a79cad23567c19833bfac9375bb0baec391da5ce66fddf4

Observation cbc04179-b8ea-434d-bf0c-fd01d3d343d5 · outbound

This paper cites http://dx.doi.org/10.21437/Interspeech.2024-90.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio http://dx.doi.org/10.21437/Interspeech.2024-90

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.127266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.127266Z digest=sha256:da24e46a2eaddb73394c4397c8c102dd9a50ec13711f9e255762a685b9545576

Observation 1b9fddd5-df7b-427b-9028-fd415c853fb5 · outbound

This paper cites an unresolved cited work.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:03:20.849305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:03:18.008529Z digest=sha256:4625e50f0490094101e50fb0af9b56e01ade0089a7d2fb89d4458739df6bf09e

Observation 75cd3cd3-c3b3-40bb-b499-1bdc7a13a2e1 · outbound

This paper cites an unresolved cited work.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.344532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.344532Z digest=sha256:8f66b592b56b5c22b368e693e9513fe8aef2faf40c929529a141d47a2bffa8cc

Pith citing papers

Observation 6a6e33d5-330d-49d2-8c46-e186c547be93 · inbound

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models cites this paper.

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:46.748745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:46.748745Z digest=sha256:d49756bf2d26012e47ae8a2ef46011b943d28d38b0f66fc912cd4b6d57664e35

Observation 34c99151-aefd-4b70-9762-9dd8993aba5f · inbound

The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge cites this paper.

The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:58:03.164719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T17:58:00.447716Z digest=sha256:9c9e39faf85895023bec612054c3290f430030bd46276e771cc64ad84aa4ce76