Pith. sign in

Paper Citation Record · LEDGER

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

As of 15 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.12672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12672 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.272359Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.073394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:50:17.584368Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact7
  • verified fuzzy15
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · outbound

This paper cites SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:979e1c7f9e28e8a7921155faca9eb30f7fe26d7a573a7e697375e7797e366cf4

Observation 405de8c9-4626-4240-a43a-a36fe2cf999c · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.942252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.078666Z digest=sha256:d9872bc0bb9fda8a246960b57471031c67fdc74bf815d76df0b80e71081d3143

Observation cb0e5e3b-925e-4fc2-b02a-54fdf9ad3b7b · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.924431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.083516Z digest=sha256:316102b0a4384c46d6e53ef36318738c5b34609f504e8c7964509cb83e7d2164

Observation 7807dc91-3110-48ba-93a3-5fe6d82d29ee · outbound

This paper cites Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28].

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.088616Z digest=sha256:2590819bd4b9a5ba1c73a30299e16b52c9b2479d2159d725031806f117343b68

Observation 119c431a-5298-420b-bf3a-0621ff9ce989 · outbound

This paper cites Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.099198Z digest=sha256:f684cfcbfc8a24a307409f86c3e11250b1d84f05c51e0c6368efc3f2e3198ca3

Observation be4c3326-ae81-42a4-97e6-6d78eb03ccf4 · outbound

This paper cites Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:50:17.859960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.104148Z digest=sha256:17d1519f31f5634def7918478cbcc8704f57d978632f9e87caa7c8866b651d59

Observation f73ff6b2-0cf5-4938-bc88-b1d7c2627bd6 · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.110271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.110271Z digest=sha256:56fb14de63f7e518f00bdd63340476ea8c89293370dc043fdee87066730ab038

Observation 4af809b9-9acb-46bb-8a38-6b1edd2a03d4 · outbound

This paper cites Sequence to multi-sequence learning via conditional chain mapping for mixture signals,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Sequence to multi-sequence learning via conditional chain mapping for mixture signals,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.800507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.151926Z digest=sha256:403b17103e7f34e527bd38518d220edab7c478f10fb38da7077421e4382eedcc

Observation 8814306e-e349-45ff-9567-2150e2b26cc3 · outbound

This paper cites Recent advances in end-to-end automatic speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recent advances in end-to-end automatic speech recognition,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.115011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.115011Z digest=sha256:4af3d11153e064023ab3db40b5f836f7d536a073b2aee62db29641fd967c7cb7

Observation f515b9f0-5c58-44cc-9342-61b398a66887 · outbound

This paper cites End-to-end speech recognition: A survey,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-end speech recognition: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.824737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.120316Z digest=sha256:1bc7c84c68485a648ef78f2529ad7169b71d02f4f49f3870a65f81b3a3ad9c41

Observation 951bc0d5-095d-4279-ae7a-28ca126fdbea · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Robust speech recognition via large-scale weak supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.125516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.125516Z digest=sha256:665475de82f52d3353e62014bcabe9d43d62d121cdf2c7651e24dc6b7d4d2905

Observation a4ed4b4e-fe46-4103-a45f-71342bddb5d0 · outbound

This paper cites Anatomy of Industrial Scale Multilingual ASR.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Anatomy of Industrial Scale Multilingual ASR

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.130311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.130311Z digest=sha256:d474cc60ce14eaea35389244100561e6b47dfd07702dcedb9fd8942e82cb172d

Observation 86e970ca-86a7-4fdb-8db1-89a9f125a60d · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.135316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.135316Z digest=sha256:6066773f77407730417486f529a40237429d70d58f3974131ef1cc2e7f2384ad

Observation ed548814-085e-400f-9406-86d11342b4be · outbound

This paper cites Recognizing Multi-talker Speech with Permutation Invariant Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recognizing Multi-talker Speech with Permutation Invariant Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.529996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.141353Z digest=sha256:cc94c4e69ff04e400304d15a40d35e0e9f90392b9e336641fd27eac2ca3c1ba7

Observation d14078c9-7074-4944-a84a-b28828a3ee89 · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.146989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.146989Z digest=sha256:ddcc36bbd46740eb793e6700f37419cd122d6eafd5fcdf6b8958f8e9b32cb483

Observation ba27bf0b-135b-4fbe-bee4-cf7f94d8c084 · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.191684Z digest=sha256:be5c36db0ef663af669f14f86efef216884499436405377f730440b867096e6b

Observation 4ba06cd0-8d43-483b-b401-36c1d57ef1ef · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.490013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.156569Z digest=sha256:3f3b01580680d90424e90161f5cb0bf3ac9d4114b433c0d93f51635ae5346d5d

Observation 4376e618-bfce-4786-8f37-1467ab96e3f7 · outbound

This paper cites Surt 2.0: Advances in transducer-based multi-talker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Surt 2.0: Advances in transducer-based multi-talker speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.784366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.161468Z digest=sha256:5f04e9ffe96fdbd39d8ba133ab5a8813d7bbb832f202ac7b0052f4ea95af850f

Observation d553db92-cce3-4a02-8e06-0f4d2aa68ca7 · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.768604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.166023Z digest=sha256:e6ff9e418bfd70a95d19a40fe17575fd51d8ed75187ae15e6f8f4459d9db2978

Observation 2952a336-5482-45c4-bda7-03457cb27a32 · outbound

This paper cites A Purely End-to-end System for Multi-speaker Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition A Purely End-to-end System for Multi-speaker Speech Recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.464512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.170341Z digest=sha256:83da48e1f627dbd65c039feda609519e92420a6b33a4b4fd3af48e5575e5b935

Observation ecf7c578-90f1-41d9-8383-166d59ebc84f · outbound

This paper cites Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.752648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.176328Z digest=sha256:d36795a953e2b8323e11a06b1338cfe245261e49624566f683847dce7c39d943

Observation e6c902b3-0659-48a7-a629-effaa25cc403 · outbound

This paper cites End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.443388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.181573Z digest=sha256:933b0d21ebd46025a3a84983f0b13d1d6c11cd878acf07ae23ce9bd7fac78e24

Observation 4017e837-3f20-4e2f-bc3e-58ec3e399815 · outbound

This paper cites Neural target speech extraction: An overview,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Neural target speech extraction: An overview,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.186901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.186901Z digest=sha256:5aa2696f4c10fa759b2f17acbf61124cb303b13994dba574de5229c2615e3836

Observation 380de5a2-eb45-4886-8511-bc397d5cc1be · outbound

This paper cites Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.657506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.232416Z digest=sha256:879f165e49879ef9e4a31294b0310b31b9161c6bd6115a57c745388dad27b63a

Observation 3c19fef3-ff92-49fa-a295-6b1fcd445f7a · outbound

This paper cites Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.711415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.196306Z digest=sha256:71056bee077731db094341b99969554ca6d4a39aaacd22d1e7941327f46f5a82

Observation 60b19474-c3ea-4af3-92ef-e7ba9ad5b911 · outbound

This paper cites Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.422697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.200864Z digest=sha256:210164f614a25163a7f8195a72cb07b68d4c1862f54037e934d3dee871232349

Observation 866845d2-be7c-46f1-ad3c-b8423cd743d0 · outbound

This paper cites Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.696412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.206891Z digest=sha256:b4a05832a8cd9491a2c7b6d66690cc4ef184f9e3861ccf6b3484add4c833ec23

Observation 7715e151-8ace-474f-a08c-23f071d90dab · outbound

This paper cites Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.682041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.211578Z digest=sha256:5895cc071992d0b21c17fc6c6d09105b6af260ffd6f6ff1b413845519e05cf9b

Observation 35cb56ab-2885-4e35-916f-a01a5f63ea7d · outbound

This paper cites The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.892600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.093997Z digest=sha256:54615f740fca13c65a80c12341085fc2ccfca0ade53599c3aeae9587ea22c7bf

Observation 09c22c77-82f4-48cd-a98f-8a49be71a03b · outbound

This paper cites Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.216547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.216547Z digest=sha256:d8a593f107c7a2de24691f77809a88a56ce38d21cc2d917738386d1d05ae99d7

Observation 1154dfe7-8571-4985-8f49-c0b593f47eeb · outbound

This paper cites Target Speaker ASR with Whisper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target Speaker ASR with Whisper

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.399004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.222508Z digest=sha256:db5446ba302c120329a40166ab2bc8d1820fe6af702866519f69f30608258d6f

Observation 33eeac7b-ef79-4c00-9143-2aa993c6071f · outbound

This paper cites DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.227570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.227570Z digest=sha256:8950cecfb3d9c99f9cffba0e3b4b2ed5d6998425e24cc362ead26ce4b6718a0e

Observation 61ec0b16-5040-4d03-b1d6-95f7c08db808 · outbound

This paper cites But/jhu system descrip- tion for chime-8 notsofar-1 challenge,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition But/jhu system descrip- tion for chime-8 notsofar-1 challenge,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.640268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.237094Z digest=sha256:7ad850f46c39eb56fad5b471bd3e76c40a00e24273b78c2ab02365664fe57c53

Observation 068f07fe-08e1-42f7-8d45-08c3a86ec4d0 · outbound

This paper cites ESPnet: End-to-End Speech Processing Toolkit.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition ESPnet: End-to-End Speech Processing Toolkit

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.242039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.242039Z digest=sha256:c9585032d02bda2e6acb58ad2986a43df0f098ff50fbeacc27c18063ae27bc85

Observation 4ddde7f5-7f1b-4edd-a0be-6f94237742de · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.625398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.246975Z digest=sha256:6246a82f9976bd2db2ab8cf9e7aad8e861b78a0b8912534cb39c5958ae3d6988

Observation 0123293d-e901-4c82-8787-7cce00d2debe · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.251656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.251656Z digest=sha256:d728100fb0db1d6b5608a78003878c969ccda60cdeb5250486d1a34c57d0c134

Observation 0fa73f01-3f40-4b90-a87c-e5527b34ac03 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.256472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.256472Z digest=sha256:6150d6c5d9f851aa780c4b73fbb7c080f6adbd5eb3c31c46e659eff18f81d083

Observation b39777e2-7e61-493e-9d59-8ffc4ff45661 · outbound

This paper cites Attention is all you need,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Attention is all you need,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.262241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.262241Z digest=sha256:fc68ad5e8ac24768b347d31f92cd96c8536833c7c36fd244784514d740979ce9

Observation 38baa58f-d614-4712-8dd3-d8bce09ac4e2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.267393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.267393Z digest=sha256:3efb1585e96135e380f69598cdb957bfa9bc502821e34b946a7ee6174066b502

Observation 63a2af0e-2c8b-4091-8761-28d663944392 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adam: A Method for Stochastic Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.272359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.272359Z digest=sha256:473ea3e8c9c76e7a5b3880f4aebdd316c69cf1b2cba81981e212329120710c01

Pith citing papers

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · inbound

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition cites this paper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:979e1c7f9e28e8a7921155faca9eb30f7fe26d7a573a7e697375e7797e366cf4