Pith. sign in

Paper Citation Record · LEDGER

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.12672.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12672 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.272359Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:17.073394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:50:17.584368Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact7
  • verified fuzzy15
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · outbound

This paper cites SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:99b0b11b8b116b8b7abd14b116bfc22eda2a08eddf54642320619e71fbd3b6ba

Observation 405de8c9-4626-4240-a43a-a36fe2cf999c · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.942252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.078666Z digest=sha256:3cc549022e6a2b1e8f7dd2c9f9ece5f5d5f41443bb34e011e7fa3b1646d24d01

Observation cb0e5e3b-925e-4fc2-b02a-54fdf9ad3b7b · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:17.924431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.083516Z digest=sha256:e0937cecc7311cec9550f428188a678ba8b2242f28473a6ce55fb3de8c176dd6

Observation 7807dc91-3110-48ba-93a3-5fe6d82d29ee · outbound

This paper cites Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28].

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Dataset The training set consists of LibriSpeech-360h [27], Libri2Mix, and Libri3Mix [28]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.909059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.088616Z digest=sha256:a7322df650e382af3d787dc8a9407a2a9758fef5d329d8a916482570f4a7fc6b

Observation 119c431a-5298-420b-bf3a-0621ff9ce989 · outbound

This paper cites Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our proposed SC-SOT models, particularly when combined with multi-task learning (MTL), demonstrate promis- ing performance gains over the conventional SOT-based model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.099198Z digest=sha256:07f49c0cf99ed655d3603bd19c02d15027ce0933a5b6da45d8e03358e2ae721b

Observation be4c3326-ae81-42a4-97e6-6d78eb03ccf4 · outbound

This paper cites Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Our initial analy- sis, through visualization of source-target attention weights in an SOT-based model, revealed that the decoder performs a de- gree of speaker separation

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:50:17.859960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.104148Z digest=sha256:5339517ca37ca39a1d62e88f9c128ebcecdb3de8ee711e70734392412f8097a2

Observation f73ff6b2-0cf5-4938-bc88-b1d7c2627bd6 · outbound

This paper cites an unresolved cited work.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.110271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.110271Z digest=sha256:351ff85976ac764cbb1e5f10382b96e3d79f3f3f98206be632b053b2c3bd3039

Observation 4af809b9-9acb-46bb-8a38-6b1edd2a03d4 · outbound

This paper cites Sequence to multi-sequence learning via conditional chain mapping for mixture signals,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Sequence to multi-sequence learning via conditional chain mapping for mixture signals,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.800507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.151926Z digest=sha256:ee1562d8f2663220c4bac863ad9d8a1cd4cd3223a932ec95c4d5c3757b6b1f93

Observation 8814306e-e349-45ff-9567-2150e2b26cc3 · outbound

This paper cites Recent advances in end-to-end automatic speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recent advances in end-to-end automatic speech recognition,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.115011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.115011Z digest=sha256:e4191532ee21fb0fc09c91555f8130c0e50ff7088c0775ddd55c6acd5b562672

Observation f515b9f0-5c58-44cc-9342-61b398a66887 · outbound

This paper cites End-to-end speech recognition: A survey,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-end speech recognition: A survey,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.824737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.120316Z digest=sha256:0460b2aa83b15b73fd2c751ed6cea758d7f0be342f1d61a7c12c9f7436ea4700

Observation 951bc0d5-095d-4279-ae7a-28ca126fdbea · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Robust speech recognition via large-scale weak supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.125516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.125516Z digest=sha256:5b57519c00a796bb0484281b6878b665585245a8204c196f8e03c56d8b3248f7

Observation a4ed4b4e-fe46-4103-a45f-71342bddb5d0 · outbound

This paper cites Anatomy of Industrial Scale Multilingual ASR.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Anatomy of Industrial Scale Multilingual ASR

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.130311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.130311Z digest=sha256:e7d25645c25223c7364d7f5018120f442f1dbe34dc4fcaade6e0245a59b2529f

Observation 86e970ca-86a7-4fdb-8db1-89a9f125a60d · outbound

This paper cites Less is More: Accurate Speech Recognition & Translation without Web-Scale Data.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Less is More: Accurate Speech Recognition & Translation without Web-Scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.135316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.135316Z digest=sha256:d6ed1899d27d0f7532a25c48bb5e338e9c78409617af382ac8d98c9cbfecf129

Observation ed548814-085e-400f-9406-86d11342b4be · outbound

This paper cites Recognizing Multi-talker Speech with Permutation Invariant Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Recognizing Multi-talker Speech with Permutation Invariant Training

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.529996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.141353Z digest=sha256:58514d995c4c0187094b120691e0d4b6cdc1d7ab1e5ee29c609279a6bdcc6f01

Observation d14078c9-7074-4944-a84a-b28828a3ee89 · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.146989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.146989Z digest=sha256:ba05265cea7242ff040bc1e2e67dd9ed80d878b7f22587139164afc168f10a2a

Observation ba27bf0b-135b-4fbe-bee4-cf7f94d8c084 · outbound

This paper cites Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Speakerbeam: Speaker aware neural network for target speaker extraction in speech mixtures,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.191684Z digest=sha256:de8e7347749255d38bb85b49c0c9d0d69f88ca293fecbfe27d282d53527dca0d

Observation 4ba06cd0-8d43-483b-b401-36c1d57ef1ef · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.490013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.156569Z digest=sha256:9f885b4c2e723e0ddb9ec1fd7e2997704db50080d9cce37805aa9b23c5e1fc88

Observation 4376e618-bfce-4786-8f37-1467ab96e3f7 · outbound

This paper cites Surt 2.0: Advances in transducer-based multi-talker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Surt 2.0: Advances in transducer-based multi-talker speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.784366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.161468Z digest=sha256:f2d8ea1873f5fad7c8c337b5033fea0c842e25eacc4da6304c8264cacbe2d9e5

Observation d553db92-cce3-4a02-8e06-0f4d2aa68ca7 · outbound

This paper cites Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Permutation invari- ant training of deep models for speaker-independent multi-talker speech separation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.768604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.166023Z digest=sha256:b7535b5898ddc81c974ad3799fc0cfe6586ad9be75a8aae93d8146af1d242a6f

Observation 2952a336-5482-45c4-bda7-03457cb27a32 · outbound

This paper cites A Purely End-to-end System for Multi-speaker Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition A Purely End-to-end System for Multi-speaker Speech Recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.464512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.170341Z digest=sha256:2fdca0b4fc1487dcbc5db4ee215aad784b75befc2c9454c6f6487d55f254cf98

Observation ecf7c578-90f1-41d9-8383-166d59ebc84f · outbound

This paper cites Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Mimo-speech: End-to-end multi-channel multi-speaker speech recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.752648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.176328Z digest=sha256:6ed928c12a5f0130e49749ab88b12ab4b6d04076e426937435a1384a8243048c

Observation e6c902b3-0659-48a7-a629-effaa25cc403 · outbound

This paper cites End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.443388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.181573Z digest=sha256:359e8afd9a2adf0baeccb72ad1e46af5d08026d356071b7276beedd1c5fe84b6

Observation 4017e837-3f20-4e2f-bc3e-58ec3e399815 · outbound

This paper cites Neural target speech extraction: An overview,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Neural target speech extraction: An overview,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.186901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.186901Z digest=sha256:a3f64cd38ae214e7409a03d293a88c4f397573a39e44c3ab3eab86c23d491264

Observation 380de5a2-eb45-4886-8511-bc397d5cc1be · outbound

This paper cites Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Transcribe-to-diarize: Neural speaker diariza- tion for unlimited number of speakers using end-to-end speaker- attributed asr,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.657506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.232416Z digest=sha256:592a01ddf5eadd7c14e3774e50601d85e179d9157e6263e496826447de32dbc9

Observation 3c19fef3-ff92-49fa-a295-6b1fcd445f7a · outbound

This paper cites Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.711415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.196306Z digest=sha256:cc908b6cd30e4f6687692e68bc3267992984452557b1c3d5254869baef553c7b

Observation 60b19474-c3ea-4af3-92ef-e7ba9ad5b911 · outbound

This paper cites Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.422697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.200864Z digest=sha256:1c6bfd73f91c265d855483f919235161a9934dcb4f81fd9d916f00b5598146e7

Observation 866845d2-be7c-46f1-ad3c-b8423cd743d0 · outbound

This paper cites Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Tar- get speaker voice activity detection with transformers and its in- tegration with end-to-end neural diarization,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.696412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.206891Z digest=sha256:9f9610625de850dcfc3ffbaed5c59fae15538c9d923fe4566bee13d7357c4474

Observation 7715e151-8ace-474f-a08c-23f071d90dab · outbound

This paper cites Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adapting self-supervised models to multi-talker speech recognition using speaker embeddings,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.682041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.211578Z digest=sha256:f824500b14bf17d9ed46b2b34bcf567e96ec7bda0354a0fb88674db3102bd9f7

Observation 35cb56ab-2885-4e35-916f-a01a5f63ea7d · outbound

This paper cites The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition The Conformer encoder has 12 layers, while the Transformer decoder is structured with 6 layers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.892600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.093997Z digest=sha256:abed80136063ff662e34ed96f7d132def0484861859b569b1692125ac7069634

Observation 09c22c77-82f4-48cd-a98f-8a49be71a03b · outbound

This paper cites Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.216547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.216547Z digest=sha256:3b4558df8f91252dcb5a072610993292e43f5de803bde1b9cc9c403783b7ca4f

Observation 1154dfe7-8571-4985-8f49-c0b593f47eeb · outbound

This paper cites Target Speaker ASR with Whisper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Target Speaker ASR with Whisper

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.399004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.222508Z digest=sha256:aff51dfaea6eac35838bf1a9dbd9f433acc43c7a94f70630b4589efbd6c8238f

Observation 33eeac7b-ef79-4c00-9143-2aa993c6071f · outbound

This paper cites DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.227570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.227570Z digest=sha256:bc22d2b2f0cfd6980ec569bf2eabbdd6531c5d7d16083f4256f4c369860a190e

Observation 61ec0b16-5040-4d03-b1d6-95f7c08db808 · outbound

This paper cites But/jhu system descrip- tion for chime-8 notsofar-1 challenge,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition But/jhu system descrip- tion for chime-8 notsofar-1 challenge,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.640268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.237094Z digest=sha256:fba0398e52fcf2154419c2f2f550046672bd932380a2ded59a1af9b290975f86

Observation 068f07fe-08e1-42f7-8d45-08c3a86ec4d0 · outbound

This paper cites ESPnet: End-to-End Speech Processing Toolkit.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition ESPnet: End-to-End Speech Processing Toolkit

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.242039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.242039Z digest=sha256:a20e7e478fb6914e200c36e2845c5ef45230029cdb623aa56b4d2f7fd11b99f8

Observation 4ddde7f5-7f1b-4edd-a0be-6f94237742de · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:17.625398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.246975Z digest=sha256:090e46d8ffe3668cd776e11d79bf36afee13142065f438293819f4ea15134226

Observation 0123293d-e901-4c82-8787-7cce00d2debe · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.251656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.251656Z digest=sha256:5ab4603ae0a389357c6fb84a34d279a8fb59feef6caf285efdac25464478e903

Observation 0fa73f01-3f40-4b90-a87c-e5527b34ac03 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.256472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.256472Z digest=sha256:40aa2fda163b5dfa8b200845ad4d884232e4f8fe39ff5bdefb93fe29798e77b8

Observation b39777e2-7e61-493e-9d59-8ffc4ff45661 · outbound

This paper cites Attention is all you need,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Attention is all you need,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.262241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.262241Z digest=sha256:7c625048956369edb9d968b140cfc8cc2b7de440748b7518924b7e4caeb267a5

Observation 38baa58f-d614-4712-8dd3-d8bce09ac4e2 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.267393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.267393Z digest=sha256:ae21e1aa48bc8d48dfc8deba44a48936b97304a423f35f5370f130224ff0c9cf

Observation 63a2af0e-2c8b-4091-8761-28d663944392 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Adam: A Method for Stochastic Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.272359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.272359Z digest=sha256:505bab5c2d0f21783eab15fd33561b1f37e77b5890067d71baea0b0427d4cf10

Pith citing papers

Observation 1ad7751f-54f4-4b45-964d-2f03033599d5 · inbound

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition cites this paper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:17.589525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:50:17.073394Z digest=sha256:99b0b11b8b116b8b7abd14b116bfc22eda2a08eddf54642320619e71fbd3b6ba