Pith. sign in

Paper Citation Record · LEDGER

Joint Speech Recognition and Speaker Diarization via Sequence Transduction

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:1907.05337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1907.05337 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T00:59:00.525481Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:27.969261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-25T01:00:09.574967Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df2148e4-efbe-47bf-9f6a-169c1c4e18ba · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.578155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:e102b85ba882694fb20c116db07f1401b741e4b127e6bb839daae6549c8b73b0

Observation dd1d165d-6cab-4e0b-89f6-5acb1a9381b2 · outbound

This paper cites Problem Formulation and Proposed Solution Many machine learning tasks can be expressed as mapping an input sequence into an output sequence.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Problem Formulation and Proposed Solution Many machine learning tasks can be expressed as mapping an input sequence into an output sequence

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.300720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:50dc2dfe3b877ca69635cbe871ccfe8782bdc5dcce734e25127d5970092ce335

Observation d88a58d1-b5af-475b-8a69-feb4e749a31c · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.362086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:98d7ee60cc63866ad70ca80d1409487e2917027c9722721c5605814668391231

Observation 7993a0b8-0a36-4f24-a645-d9ff7c842276 · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.240446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:244b590e18abc53843050d8c1ca3a9b2f1920529f34b41c0daf90decf2c0481f

Observation 29d16264-0740-4a61-bd05-4d6862a3a615 · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.381972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:6b1a8e32236c2dd556c4a37a142269a84382978b4518680ddf489556330aabb7

Observation 97e3ca92-1e81-4f84-bc05-55482e9c14ab · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.228658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:f76839db047c9bb4f90082a68a9e7dbda512a3e98cbe06a7e808b04f54ac1995

Observation 98cd67ce-d600-4ad2-b90a-5ac506107060 · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.236456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:7241383558fe65bc2ea3c13ad640e93aee57e66dccc4ad150ba72d0d300992ff

Observation d3f6adbc-08e1-4b4d-8dc7-68c887a5b5a9 · outbound

This paper cites We demonstrated the performance of our approach by evaluating it on a large corpus of clinical conversa- tions between physicians and patients.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction We demonstrated the performance of our approach by evaluating it on a large corpus of clinical conversa- tions between physicians and patients

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.292305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:eaa9804fcb334c3092293eace86556c9fcbb50adc42db33368dcc6f1a74f927f

Observation 07b88adf-565a-46cb-a228-2d068797be7f · outbound

This paper cites an unresolved cited work.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-25T01:00:10.366683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:771ae5de4114728826d26a5d74fd456633c00c5b67b4526c7e08edde9a32bef2

Observation 175393fa-3556-4888-ad78-2b18e7e2507a · outbound

This paper cites An overview of auto- matic speaker diarization systems.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction An overview of auto- matic speaker diarization systems

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.375332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:4412e4dda3dadc68fa58f7e0e40deac926ca45e1234b929c6de7c426d76800a4

Observation 5e9e5dc7-7e68-4d98-a67c-e79a4aedee57 · outbound

This paper cites Speaker diarization: A review of recent research.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization: A review of recent research

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.378649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:2fb4d35dda9f3637b7a1e0c524fee2dc2e0e62ada861baf8d616074c2da7080a

Observation e92e4095-589b-462b-a11f-c7b73fd89946 · outbound

This paper cites A robust speaker clustering algo- rithm.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction A robust speaker clustering algo- rithm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.350458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:6bc65b7c9a80f1338f734fa24e217b732efaa0689ad59280cfa7275e9935a600

Observation 23ca0396-8c5c-4579-b7e9-b1e22ac3d0ae · outbound

This paper cites Multistage speaker diarization of broadcast news.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Multistage speaker diarization of broadcast news

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.334842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:bb1ddcf80e57902bec77ea3e436925298db11cb5bf618e00ffd5cc41c3d912bb

Observation 8fd2b52d-a994-4177-ba77-294dc10d82d6 · outbound

This paper cites Speaker diarization with PLDA i-vector scoring and unsupervised calibration.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization with PLDA i-vector scoring and unsupervised calibration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.287755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:b8817aa518f92436dcc90585130833e2d4e42a2f5f2143f3efd39f8fa5faf88f

Observation e487747d-1160-474c-b7ad-dbc58b25a482 · outbound

This paper cites Speaker diarization using deep neural network embeddings.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization using deep neural network embeddings

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.393043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:a77b41fd573a34fff3f861be927b9f388c4a4f68bac6cf79879a606f3c4b439b

Observation bea5c21e-eceb-46a1-9788-309e58dfb63e · outbound

This paper cites Speaker diarization with LSTM.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker diarization with LSTM

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.338336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:b28a10adb3afad721b830287005d28097283bf36cdc9f67519722bed238f3247

Observation fc8682a6-ce90-476f-9257-60ca9c24f8b0 · outbound

This paper cites X-vectors: Robust DNN embeddings for speaker recogni- tion.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction X-vectors: Robust DNN embeddings for speaker recogni- tion

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.358453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:cf3adf928b1ddfba320f057de937bab43a201cbcaaafbce17aa7a8f1f6cb041b

Observation 8f3d676e-8def-4d48-82c4-8756d3ccabc1 · outbound

This paper cites Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD chal- lenge.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Diarization is hard: Some experiences and lessons learned for the JHU team in the inaugural DIHARD chal- lenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.371970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:8628b7f8f8f52f533baa6e56b254e4d07651a968e782629b23fc934d8d2ed9b4

Observation 8805d2e1-9e96-4d0e-b8a6-5d50f4562c44 · outbound

This paper cites Tristounet: Triplet loss for speaker turn embedding.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Tristounet: Triplet loss for speaker turn embedding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.389529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:4619c028a8a9ad9f98cb640d119e9f2f1b974c2a04838f1eb0ca2fcdd8f7d01b

Observation 7d2a9063-ea67-4c88-84bb-a4d65a32540a · outbound

This paper cites Fully Supervised Speaker Diarization.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Fully Supervised Speaker Diarization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.565790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:39b8b47d8389ad065bbf5d11ae5745325af53491b01e00ce593e50452b5f4413

Observation cc1057fd-5d2d-46d6-97cc-be24bc01a83a · outbound

This paper cites Speaker di- arization from speech transcripts.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speaker di- arization from speech transcripts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.385332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:a98ce185b5fe3d48c7743a41fa55d873519aca218bbe9a545df66d0394fad503

Observation 7ec0a56b-c577-4a66-be1c-ac97361682bc · outbound

This paper cites Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to se- quence neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to se- quence neural networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.354194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:5989d226f8352b26154677d6c4406421c466a91f7119db87280ae35711a18a7b

Observation 42913212-c0ef-4597-b76e-e529b5aecee3 · outbound

This paper cites The use of recurrent neural networks in continuous speech recognition.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction The use of recurrent neural networks in continuous speech recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.273558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:1fea1bc878ca5e73a95b378bb88255f35e53b97c7d6f718ce6345e8c9dd91468

Observation f178c1d3-fdcf-4ecd-b75e-b7e18570297d · outbound

This paper cites Connectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Connectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.277767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:7dacd844df9ac5b419d40d9dea309b446d530bd11ed5244baccc1e252f3430d2

Observation bf21111b-9427-4ff4-b75b-b0132871c77a · outbound

This paper cites Sequence transduction with recurrent neural net- works.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Sequence transduction with recurrent neural net- works

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.282214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:4f88f0e648f6d30417ed2fec9ba0370722b29246a1f968ff09375a0766145f2d

Observation 197d9a67-66ea-442a-80cc-916a7cfb3806 · outbound

This paper cites Speech recognition with deep recurrent neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Speech recognition with deep recurrent neural networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.232528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:e1c816bd7990f3902c9b7afe8dc462061eb24a7c7f215795f63f5716e724c894

Observation 6eb82b7c-d5fa-4d52-a9cc-cd65ea5ee047 · outbound

This paper cites Streaming End-to-end Speech Recognition For Mobile Devices.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Streaming End-to-end Speech Recognition For Mobile Devices

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.571824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:61df2ec508b325081d2aa9aaaa0d809b1b768ef8251504303a8ee1f93968741d

Observation eacc9491-ce87-4ba4-a3cd-4c130cd3a670 · outbound

This paper cites In- datacenter performance analysis of a tensor processing unit.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction In- datacenter performance analysis of a tensor processing unit

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.342565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:c05c1738c2b9fc46ca29ea47933051da6551fcbe202c66c095221c28482542d4

Observation c5f8e694-5f06-469f-a95c-4dae748b3ed5 · outbound

This paper cites Improving the efficiency of forward-backward algorithm using batched computation in tensorflow.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Improving the efficiency of forward-backward algorithm using batched computation in tensorflow

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.327322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:ed85f9916b0843befad3f66498a217afc2dcdaa23d8f7df4e3f2df2aed97806c

Observation 04feb0f1-01c7-4c3c-b3ac-b1108f51d59b · outbound

This paper cites Efficient implementation of recurrent neu- ral network transducer in tensorflow.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Efficient implementation of recurrent neu- ral network transducer in tensorflow

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.305342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:e5748e17c7d448e2df2f4de34254edba659715ec35a5ee8bbdbf603ccdd750b9

Observation f513f50c-df4a-451d-af65-857af8e9acf2 · outbound

This paper cites Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.396781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:9e64e34a9bc6351c95b63ecb3171d96444bf6eb4b405198d85eff57e86b972ea

Observation 9ec9e717-29f6-49a8-94b9-7bdfed362b15 · outbound

This paper cites Morfessor 2.0: Python implementation and extensions for morfessor base- line.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Morfessor 2.0: Python implementation and extensions for morfessor base- line

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.323290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:44cc8e7f60dee4e9dc1c3a2d7d92e27b92560071be4206543b22cc319212d49c

Observation ddc453d7-d7fe-4be0-8c52-ea030f679765 · outbound

This paper cites Phoneme recognition using time-delay neural networks.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Phoneme recognition using time-delay neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.296298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:ba9637eb7043dba655fcd65eeda936da53991a1da4b30b817fecf15cbc695742

Observation 77f9fa79-f727-49da-9c24-555afded2daa · outbound

This paper cites Reducing the computational complexity for whole word models.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Reducing the computational complexity for whole word models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.331126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:803d349a2fc0a185bd8544ef57532413293716c9b89619696a369e5f9e1bfa13

Observation b0b87f38-9eb6-403b-9f69-ff41a168d1d7 · outbound

This paper cites Long short-term memory.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Long short-term memory

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.346710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:5572bdacdfac7d97f584d7461c1dd7eff093afd5152d8f6d0b61e0c5179593b1

Observation 86aae55a-6dbc-4151-b93a-d76934d4abf6 · outbound

This paper cites Adam: A method for stochastic opti- mization.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Adam: A method for stochastic opti- mization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.268808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:130cb827579653e80a7c9afc0056a2a366667357b11abfb1e6533861be3ccb87

Observation 9cbfc75a-8ef7-479f-9943-34efdec55b15 · outbound

This paper cites The Rich Transcription Fall 2003 (RT-03F) Evalu- ation Plan.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction The Rich Transcription Fall 2003 (RT-03F) Evalu- ation Plan

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.244402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:60670c6c2e0cc01bb6670ad32edcd7f8206e15729e7d3311f8c5175d4c83e522

Observation 345aaecc-4234-4cd8-b7af-20877cdb534e · outbound

This paper cites Feature learn- ing with raw-waveform cldnns for voice activity detection.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Feature learn- ing with raw-waveform cldnns for voice activity detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.310789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:b61e94809381e07b3e1d3f52dbe34f0b06b5056e74fa5a7bcd4bf3f9cb2513b1

Observation 2d39334c-f8b3-4a0e-befd-0d94e98cd608 · outbound

This paper cites End-to- end text-dependent speaker verification.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction End-to- end text-dependent speaker verification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.314975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:7ac9893cb954230a0b21d278facd9f6aa22c2ba3eb4b29d7a134a5ab8ce17602

Observation 7f4a7fb6-5a2c-4945-8ac6-00b760f5b051 · outbound

This paper cites V oxceleb2: Deep speaker recognition.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction V oxceleb2: Deep speaker recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T01:00:10.318901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:42dfb8512914b71ace50305606abcca70995468840b8eb1e9f8ea12abff346cd

Pith citing papers

Observation df2148e4-efbe-47bf-9f6a-169c1c4e18ba · inbound

Joint Speech Recognition and Speaker Diarization via Sequence Transduction cites this paper.

Joint Speech Recognition and Speaker Diarization via Sequence Transduction Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:00:09.578155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T00:59:00.525481Z digest=sha256:e102b85ba882694fb20c116db07f1401b741e4b127e6bb839daae6549c8b73b0

Observation d9b4c47a-ad4d-405a-b808-72a72c24843f · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.969261Z digest=sha256:0f34cfc6ecab47f53f3cd2505ad8a5a2e89523eb19b1b1c53ed4d8bd29771196

Observation 80d5c565-14c0-4f84-9a26-578fc0102d5e · inbound

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models cites this paper.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.818718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.818718Z digest=sha256:6a1439e2c16c3943834c7f7c041a1ba8dcb435335a7b46d531cfb026ab32e27e

Observation 612daea9-4d17-4202-b167-cc2aa4c16239 · inbound

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation cites this paper.

TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:27:29.255358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:27:29.255358Z digest=sha256:e246502fb322d958918f69ed875309b1330d10432db5f7072194aed638a6e00d