Pith. sign in

Paper Citation Record · LEDGER

Joint ASR and Speaker Role Tagging with Serialized Output Training

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2506.10349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10349 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:32:30.776033Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a56eb7fb-e67d-4b8a-ba72-b8af49fc8c6e · outbound

This paper cites End-to-end speech recognition: A survey,.

Joint ASR and Speaker Role Tagging with Serialized Output Training End-to-end speech recognition: A survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.220372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:27.650243Z digest=sha256:f4f056749efe393acd07e4e2cbd878a46439457d558cd450538facdbd82801c9

Observation a499c853-59cd-44d6-95fc-4aa43855ff14 · outbound

This paper cites A review of speaker diarization: Recent advances with deep learning,.

Joint ASR and Speaker Role Tagging with Serialized Output Training A review of speaker diarization: Recent advances with deep learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.749181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.749181Z digest=sha256:9a1ea3c5752353f633669a6e3e0238d0c8c0d9549ee67050483cd3801db1750f

Observation c6aa7b7a-10db-4bbd-a7ea-4ac0cfda46f3 · outbound

This paper cites Joint vs sequential speaker- role detection and automatic speech recognition for air-traffic control,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint vs sequential speaker- role detection and automatic speech recognition for air-traffic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.201669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:27.859759Z digest=sha256:a6852d04f68a9669310cba8a0a0f228b859886f1cc0a592562ca7da88870a3fd

Observation d9b4c47a-ad4d-405a-b808-72a72c24843f · outbound

This paper cites Joint Speech Recognition and Speaker Diarization via Sequence Transduction.

Joint ASR and Speaker Role Tagging with Serialized Output Training Joint Speech Recognition and Speaker Diarization via Sequence Transduction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:27.969261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:27.969261Z digest=sha256:af12e02f03f237a729d92b3f9f9520a7d02e4d6a0a8026acda5193647775b3ef

Observation d2e40b0a-8d3b-402b-a795-1c6e9919c25c · outbound

This paper cites One model to rule them all? towards end-to-end joint speaker diarization and speech recognition,.

Joint ASR and Speaker Role Tagging with Serialized Output Training One model to rule them all? towards end-to-end joint speaker diarization and speech recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.191045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:28.088900Z digest=sha256:efeb5411c7a4a7fbc681665312e3a5966fb59bd2ac64854537991fa7195e19a8

Observation df57b59b-7deb-4786-a949-d47497047a53 · outbound

This paper cites Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems.

Joint ASR and Speaker Role Tagging with Serialized Output Training Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.219916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.219916Z digest=sha256:ff16d919550889585db89ed5a78885bc8b23ba28c3796c7ed795675d1702db6a

Observation 88a4f7d2-0faa-4bff-8e09-ef6eadc0959e · outbound

This paper cites Serialized Output Training for End-to-End Overlapped Speech Recognition.

Joint ASR and Speaker Role Tagging with Serialized Output Training Serialized Output Training for End-to-End Overlapped Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.346093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.346093Z digest=sha256:8541eba1ccd4afd3a2bcda38db4315379fc7f326c7338383451721c0672db73e

Observation 02be410a-d15b-4189-acfc-7e06fdf39b44 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.501651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.501651Z digest=sha256:8820824cdf0e33eec6a7b2296ca05e4ac1782ac6e8eaa2ff04c124c22857972d

Observation 037f3b84-3a70-477b-82b3-41105d6fdbd0 · outbound

This paper cites Large Language Models based ASR Error Correction for Child Conversations.

Joint ASR and Speaker Role Tagging with Serialized Output Training Large Language Models based ASR Error Correction for Child Conversations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:32:31.149009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:28.582864Z digest=sha256:e39102d228fc350191de0da94cdd1aabb83a1a2a459bd9e405ca7a09b5c3392c

Observation c398d687-7946-4a49-88de-b288647c4b63 · outbound

This paper cites Whislu: End-to-end spoken language under- standing with whisper,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Whislu: End-to-end spoken language under- standing with whisper,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.174849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:28.696382Z digest=sha256:2cdd0f308fc38693b2ad4fe04ba3d494ae6d6411a19a396969f75ed2c04f6b4d

Observation 1590c83d-fced-4424-82f2-2587dd72ea34 · outbound

This paper cites Attention is all you need,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Attention is all you need,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:28.862742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:28.862742Z digest=sha256:b899236f6f9c1f8e7e0b302b6cef2efb5b784a2dca71f931a30d3853d8a7c529

Observation dd031743-beb8-4c7e-ae96-72d9ab3117c5 · outbound

This paper cites Directional speech recognition for speaker disambiguation and cross- talk suppression,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Directional speech recognition for speaker disambiguation and cross- talk suppression,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.158696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:28.959051Z digest=sha256:d6fffa86ba5868acd24d0393df8abd3e842124d8b85662fabc73d62d6441b963

Observation f85097fe-00ee-45bf-aa30-c361fc1584e9 · outbound

This paper cites Agadir: Towards array-geometry agnostic directional speech recognition,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Agadir: Towards array-geometry agnostic directional speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.148804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:29.137476Z digest=sha256:71b51f9cc9b281ff6876e2b47b0aee4d13cc71ce2b5935d90fd9ca7db592370a

Observation a856c219-40db-438d-b3ae-ed6df1be813d · outbound

This paper cites Directional source separation for robust speech recognition on smart glasses,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Directional source separation for robust speech recognition on smart glasses,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.138485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:29.257872Z digest=sha256:a668cd6ba942e02aa77d0305a77cf6565edb8e07db28591359a7cfccfdfafea9

Observation 09ba88e8-626d-4d8d-8e53-1a54f04deab1 · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

Joint ASR and Speaker Role Tagging with Serialized Output Training Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.388286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.388286Z digest=sha256:8b353a402ce3df71f690e8a9009e2d5de5eb678156748db60eb08f646d591f5c

Observation e8f2ff60-b511-4cd0-b851-ba330a7c4caa · outbound

This paper cites Bertraffic: Bert-based joint speaker role and speaker change detection for air traffic control com- munications,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Bertraffic: Bert-based joint speaker role and speaker change detection for air traffic control com- munications,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.126102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:29.500842Z digest=sha256:0cf7cf56ac32c91b20ea92391fa867777d36d70a49d203517e1f4fb46c9797d4

Observation 22877d80-cfea-492a-94bf-01c534ad243c · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.676078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.676078Z digest=sha256:fbc4ce04ca1e09192ca1ae2b1fe9d7f3d506d32476588aa1eb23456e6a6c6454

Observation 10c32fc3-2a65-47b7-ba7d-e6d474f1b1c4 · outbound

This paper cites Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:32:31.012212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:29.756704Z digest=sha256:ff8be7500375488de0a85f950481af92fea4350c7b0df2a7caae93cab382e585

Observation 4d38ee8c-7cfe-404d-86ed-04b746ea24ed · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Joint ASR and Speaker Role Tagging with Serialized Output Training wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:29.844757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:29.844757Z digest=sha256:b284874c94a16b253bf86d39c6404571252919ac8aa795b7f624b4c968d366d9

Observation 4b44b5da-5781-4bf6-a0fe-b697476cb03e · outbound

This paper cites Exploring speech foundation models for speaker diariza- tion in child-adult dyadic interactions,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Exploring speech foundation models for speaker diariza- tion in child-adult dyadic interactions,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:32.023994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:29.972442Z digest=sha256:0b68735849c16385be5eaca5dca10a4dad72c9c26c5c11076121d18aaff7d527

Observation a6e47252-1fb9-4101-80f5-121cfbd47230 · outbound

This paper cites Data efficient child-adult speaker diarization with simulated conversations,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Data efficient child-adult speaker diarization with simulated conversations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.877891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:30.007711Z digest=sha256:80136b4faac20e09bcdd3c89128620e7de50937314013b2f7d1241f8257aa580

Observation 65de492a-1c5c-4f55-874d-43f4866d19d2 · outbound

This paper cites Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits.

Joint ASR and Speaker Role Tagging with Serialized Output Training Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.081659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.081659Z digest=sha256:a4d1770075a8738c42b8b89499634731ea15be84e91eb0d2812bd5ae660ebab6

Observation 74354199-cc9d-4234-8519-b698d0098208 · outbound

This paper cites The chime-8 mmcsg chal- lenge: Multi-modal conversations in smart glasses,.

Joint ASR and Speaker Role Tagging with Serialized Output Training The chime-8 mmcsg chal- lenge: Multi-modal conversations in smart glasses,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.637652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:30.179128Z digest=sha256:e9471c00fce13b47ab30efbd76577c6fa8440e003e81b78c753d6f392715ac7a

Observation 97d95895-7355-4b31-83b6-d50e6d76d3c4 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Joint ASR and Speaker Role Tagging with Serialized Output Training XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.227311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.227311Z digest=sha256:948c685a955bbc0dc608442d28c48c9d4b9e7e2d150e72f8fc6d2eba47d3b687

Observation d2271d5a-ea78-494c-9b26-a5c82d23040f · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.318273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.318273Z digest=sha256:39906d7e60127e03f506b8deebc54931393fd809cabc687ce9807e36c7edb805

Observation d8de2a0e-6168-4d88-ab94-a6a625aaacae · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Joint ASR and Speaker Role Tagging with Serialized Output Training SUPERB: Speech processing Universal PERformance Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.437423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.437423Z digest=sha256:a5b463a8c53c13aab5899b82f0605eef0cbf6499108e3a42cadc78df5ea608c3

Observation 353ba2e7-ac54-4706-b7d4-d03b74d90228 · outbound

This paper cites Playlogue: Dataset and benchmarks for analyzing adult- child conversations during play,.

Joint ASR and Speaker Role Tagging with Serialized Output Training Playlogue: Dataset and benchmarks for analyzing adult- child conversations during play,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.412456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:30.521621Z digest=sha256:6154b5f955fd423dcaff22ac6ffd4dcb00dd102b28998c9d6db8a31ed950371e

Observation 5ddb2afe-44ea-4564-848f-49f535c1b8ef · outbound

This paper cites The talkbank project,.

Joint ASR and Speaker Role Tagging with Serialized Output Training The talkbank project,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:32:31.273720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T04:32:30.571547Z digest=sha256:7b8aa978b981ef85ec8d55c5c0fe9f3a297b9d9a188710ac9382965c52533834

Observation 3c41ea32-f349-4914-95bf-a58fa6aec58b · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

Joint ASR and Speaker Role Tagging with Serialized Output Training NeMo: a toolkit for building AI applications using Neural Modules

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.682646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.682646Z digest=sha256:bae4be8392e1ffac31deb10d8e923237096b4863fedb30ac6ca85c718a6de656

Observation 455c2cf4-d61f-4173-84a1-f0b760fedc38 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Joint ASR and Speaker Role Tagging with Serialized Output Training HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.776033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.776033Z digest=sha256:88707849cc1d50a2f0361d2a2271756ce4ff6d57293d156a1e1729e742e70064

Pith citing papers

No inbound Pith citation observations are available.