Pith. sign in

Paper Citation Record · LEDGER

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition

As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2411.17537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17537 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:13:11.939403Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1fb53bf-ccba-4b84-8587-837ac432bed3 · outbound

This paper cites Improving proper noun recog- nition in end-to-end asr by customization of the mwer loss criterion,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Improving proper noun recog- nition in end-to-end asr by customization of the mwer loss criterion,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.383662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.816274Z digest=sha256:e4b27af0e0096814d1630681f28c07ece0e3c6b7e5e5b592cea3430110ac3956

Observation 171e2db0-00b3-48bc-b65f-c505cbef4ef0 · outbound

This paper cites Personalization of end-to-end speech recognition on mobile devices for named entities,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Personalization of end-to-end speech recognition on mobile devices for named entities,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.367127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.821661Z digest=sha256:f6bb43b3bb3e2b3bca5e256ad51894e6f7da32b2e545c11ea1fcb2b4841e72e8

Observation b12b8d2b-9278-414c-878c-72d4e96a80d5 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.826465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.826465Z digest=sha256:a993caac69d2c764348d66022cbf6fddc453fce0abd67caf3566f99f1d47e7ca

Observation a600b9bd-bfdb-40d5-8f86-e50a9f9d4cc9 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.831434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.831434Z digest=sha256:af079a53ce4859ba849bab744704f5ebbde7068dec65afdfe4050d59cfa44cc2

Observation 7cad08d1-e2ab-45dd-aa9a-37f66686eb3a · outbound

This paper cites A better and faster end-to- end model for streaming asr,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition A better and faster end-to- end model for streaming asr,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.350590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.836718Z digest=sha256:8d2f3ccca74642de1b5f3558475b175063af8b45b90b76a4f0c5d619a2e33d86

Observation 8d1623b1-1960-4911-9663-a6a033a65439 · outbound

This paper cites Cascaded encoders for unifying streaming and non-streaming asr,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Cascaded encoders for unifying streaming and non-streaming asr,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.335630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.841431Z digest=sha256:cbc6982b52011df3f8b0e9dc2b1c996ff4f21d581ea996347e5a638c6c21741a

Observation 6f2010c9-a873-4e92-b3ba-2c06ffbcbeff · outbound

This paper cites Dual-mode asr: Unify and improve streaming asr with full-context modeling,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Dual-mode asr: Unify and improve streaming asr with full-context modeling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.320471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.846582Z digest=sha256:269c48c2cec62b87d0ca7d192ca0c468798b4e422b4e2407f16f153f6282ef11

Observation 5a344ed2-2b68-45de-9f4d-a90fc9b05648 · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.851129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.851129Z digest=sha256:aaf639943c8f415131ff96c7a5862461a774de8c5863dc368deafb67f608bbb9

Observation b46cab54-a763-478d-a4e9-e98026def95a · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recur- rent neural networks,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Connectionist temporal classification: labelling unsegmented sequence data with recur- rent neural networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.294940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.855609Z digest=sha256:3e67425cd4b80fd9fb984216c75ee9989cb3724b8e7db6f65916ec7c56864912

Observation d4108838-eac7-4adb-9417-b30939bcd4d3 · outbound

This paper cites Sequence transduction with recurrent neural networks,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Sequence transduction with recurrent neural networks,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.279887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.860027Z digest=sha256:0c592267605a8440d1caab7fe944060dd62ba5557fc2a13acd6b0526b6daec56

Observation 0b875d82-d58c-4c7a-9d4b-53ce49212ab7 · outbound

This paper cites End-to-end attention-based large vocabulary speech recognition,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition End-to-end attention-based large vocabulary speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.263759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.864442Z digest=sha256:498458080b3b01cdb070b176dae98eb793bfc2280a7caf7eaeba1d0dba9dbd52

Observation 3e4a5c99-a609-41d8-b611-ac70dfbab78f · outbound

This paper cites Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.249338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.870010Z digest=sha256:48c540cc91c9690bbec20c4492cdb492d474b1e858ea3a2e624e48366a432d25

Observation 4d9306e7-1ed0-4a94-89ec-db00dc4286a2 · outbound

This paper cites CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:13:12.016770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.875107Z digest=sha256:4b28fa385ade8e7535b094b50c06b322f8955316d0803aa9e87dc33180b5ef8f

Observation 10e80c2f-73f4-4c60-b04a-aa19bcfc006b · outbound

This paper cites Crf-based single-stage acoustic modeling with ctc topology,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Crf-based single-stage acoustic modeling with ctc topology,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.234353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.879782Z digest=sha256:6dd83673c0d85957962f5ac988b9ae54e3ea034d815468eeca0dd39ff74da70a

Observation 5903ca51-3a82-4b88-a792-a8a449583128 · outbound

This paper cites Imputer: Sequence modelling via imputation and dynamic programming,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Imputer: Sequence modelling via imputation and dynamic programming,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.219564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.884138Z digest=sha256:fcdee77087c14229cec2719813ca04fcb2c4793c761529c2a0fa4fb064fca435

Observation e9b2fdcd-caf4-44c9-880f-b99446e699a6 · outbound

This paper cites Global normalization for streaming speech recognition in a modular framework,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Global normalization for streaming speech recognition in a modular framework,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.204049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.888434Z digest=sha256:bb546247d7e3eb1a45848f99976831737df4f793d067869b22f86001c9a7572e

Observation c55bfd57-f5ec-470f-bc88-34313b235760 · outbound

This paper cites An unsupervised autoregressive model for speech representation learning,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition An unsupervised autoregressive model for speech representation learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.189145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.893079Z digest=sha256:49ec21bbd9d565d203ac8cdb594bd3269df1ee342f4e8b81dcfd5d4a1e49f433

Observation 7d281fd7-cbe5-4a1f-9e41-fb57c8f17546 · outbound

This paper cites Variational inference with normalizing flows,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Variational inference with normalizing flows,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.174075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.897533Z digest=sha256:46c63966da2fed16956d4c1f609ae0df6d72997044a9d81850c38b66919cf03f

Observation 0d4597f1-bc0a-447f-9fe8-b7d039341609 · outbound

This paper cites https://github.com/k2-fsa/icefall.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition https://github.com/k2-fsa/icefall

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.158985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.901971Z digest=sha256:4e50959905067afc7c1137cb0535ebc9efc423a88de7c15ff3068b34a986d6b4

Observation 7eef7b32-dfc2-4500-b436-de80460cd04d · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Librispeech: an asr corpus based on public domain audio books,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.143688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.906742Z digest=sha256:f2e4f8bc050f47bf056688e8f63b20d6a0ea06d78aa25365f3a9dff5902d9221

Observation 343bbec1-16af-49d8-8430-fb1ee3a486aa · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.127432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.911102Z digest=sha256:112dd5ad938789a0af2f0d7e55ca3d8a339fc78145a47f757ebc480609200e02

Observation 74753a00-dfa4-4b74-83f1-23b4aadd3664 · outbound

This paper cites A new algorithm for data compression,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition A new algorithm for data compression,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.112275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.915698Z digest=sha256:d0bf231123290b487bddf5a54e10b86de5a4b493c5e7ac49b3f10102d9d1bb07

Observation 1f671f66-86ad-4e7b-84ce-7b688cad4f50 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Neural Machine Translation of Rare Words with Subword Units

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.920124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.920124Z digest=sha256:24f1f79d76cb9f9da2ff1cc461ce933b40e4a6b645aebfa41b42da26640f1f52

Observation 28b369b5-17ba-4972-859f-618ccaef638b · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Zipformer: A faster and better encoder for automatic speech recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:11.925137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:11.925137Z digest=sha256:534a8e84eaa4f407718b238244c47ff1779b37dc1c54814bb72e024e8b89574c

Observation 250cd683-1af0-4a57-91ea-0b7d501eabd6 · outbound

This paper cites Rnn- transducer with stateless prediction network,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Rnn- transducer with stateless prediction network,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.097217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.930273Z digest=sha256:ffc29ef204798ed09929656fe1dc8f41fb8e38863b036722fc5fe94038d5c089

Observation 011b3b63-fd94-4e61-8ad7-803d89938d9e · outbound

This paper cites Made: Masked autoencoder for distribution estimation,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Made: Masked autoencoder for distribution estimation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.081969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.934933Z digest=sha256:b4198781c3dc87b8697265408a35449302250fb41d8b2c9753a93b13fd30a145

Observation 42beed4c-1604-49a0-8c7b-e468d514b2cb · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:13:12.066442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:13:11.939403Z digest=sha256:46609ac8e98a5bdc00bc304972ad2d71e26bc85045163e1353964283f82e5029

Pith citing papers

No inbound Pith citation observations are available.