Pith. sign in

Paper Citation Record · LEDGER

WhisperFlow: speech foundation models in real time

As of 15 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2412.11272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11272 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:11:57.955428Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T22:16:51.917336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T22:21:53.649506Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved37
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1fee2c3-e370-41a2-9439-93222cf0253f · outbound

This paper cites Accessed: 2024-11-3.

WhisperFlow: speech foundation models in real time Accessed: 2024-11-3

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.936833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.731021Z digest=sha256:bc45a9f1d84d2f4dfd5abf62fded1cf9b8e0274c7d481c162360d9dceafbc895

Observation dc2052ab-9b74-4902-b738-bf19f9a44b52 · outbound

This paper cites GPT-4 Technical Report.

WhisperFlow: speech foundation models in real time GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.735305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.735305Z digest=sha256:a6198446a25afef2c7ea14b27666f376278102067547bd27e0573fc8a138418a

Observation a17cf4bb-74a5-464f-ba1c-3fa280105781 · outbound

This paper cites Did you hear that? Adversarial Examples Against Automatic Speech Recognition.

WhisperFlow: speech foundation models in real time Did you hear that? Adversarial Examples Against Automatic Speech Recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.738385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.738385Z digest=sha256:52cbb1a6ac530d4f80eb6e6a2c5261479ccffe09d91a728fca7f1922b4eeda39

Observation 30577563-e2bb-4df3-8748-a55b7caaaadf · outbound

This paper cites Apple macbook air tech specs, 2024.

WhisperFlow: speech foundation models in real time Apple macbook air tech specs, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.926431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.742310Z digest=sha256:73255624518185018f12d63e17ff3cb968be00cfbf9bcaeefc7e49f449063430

Observation 7ca7c6ef-9f0b-4613-9aee-f4876768af59 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

WhisperFlow: speech foundation models in real time Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.745362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.745362Z digest=sha256:86e16d39099ce68bb0df7e01a1d0fb3a26973accd4f701aed5af41c9a2c6c1bc

Observation 246f45ca-6e4b-40d1-b74b-08e9cbebf41d · outbound

This paper cites an unresolved cited work.

WhisperFlow: speech foundation models in real time Unresolved cited work

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T15:11:58.596857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.748694Z digest=sha256:b7d91adefcc2b83db5e142d5c6376b9c6ba75f03c7c42460b225c1fb9fb42bbc

Observation 62593c14-0bf6-46d0-b9da-25d63e01403c · outbound

This paper cites A mathematical theory of adaptive control processes.

WhisperFlow: speech foundation models in real time A mathematical theory of adaptive control processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.751674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.751674Z digest=sha256:eeb07dd850bd2d9053409f6f11658158bdb9cc4f14389ab6ff2fa686b0e87a78

Observation 090f84b0-c612-4419-84ef-3452f7ae6132 · outbound

This paper cites Speech recognition for clinical documentation from 1990 to 2018: a systematic review.

WhisperFlow: speech foundation models in real time Speech recognition for clinical documentation from 1990 to 2018: a systematic review

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.913193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.755262Z digest=sha256:f172da55ada648dc05b9bdf91b93b6780973e0488f28e1c2e7ec580421b52845

Observation 494cbbb8-5a55-4348-8ca2-80fa349759ec · outbound

This paper cites Language models are few-shot learners.

WhisperFlow: speech foundation models in real time Language models are few-shot learners

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.905237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.757951Z digest=sha256:1787240315623c64595f2e70e26515515d6c63f6727c4c6457ba7867a2d6504d

Observation 094d6cd5-e6f3-48a0-adb3-52eb0e1797be · outbound

This paper cites Audio adversarial examples: Targeted attacks on speech-to-text.

WhisperFlow: speech foundation models in real time Audio adversarial examples: Targeted attacks on speech-to-text

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.896876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.760645Z digest=sha256:0655f9edf655d1063cf5976bd457617cfb7466b52eb7b69d43b1e1f1aaeb4f12

Observation 725460a3-7847-4591-949e-10a6cb67273c · outbound

This paper cites https://github.com/corsix/amx, 2022.

WhisperFlow: speech foundation models in real time https://github.com/corsix/amx, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.888962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.763351Z digest=sha256:7b6b0745a9db42183b69227de8d26d2a5bd467735b501482ef70c3632a2f643d

Observation 3877e876-ef45-4874-afa7-83e9d4a5c606 · outbound

This paper cites In 29th USENIX Security Symposium (USENIX Security 20) , pages 2667–2684, 2020.

WhisperFlow: speech foundation models in real time In 29th USENIX Security Symposium (USENIX Security 20) , pages 2667–2684, 2020

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.881328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.766015Z digest=sha256:5550b7dcf11d9b6d26ff1295ee333b279380a18dbf6d86ec41aae95049183fb6

Observation 1e62c57e-2122-4f3f-a686-70b5754b17a6 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech.

WhisperFlow: speech foundation models in real time Fleurs: Few-shot learning evaluation of universal representations of speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.768592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.768592Z digest=sha256:8ebad7c0c6204461972f31156370e056e69d8800603074982e7772e140fb5127

Observation 50cbdccf-acf2-4b36-9cee-f03a241e06f0 · outbound

This paper cites Automatic recognition of spoken digits.

WhisperFlow: speech foundation models in real time Automatic recognition of spoken digits

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.869693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.771062Z digest=sha256:7328e2d7231d67d66361912bd5599ea4c10320c6b460c7fb690dadc69cec2f35

Observation e23ba16d-0950-4668-a35b-636f32acc0eb · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

WhisperFlow: speech foundation models in real time Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.862092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.774604Z digest=sha256:31db2de58eaf8a71631fa57899b5c739f153f407733fc7618a2be935c1529612

Observation 2040508e-5056-4dd7-9985-87d990395ede · outbound

This paper cites Speculative decoding for 2x faster whisper inference, 2023.

WhisperFlow: speech foundation models in real time Speculative decoding for 2x faster whisper inference, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.854362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.777321Z digest=sha256:15ff126c9ddeef72259cf6dc5ae2ad31d5d7983a20a08dfa821bff053bafe1bf

Observation 1ac83dde-2c25-45d3-b37c-06b8261f611e · outbound

This paper cites ggerganov/llama.cpp, 2022.

WhisperFlow: speech foundation models in real time ggerganov/llama.cpp, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.845564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.780439Z digest=sha256:f0f810b0ef27ccc46c3d96e131496dd5546d6593b09db906c0e71c084ae277d8

Observation 26fb2ca9-a17f-4164-a9f1-041b2550ceb3 · outbound

This paper cites ggerganov/whisper.cpp, 2022.

WhisperFlow: speech foundation models in real time ggerganov/whisper.cpp, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.838523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.783040Z digest=sha256:0f70cb208cec8c5c1dac30ec124c36a592a4b37bb9113fa2c27374b60699e7bc

Observation e82e09b5-dc9e-45b2-9b86-ac3f1a5d35df · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

WhisperFlow: speech foundation models in real time Sequence Transduction with Recurrent Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.785464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.785464Z digest=sha256:05695d937b4afcf6a5d58dad5cc18e577f21f13beaa3f7766ed8598949c957e3

Observation 468ba433-b7d5-4625-9837-42621a8a68e5 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

WhisperFlow: speech foundation models in real time Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.788464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.788464Z digest=sha256:0e67d1ade1af313d4165619b7a3d4414f0326405e9c7d9de05da5a85518d1940

Observation bef0cb85-d03c-4b74-af88-aaaacca45f7f · outbound

This paper cites MLX: Efficient and flexible machine learning on apple silicon, 2023.

WhisperFlow: speech foundation models in real time MLX: Efficient and flexible machine learning on apple silicon, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.829653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.791179Z digest=sha256:01ed4761f4adb87077b2012aaa3532977b1f59d6fbc8691cee0664788245dfc4

Observation 645054e9-16f2-4b42-9754-79d1b3eac5a9 · outbound

This paper cites Speech understanding systems: Summary of results of the five-year research effort, 1976.

WhisperFlow: speech foundation models in real time Speech understanding systems: Summary of results of the five-year research effort, 1976

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.821807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.793723Z digest=sha256:cda0353cfd1367edeb11992d630d270ac19e23436912a537c15144aa2f9b2fe6

Observation 20530a75-8258-4a5c-9c6a-e56172710f77 · outbound

This paper cites Streaming end-to-end speech recognition for mobile devices.

WhisperFlow: speech foundation models in real time Streaming end-to-end speech recognition for mobile devices

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.814296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.796197Z digest=sha256:717ce138b86df5c3da0960c233779e7a670b21e792b0ad8fbd784f5acbdd0dd2

Observation 2344ecb9-9d73-415c-a5be-0642d983e51b · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation.

WhisperFlow: speech foundation models in real time Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.805651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.798672Z digest=sha256:f40a2530141ee18e1dfab5a5c0c6f11b2d2dcc3d570c4ddded8d951139b8d125

Observation c88a3b4b-3859-48d0-9a3c-40afca78a3ff · outbound

This paper cites Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM.

WhisperFlow: speech foundation models in real time Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.801247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.801247Z digest=sha256:1b5d47aabafb6f6ed393fb19fa99428fa41abc53f152df981910c0877a99a61f

Observation e55ac934-d788-4dcf-988d-6884d4947301 · outbound

This paper cites Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping.

WhisperFlow: speech foundation models in real time Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.804177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.804177Z digest=sha256:94f010a89285afb345a283415e7f03f93db8b05dd476fda13c5550f464257e97

Observation 9401c046-5b23-430e-af94-ab6774ffb69f · outbound

This paper cites Deepum: Tensor migration and prefetching in unified memory.

WhisperFlow: speech foundation models in real time Deepum: Tensor migration and prefetching in unified memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.806689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.806689Z digest=sha256:d9fb1f2ddf116884b495e3e1387ccf5982c789bce21cb73e06f9a80dcee92f0c

Observation 091c9561-37b9-4a0e-90de-d888f886085b · outbound

This paper cites Speech and language processing, 2000.

WhisperFlow: speech foundation models in real time Speech and language processing, 2000

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.797821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.809268Z digest=sha256:f61f6e40c71ff0bcff7ffde6654d6a5dadcac0bca0218811636d46b2499090dc

Observation 2a4a9b6a-28c3-4fe8-bff0-2a03ec80140a · outbound

This paper cites Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model.

WhisperFlow: speech foundation models in real time Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.811984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.811984Z digest=sha256:e218af8e947634129ad986a6b5f8c7c1af331e3b476c24c56832d8a5a45b966b

Observation af25c6a9-e662-4d0c-b7ab-f308a048cb80 · outbound

This paper cites Scaling Laws for Neural Language Models.

WhisperFlow: speech foundation models in real time Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.815931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.815931Z digest=sha256:e7789c104ede6a73538dfe45f449fed4494db8093f9e4c14a4d34241b61a8d57

Observation a8b0faf0-90df-446a-9a85-f6e0d22b16b4 · outbound

This paper cites Convolution-augmented parameter-efficient fine-tuning for speech recognition.

WhisperFlow: speech foundation models in real time Convolution-augmented parameter-efficient fine-tuning for speech recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.789991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.819593Z digest=sha256:4fcb5a9a5c05daf088b261f1d7eb22f1a8a5e63a91005a24145dbdcb1f413796

Observation 8133820a-9a97-48b7-a7b1-341eb3ebde10 · outbound

This paper cites Speculative Decoding with Big Little Decoder.

WhisperFlow: speech foundation models in real time Speculative Decoding with Big Little Decoder

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.822034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.822034Z digest=sha256:3c36aa3059fd5e13a9e71dfa8d4a6d66fc6bd9afd35a390ae8967e3733214205

Observation 3cbfb0bb-9383-4450-a179-4372dfc16d32 · outbound

This paper cites Low-latency sequence-to- sequence speech recognition and translation by partial hypothesis selection.

WhisperFlow: speech foundation models in real time Low-latency sequence-to- sequence speech recognition and translation by partial hypothesis selection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.781953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.825115Z digest=sha256:857e588cd28288f4d81cef0b9980b845fdd3219835f4f8113398b2d958de81c5

Observation c0e2a732-22d5-4cbc-95cf-eebe0733d498 · outbound

This paper cites Turning Whisper into Real-Time Transcription System.

WhisperFlow: speech foundation models in real time Turning Whisper into Real-Time Transcription System

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.828381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.828381Z digest=sha256:ab15efe5497ee8976e53cd3b09db183ae429b2fa0bdfb685da0afe7435e71efa

Observation e04c2db0-c5c3-4d73-96f9-8ca1ec3b96b6 · outbound

This paper cites Streaming automatic speech recog- nition with the transformer model.

WhisperFlow: speech foundation models in real time Streaming automatic speech recog- nition with the transformer model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.772640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.831725Z digest=sha256:d427016ab724d6ca83cd5447abd27f8debcdb08319e18fe462b08ae8a1597218

Observation 76335910-b08d-4256-9697-772306a91042 · outbound

This paper cites Universal Adversarial Perturbations for Speech Recognition Systems.

WhisperFlow: speech foundation models in real time Universal Adversarial Perturbations for Speech Recognition Systems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.834442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.834442Z digest=sha256:54621376d95423ef999dde5f2cef8063d21035d6acbb02615f935e37d86e1b0f

Observation 00224c54-9fb2-4696-b3fd-d2add7375363 · outbound

This paper cites There is more than one kind of robustness: Fooling Whisper with adversarial examples.

WhisperFlow: speech foundation models in real time There is more than one kind of robustness: Fooling Whisper with adversarial examples

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.837979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.837979Z digest=sha256:b5e5a3dc74deadb5273ceeb358f2feec9d87af2eb4655de27782f197de29434a

Observation b053425f-17b8-482c-aebc-27494d4481a5 · outbound

This paper cites Train- ing language models to follow instructions with human feedback.

WhisperFlow: speech foundation models in real time Train- ing language models to follow instructions with human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.764019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.841017Z digest=sha256:70573d68eb1b28fdbdc4cee92f0fcda85b0ede70fe3585b2e1a704488d53e900

Observation f28c8d69-25ac-4a81-a9ed-dec47100c045 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books.

WhisperFlow: speech foundation models in real time Lib- rispeech: an asr corpus based on public domain audio books

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.756208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.843498Z digest=sha256:599bf298ef2ee3f411d491eb826ebe9915af2df7f4aaa66dd269dc31527dbb8a

Observation 6b295caa-a991-45bb-a78a-0597aaed94e0 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

WhisperFlow: speech foundation models in real time Splitwise: Efficient generative llm inference using phase splitting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.846091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.846091Z digest=sha256:80380d710391ca93714e6b597714f5cbad9d3e8102e6fe4f435a8edf03ada3f0

Observation 698b7f08-3809-4cb9-aba8-0c5b0a5ee70a · outbound

This paper cites Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding.

WhisperFlow: speech foundation models in real time Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.742539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.848604Z digest=sha256:ca0acb22457dc721bad62d275cfebb67444896e4e36cbcd7e68005771e9ec91b

Observation b75de65d-c04d-42fb-9686-e1c49c643009 · outbound

This paper cites OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification.

WhisperFlow: speech foundation models in real time OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.851647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.851647Z digest=sha256:8db4f1466c8d5c7bc83b0e9f9ef37c7a8fc28a0b43a76c4654ecbe92b95779a3

Observation 138f80b3-879a-439e-b263-28b3ba8a0e7f · outbound

This paper cites OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer.

WhisperFlow: speech foundation models in real time OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.855407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.855407Z digest=sha256:109934ab04a14ebc0b9ccd8288833c334ac9a4c49b080b6b3ddf1ea8dfbbc898

Observation 98f7d38c-4830-41c1-913c-6c03fff58770 · outbound

This paper cites Reproducing whisper-style training using an open-source toolkit and publicly available data.

WhisperFlow: speech foundation models in real time Reproducing whisper-style training using an open-source toolkit and publicly available data

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.733664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.858375Z digest=sha256:3405458077c21c725089aa28c87e20ca7cc45e44f4c4d2d14ad70f0a902854ee

Observation d94484bb-c132-4fd9-b332-6f766f875476 · outbound

This paper cites Speech percep- tion at the interface of neurobiology and linguistics.

WhisperFlow: speech foundation models in real time Speech percep- tion at the interface of neurobiology and linguistics

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.724628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.860881Z digest=sha256:8301a7d7a472fa41cd71d963c01537bf4c109e8ee1bb8d236ffc05e8c64ccd5a

Observation 251e98f7-4a70-411e-b3c8-3de52bddfbf9 · outbound

This paper cites Improving language understanding by generative pre-training.

WhisperFlow: speech foundation models in real time Improving language understanding by generative pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.863590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.863590Z digest=sha256:a755b6e0efef9b9d394841a4cf97223134b7cc4a34be50a3f414dbf0796966a2

Observation 6cb8518c-9b3a-4ebd-81fb-8dd3c92faf3f · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

WhisperFlow: speech foundation models in real time Robust speech recognition via large-scale weak supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.866274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.866274Z digest=sha256:541597458270741e3e7ca0dbb5d624a7bc83f1a6872a5646d5b6c4ccd2930863

Observation af9b2efb-3254-4f00-91fc-7315dbf3a9db · outbound

This paper cites Language models are unsupervised multitask learners.

WhisperFlow: speech foundation models in real time Language models are unsupervised multitask learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.868942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.868942Z digest=sha256:74d2ed600cb4f9f3d3298d68892c5b62b6d999baddc4246c022ae89869cdfd7b

Observation 07330927-76d8-4dae-a04f-5db36bfcee3a · outbound

This paper cites Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models.

WhisperFlow: speech foundation models in real time Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.874444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.874444Z digest=sha256:4d63d51d300d3e914e658ea99fe5eb59e4f65faa200651c3a4b37aff77bec7f8

Observation 392bc3de-f999-4476-b06d-b3ef08c5a579 · outbound

This paper cites Muting whisper: A universal acoustic adversarial attack on speech foundation models.

WhisperFlow: speech foundation models in real time Muting whisper: A universal acoustic adversarial attack on speech foundation models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.877075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.877075Z digest=sha256:9414e88df4923811c57e9e9af809b4728aad2156efc1309cb62af451999f0e67

Observation 006ed930-ce33-42b2-9c7b-ad34ce27e3dc · outbound

This paper cites Paul Robinson, and Bradley S.

WhisperFlow: speech foundation models in real time Paul Robinson, and Bradley S

Reference 52

Resolution
verified exact
doi, observed 2026-08-11T15:11:57.982976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.879758Z digest=sha256:83ca77a696e45acc494e6ef77a949c534a64e4eada4757485bdbff6865bd0aae

Observation c359b220-7008-4c8f-be90-4b7f55c4234c · outbound

This paper cites Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer.

WhisperFlow: speech foundation models in real time Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.704061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.882478Z digest=sha256:f357c6fd03b5b0617935c62998ee4ad795a7d1963d76bff23b9e5036b3c7ca07

Observation 93db44f8-3349-4a37-86e6-d80025cece7c · outbound

This paper cites ZeRO-Offload: De- mocratizing Billion-Scale model training.

WhisperFlow: speech foundation models in real time ZeRO-Offload: De- mocratizing Billion-Scale model training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.695451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.885034Z digest=sha256:71a7ac4a1094929eb203860950cc62871c22841978ce062e025a160c4566491b

Observation 1a6f2e7e-3144-4ebd-98a7-85a1084640b2 · outbound

This paper cites Enhancing the ted-lium corpus with selected data for language modeling and more ted talks.

WhisperFlow: speech foundation models in real time Enhancing the ted-lium corpus with selected data for language modeling and more ted talks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.687304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.887689Z digest=sha256:0fec7b4d1d2ab05974fbca776e0306d5b3dc89cefb2aebfcdff2c742b1e9a0d8

Observation 34169d80-0d54-47e2-a1af-2b5148244e3a · outbound

This paper cites Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding.

WhisperFlow: speech foundation models in real time Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.890946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.890946Z digest=sha256:bd2cc39c4ae78b5affdeeeba11608334e73994cbe5c89af115275e4b7810a40c

Observation 7eb371e8-3569-4b44-93e6-ce397c7d49fe · outbound

This paper cites Review of speech-to-text recognition technology for enhancing learning.

WhisperFlow: speech foundation models in real time Review of speech-to-text recognition technology for enhancing learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.679625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.893794Z digest=sha256:b9ed2e51ea4c5f21687acf9760636e477118b6921454263cfd4b3283a9949b5d

Observation ff1dba35-b561-490c-8c76-5ade1e72c3a1 · outbound

This paper cites Dissecting User-Perceived Latency of On-Device E2E Speech Recognition.

WhisperFlow: speech foundation models in real time Dissecting User-Perceived Latency of On-Device E2E Speech Recognition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.896422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.896422Z digest=sha256:3397702a15fc3b33166806848879a1d95d3bc0a1938d1c9a3dc682e06fb082be

Observation ada30cfe-0565-4506-b4a1-d835bfc7a6b6 · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

WhisperFlow: speech foundation models in real time PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.899379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.899379Z digest=sha256:6d7a52a87e731e858f8fb32541d95d68bd138a07e0c5e849da7c1be83465e56f

Observation b63a376d-5d78-4531-94ae-2497a4566db1 · outbound

This paper cites Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding.

WhisperFlow: speech foundation models in real time Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.904845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.904845Z digest=sha256:b196cc277f544e53e71129f69e859c14754d24e81948fd89dd332e9c38df46a3

Observation 73da8f7c-9092-44f6-8757-4afae33bd75c · outbound

This paper cites Intriguing properties of neural networks.

WhisperFlow: speech foundation models in real time Intriguing properties of neural networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.907702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.907702Z digest=sha256:6e527951d659ceec1dda94883110ceccee4efa14783a7d558f24338249c71db2

Observation 4fff794c-ac5e-42d5-b128-04ba5146e1a8 · outbound

This paper cites Streaming trans- former asr with blockwise synchronous beam search.

WhisperFlow: speech foundation models in real time Streaming trans- former asr with blockwise synchronous beam search

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-11T15:11:57.910917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.910917Z digest=sha256:58294369e34e066bb004fc6a2b35ccd5cb4761e223c4618a84a8bd0c3b87bada

Observation 991113f4-2e3a-40c0-b0c3-38b0e68e9f90 · outbound

This paper cites The calo meeting speech recognition and understanding system.

WhisperFlow: speech foundation models in real time The calo meeting speech recognition and understanding system

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.671050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.913485Z digest=sha256:16166bfa77e6f95ad1462557c7042ab60e189baa605a6a095a80b7954bc4d5b5

Observation 3dc58165-3aab-40f6-a3b2-0a83fc81a30f · outbound

This paper cites Attention is all you need.

WhisperFlow: speech foundation models in real time Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.916319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.916319Z digest=sha256:3573c55591567304e662947c1f135f54ad5e9d0297870db5194381df92509047

Observation d6fff2f0-5bbf-43a4-b963-9db803f49ac5 · outbound

This paper cites Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection.

WhisperFlow: speech foundation models in real time Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.919068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.919068Z digest=sha256:f28f4be8473cdddba99cca689b15e2f070dbf7f7a045b53a59a041a558390375

Observation 03387295-7265-473b-821a-24b96bc67d39 · outbound

This paper cites Turbocharge speech understanding with pilot inference.

WhisperFlow: speech foundation models in real time Turbocharge speech understanding with pilot inference

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-11T15:11:57.921997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.921997Z digest=sha256:5aa75d17945db342444a9ff18a9de38ce224f2046e9a6029f27276ef2e7bf9f8

Observation 03455b10-86cd-4b1a-ad0f-9a48a940dd5f · outbound

This paper cites ESPnet: End-to- end speech processing toolkit.

WhisperFlow: speech foundation models in real time ESPnet: End-to- end speech processing toolkit

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.658540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.924604Z digest=sha256:4e409d1da200c31fb4eed6c094045cf74fd04138e40713545737906d7ca57f72

Observation 42c84d03-9e8b-4e7b-a97f-b066a9915e7a · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism.

WhisperFlow: speech foundation models in real time Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.649972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.930867Z digest=sha256:b40c6e77bc3b1ea7497e9d575d74d0dcb341218c0530d28be1110f6d07c799a9

Observation f28cfe96-0ce5-4b57-be1b-aa55c22217d2 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

WhisperFlow: speech foundation models in real time Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.936653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.936653Z digest=sha256:e8b8a1b1dfd4a12e19224e1feae1ff7d51bfb0f6728661b07d88d60ccb1f123c

Observation 5a5b0c0d-2af8-4d98-8484-c827f9c37b3f · outbound

This paper cites Toward human parity in con- versational speech recognition.

WhisperFlow: speech foundation models in real time Toward human parity in con- versational speech recognition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.641768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.940291Z digest=sha256:639a518f46329b80a51fba70af105fffbae1a97b755abc49876bffb18c632a49

Observation 0f39021f-fe3e-452b-af6b-d3d1fd2dd17c · outbound

This paper cites Fast On-device LLM Inference with NPUs.

WhisperFlow: speech foundation models in real time Fast On-device LLM Inference with NPUs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.943619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.943619Z digest=sha256:bf6461e64d8879aed4a61b58b5152a20fe7c480c543bf4f2235ec4c851b1d359

Observation ad35efc5-74ed-4041-9ed5-32b43527ef68 · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

WhisperFlow: speech foundation models in real time Inference with Reference: Lossless Acceleration of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.946433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.946433Z digest=sha256:2a2913afb940552bd1b92af1605d18b05d7c1e07ce7e361cde04970098c26765

Observation ad388ebc-0738-4f51-a6b9-13e26e4f2cbe · outbound

This paper cites Transformer transducer: A streamable speech recog- nition model with transformer encoders and rnn-t loss.

WhisperFlow: speech foundation models in real time Transformer transducer: A streamable speech recog- nition model with transformer encoders and rnn-t loss

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.949322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.949322Z digest=sha256:71140ccb9134d8e95c39336cbdc6f9aec41902695b27f3c6d1df09c9beaf59c1

Observation 25c14384-0ad6-4aa9-bafd-b20256e9b58d · outbound

This paper cites Black-box adversarial attacks on commercial speech platforms with minimal information.

WhisperFlow: speech foundation models in real time Black-box adversarial attacks on commercial speech platforms with minimal information

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.633008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.951868Z digest=sha256:34fbb5231c1b55da240f9f2ff8ba192f39d51a0e01506398fcef5f010f3f5580

Observation 85c56d45-b835-4286-8d99-be546aa2d969 · outbound

This paper cites In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , pages 193–210, 2024.

WhisperFlow: speech foundation models in real time In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , pages 193–210, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.624249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.955428Z digest=sha256:ab73b96f9bb341a15d04a27a77773e13426bea5ab5d1f5e81e832f9d4a74ac21

Observation 6a3efb1f-8404-416b-9c99-cd82ef5f4442 · outbound

This paper cites an unresolved cited work.

WhisperFlow: speech foundation models in real time Unresolved cited work

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.927426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.927426Z digest=sha256:4eb41c620bed7f3cee0414664c0c007714b19939ba52db808e7d04a6811ab1c1

Observation d164851d-0a5a-4698-92eb-9e96d9578ee8 · outbound

This paper cites doi:10.1145/3694715.3695948.

WhisperFlow: speech foundation models in real time doi:10.1145/3694715.3695948

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.933707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.933707Z digest=sha256:6ff53f5549bcbcd2d412612ad572379ba10c2e998ad41f1b4424fd9920703439

Pith citing papers

Observation 818eced5-2f8e-4812-9425-a9adefa861aa · inbound

WhisperRT -- Turning Whisper into a Causal Streaming Model cites this paper.

WhisperRT -- Turning Whisper into a Causal Streaming Model WhisperFlow: speech foundation models in real time

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:21:53.652242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T22:16:51.917336Z digest=sha256:356491f4ff46f852a8f50e4c4ea194ce2076ffa557b2e360ef6368973ba11e86

Observation 4440a6c6-e16d-4a46-98e2-82eba009ed46 · inbound

Sink or SWIM: Tackling Real-Time ASR at Scale cites this paper.

Sink or SWIM: Tackling Real-Time ASR at Scale WhisperFlow: speech foundation models in real time

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:10:53.532913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T12:09:17.087477Z digest=sha256:1bb29e0aee06de65d84687d5f29351ec38f115ca499bc2746f33de0122efa230