Pith. sign in

Paper Citation Record · LEDGER

WhisperFlow: speech foundation models in real time

As of 16 August 2026, this Paper Citation Record lists 76 of 76 outbound references and 2 inbound Pith citation observations for arXiv:2412.11272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11272 v2

Coverage vector

measured 76 of 76 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:11:57.955428Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T22:16:51.917336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T22:21:53.649506Z

Reference resolution

76 of 76 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved37
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1fee2c3-e370-41a2-9439-93222cf0253f · outbound

This paper cites Accessed: 2024-11-3.

WhisperFlow: speech foundation models in real time Accessed: 2024-11-3

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.936833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.731021Z digest=sha256:1781ae81aa95cf74f3d68118faedcf5ce05874f09a2b0a5ae96777e11960eb94

Observation dc2052ab-9b74-4902-b738-bf19f9a44b52 · outbound

This paper cites GPT-4 Technical Report.

WhisperFlow: speech foundation models in real time GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.735305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.735305Z digest=sha256:03a4279329c0cd1f4cfd4ed4196071f2f312636885fa689f113a21cdb9d0e8c9

Observation a17cf4bb-74a5-464f-ba1c-3fa280105781 · outbound

This paper cites Did you hear that? Adversarial Examples Against Automatic Speech Recognition.

WhisperFlow: speech foundation models in real time Did you hear that? Adversarial Examples Against Automatic Speech Recognition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.738385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.738385Z digest=sha256:82ec8d538b61422df94545631d213c609f6c844114b9bb5acd9fad81b22e39ba

Observation 30577563-e2bb-4df3-8748-a55b7caaaadf · outbound

This paper cites Apple macbook air tech specs, 2024.

WhisperFlow: speech foundation models in real time Apple macbook air tech specs, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.926431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.742310Z digest=sha256:954e59c9f55a6b4f85d17f36554e3c6ad368f4fc7409b5f9e4066a1062dde56e

Observation 7ca7c6ef-9f0b-4613-9aee-f4876768af59 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

WhisperFlow: speech foundation models in real time Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.745362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.745362Z digest=sha256:1ebbf4f949300683907af08234a8a5a1fd9cadc3fc47668e032b67dc134b6e56

Observation 246f45ca-6e4b-40d1-b74b-08e9cbebf41d · outbound

This paper cites an unresolved cited work.

WhisperFlow: speech foundation models in real time Unresolved cited work

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T15:11:58.596857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.748694Z digest=sha256:56497340d545a83640b95181b1b49517d924084ffbde040b84af6a9d64410b75

Observation 62593c14-0bf6-46d0-b9da-25d63e01403c · outbound

This paper cites A mathematical theory of adaptive control processes.

WhisperFlow: speech foundation models in real time A mathematical theory of adaptive control processes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.751674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.751674Z digest=sha256:fcc13c80a6a4f42bc25d8997c92bc957e883b3aad97b528ad9fa5c79cbd0353b

Observation 090f84b0-c612-4419-84ef-3452f7ae6132 · outbound

This paper cites Speech recognition for clinical documentation from 1990 to 2018: a systematic review.

WhisperFlow: speech foundation models in real time Speech recognition for clinical documentation from 1990 to 2018: a systematic review

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.913193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.755262Z digest=sha256:2f65a9054e37f1ace82659678e9fd5232453714110716073ee5bedd346ca7382

Observation 494cbbb8-5a55-4348-8ca2-80fa349759ec · outbound

This paper cites Language models are few-shot learners.

WhisperFlow: speech foundation models in real time Language models are few-shot learners

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.905237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.757951Z digest=sha256:7f34ea51603e1d50b98e94d25d65ef0dddb12431563aae5836775adb5aa18db0

Observation 094d6cd5-e6f3-48a0-adb3-52eb0e1797be · outbound

This paper cites Audio adversarial examples: Targeted attacks on speech-to-text.

WhisperFlow: speech foundation models in real time Audio adversarial examples: Targeted attacks on speech-to-text

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.896876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.760645Z digest=sha256:21f90115df1dcef734114d91437b1fb33d18aed4ad12a115d45194edac8e5094

Observation 725460a3-7847-4591-949e-10a6cb67273c · outbound

This paper cites https://github.com/corsix/amx, 2022.

WhisperFlow: speech foundation models in real time https://github.com/corsix/amx, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.888962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.763351Z digest=sha256:5e09ef2167a50df8653061a76310fda32c9abb6c4382a4c6b89582bd7418ee0a

Observation 3877e876-ef45-4874-afa7-83e9d4a5c606 · outbound

This paper cites In 29th USENIX Security Symposium (USENIX Security 20) , pages 2667–2684, 2020.

WhisperFlow: speech foundation models in real time In 29th USENIX Security Symposium (USENIX Security 20) , pages 2667–2684, 2020

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.881328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.766015Z digest=sha256:d3cbd9eb2388245cd149d6cb05797e6686195c2006fa9e1481f2d92b668199e6

Observation 1e62c57e-2122-4f3f-a686-70b5754b17a6 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech.

WhisperFlow: speech foundation models in real time Fleurs: Few-shot learning evaluation of universal representations of speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.768592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.768592Z digest=sha256:b27f9b24239319a41e6f9b2c9c4f064e087896967cf515b49157b36a025cc63d

Observation 50cbdccf-acf2-4b36-9cee-f03a241e06f0 · outbound

This paper cites Automatic recognition of spoken digits.

WhisperFlow: speech foundation models in real time Automatic recognition of spoken digits

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.869693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.771062Z digest=sha256:fddd2c779411213cfe551b11c68553f0403c49b90f5512f43657d986c7d4f865

Observation e23ba16d-0950-4668-a35b-636f32acc0eb · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

WhisperFlow: speech foundation models in real time Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.862092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.774604Z digest=sha256:7035c2683b011a08c3daf75ed31d882a06f651400e1d6a71116df7406a34d6fc

Observation 2040508e-5056-4dd7-9985-87d990395ede · outbound

This paper cites Speculative decoding for 2x faster whisper inference, 2023.

WhisperFlow: speech foundation models in real time Speculative decoding for 2x faster whisper inference, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.854362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.777321Z digest=sha256:b3101da4b9ab3ac89419a004891db7e838b05fb399c31b9a4c2d4af7253f999d

Observation 1ac83dde-2c25-45d3-b37c-06b8261f611e · outbound

This paper cites ggerganov/llama.cpp, 2022.

WhisperFlow: speech foundation models in real time ggerganov/llama.cpp, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.845564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.780439Z digest=sha256:121788978a738f0825e1e79d87b5034ad0faf40b21aa5b51a77de8f328300441

Observation 26fb2ca9-a17f-4164-a9f1-041b2550ceb3 · outbound

This paper cites ggerganov/whisper.cpp, 2022.

WhisperFlow: speech foundation models in real time ggerganov/whisper.cpp, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.838523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.783040Z digest=sha256:faa556548fcf07cc20aa7ceb6602c27ec25d134acf935377f0297c5b44b2f8b1

Observation e82e09b5-dc9e-45b2-9b86-ac3f1a5d35df · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

WhisperFlow: speech foundation models in real time Sequence Transduction with Recurrent Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.785464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.785464Z digest=sha256:02cb1ddb70f73cd35ad68f1e34fd2f6e802f1edbbaf27bcabc428efaf51a3524

Observation 468ba433-b7d5-4625-9837-42621a8a68e5 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

WhisperFlow: speech foundation models in real time Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.788464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.788464Z digest=sha256:5562aa89227713aea5bfb16e552d568d500222de915062e58070af001a2f487c

Observation bef0cb85-d03c-4b74-af88-aaaacca45f7f · outbound

This paper cites MLX: Efficient and flexible machine learning on apple silicon, 2023.

WhisperFlow: speech foundation models in real time MLX: Efficient and flexible machine learning on apple silicon, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.829653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.791179Z digest=sha256:8879c7e6d485b9a8d295c1653eb314d39b5e0754e06900dad741f235b0e3ed58

Observation 645054e9-16f2-4b42-9754-79d1b3eac5a9 · outbound

This paper cites Speech understanding systems: Summary of results of the five-year research effort, 1976.

WhisperFlow: speech foundation models in real time Speech understanding systems: Summary of results of the five-year research effort, 1976

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.821807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.793723Z digest=sha256:6d9ab7caf3fc3e6f8501d7eaa94d09fadb7cea5af820827de3683b15707cc040

Observation 20530a75-8258-4a5c-9c6a-e56172710f77 · outbound

This paper cites Streaming end-to-end speech recognition for mobile devices.

WhisperFlow: speech foundation models in real time Streaming end-to-end speech recognition for mobile devices

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.814296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.796197Z digest=sha256:e5868aedfcc1f1b7aee81b2363d75a3e34287884368d71b99014d05449559cd1

Observation 2344ecb9-9d73-415c-a5be-0642d983e51b · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation.

WhisperFlow: speech foundation models in real time Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.805651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.798672Z digest=sha256:1f852196a6edc5437c29207dfbb077c14df63f525f8b80a047aa79e9fa7ffb0e

Observation c88a3b4b-3859-48d0-9a3c-40afca78a3ff · outbound

This paper cites Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM.

WhisperFlow: speech foundation models in real time Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.801247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.801247Z digest=sha256:4a3e6f9396bad237c57cafa847570871b0ccaa6e84cededde421ee0f89c8058f

Observation e55ac934-d788-4dcf-988d-6884d4947301 · outbound

This paper cites Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping.

WhisperFlow: speech foundation models in real time Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.804177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.804177Z digest=sha256:11821d1a0b1e1e700dd9440e952c0094bda7e727afe137256dd7960b4ac9617b

Observation 9401c046-5b23-430e-af94-ab6774ffb69f · outbound

This paper cites Deepum: Tensor migration and prefetching in unified memory.

WhisperFlow: speech foundation models in real time Deepum: Tensor migration and prefetching in unified memory

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.806689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.806689Z digest=sha256:8ba6870357a2e1ffc411e94d03be76d812c1a90bd9fc79e09df90cd4e678bdfb

Observation 091c9561-37b9-4a0e-90de-d888f886085b · outbound

This paper cites Speech and language processing, 2000.

WhisperFlow: speech foundation models in real time Speech and language processing, 2000

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.797821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.809268Z digest=sha256:26c47867fa67d8f387b6fe011069fd0909e1dc345b8898085090696dc055fc9d

Observation 2a4a9b6a-28c3-4fe8-bff0-2a03ec80140a · outbound

This paper cites Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model.

WhisperFlow: speech foundation models in real time Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.811984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.811984Z digest=sha256:f68179d722517682942a5a7a99a34821e34925ae3045bb3a01802e315fc5729f

Observation af25c6a9-e662-4d0c-b7ab-f308a048cb80 · outbound

This paper cites Scaling Laws for Neural Language Models.

WhisperFlow: speech foundation models in real time Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.815931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.815931Z digest=sha256:720d7065a7f5455c270234fffb10ea40fe3a538fcb596a8cb1677e1ebc175ef4

Observation a8b0faf0-90df-446a-9a85-f6e0d22b16b4 · outbound

This paper cites Convolution-augmented parameter-efficient fine-tuning for speech recognition.

WhisperFlow: speech foundation models in real time Convolution-augmented parameter-efficient fine-tuning for speech recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.789991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.819593Z digest=sha256:51a5f7483e39d5937d9896d3a88beec529a22d6d50ce265cbe4dc5f5a16960c6

Observation 8133820a-9a97-48b7-a7b1-341eb3ebde10 · outbound

This paper cites Speculative Decoding with Big Little Decoder.

WhisperFlow: speech foundation models in real time Speculative Decoding with Big Little Decoder

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.822034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.822034Z digest=sha256:08afc0a4c1b278dc802a93dfc7f62d4758b95572eb5224be638c26e588153b1a

Observation 3cbfb0bb-9383-4450-a179-4372dfc16d32 · outbound

This paper cites Low-latency sequence-to- sequence speech recognition and translation by partial hypothesis selection.

WhisperFlow: speech foundation models in real time Low-latency sequence-to- sequence speech recognition and translation by partial hypothesis selection

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.781953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.825115Z digest=sha256:8bf44642719d452ec6963d29fff5408c73440e1b9d9eaf12d75f56569b331e96

Observation c0e2a732-22d5-4cbc-95cf-eebe0733d498 · outbound

This paper cites Turning Whisper into Real-Time Transcription System.

WhisperFlow: speech foundation models in real time Turning Whisper into Real-Time Transcription System

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.828381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.828381Z digest=sha256:bea3b692481aadfba3afd2f7633ce5e214f5a77ecbc8e3e1e4a997f3f0773800

Observation e04c2db0-c5c3-4d73-96f9-8ca1ec3b96b6 · outbound

This paper cites Streaming automatic speech recog- nition with the transformer model.

WhisperFlow: speech foundation models in real time Streaming automatic speech recog- nition with the transformer model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.772640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.831725Z digest=sha256:7a98a87377499d1479d453a43b8cded08f4e6f14bfbe19ef7330da5d5f62180d

Observation 76335910-b08d-4256-9697-772306a91042 · outbound

This paper cites Universal Adversarial Perturbations for Speech Recognition Systems.

WhisperFlow: speech foundation models in real time Universal Adversarial Perturbations for Speech Recognition Systems

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.834442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.834442Z digest=sha256:8c5d3c251c06bd02f205d1e37cc037e26a90ef600e3596d75d46814724286dc8

Observation 00224c54-9fb2-4696-b3fd-d2add7375363 · outbound

This paper cites There is more than one kind of robustness: Fooling Whisper with adversarial examples.

WhisperFlow: speech foundation models in real time There is more than one kind of robustness: Fooling Whisper with adversarial examples

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.837979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.837979Z digest=sha256:059a4b4fa69abf076c9672aad93830a14480db455cbe035e7c39b48c5a7144c0

Observation b053425f-17b8-482c-aebc-27494d4481a5 · outbound

This paper cites Train- ing language models to follow instructions with human feedback.

WhisperFlow: speech foundation models in real time Train- ing language models to follow instructions with human feedback

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.764019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.841017Z digest=sha256:3dd75df477e80e759a6b72203e5c943882507022fcce66382238152ace45306f

Observation f28c8d69-25ac-4a81-a9ed-dec47100c045 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books.

WhisperFlow: speech foundation models in real time Lib- rispeech: an asr corpus based on public domain audio books

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.756208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.843498Z digest=sha256:6b4734ab4f604364df66afd88db276356e4d8f3483fc3627e20cd08e10a96cc9

Observation 6b295caa-a991-45bb-a78a-0597aaed94e0 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

WhisperFlow: speech foundation models in real time Splitwise: Efficient generative llm inference using phase splitting

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.846091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.846091Z digest=sha256:0ec978b82d62dd7a1b6be8a4cc04495c94f90be530e86a1f5cb8d8a353025814

Observation 698b7f08-3809-4cb9-aba8-0c5b0a5ee70a · outbound

This paper cites Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding.

WhisperFlow: speech foundation models in real time Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.742539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.848604Z digest=sha256:caeee920f9f744a98b82004fae9796327861d028e51a4761a2520eb5b5b08848

Observation b75de65d-c04d-42fb-9686-e1c49c643009 · outbound

This paper cites OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification.

WhisperFlow: speech foundation models in real time OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.851647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.851647Z digest=sha256:f4939f0f1fa1396dbcbd37b7ff12987200c468f05c3103068cebb6214b7d181f

Observation 138f80b3-879a-439e-b263-28b3ba8a0e7f · outbound

This paper cites OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer.

WhisperFlow: speech foundation models in real time OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.855407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.855407Z digest=sha256:8d44689957aafb34e4e6cd7bbe710a5d5f5a65c7382444ed3b98249b4126bfac

Observation 98f7d38c-4830-41c1-913c-6c03fff58770 · outbound

This paper cites Reproducing whisper-style training using an open-source toolkit and publicly available data.

WhisperFlow: speech foundation models in real time Reproducing whisper-style training using an open-source toolkit and publicly available data

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.733664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.858375Z digest=sha256:a2910b61a158ef7ca8ebbb59f62bcb5e50d9d7aaae336fa6cdaed111f3a1b38c

Observation d94484bb-c132-4fd9-b332-6f766f875476 · outbound

This paper cites Speech percep- tion at the interface of neurobiology and linguistics.

WhisperFlow: speech foundation models in real time Speech percep- tion at the interface of neurobiology and linguistics

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.724628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.860881Z digest=sha256:5c29fc2b9fad0348e33a01910900aff9178b47018bd6e95571c4a34a184aa94f

Observation 251e98f7-4a70-411e-b3c8-3de52bddfbf9 · outbound

This paper cites Improving language understanding by generative pre-training.

WhisperFlow: speech foundation models in real time Improving language understanding by generative pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.863590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.863590Z digest=sha256:181c91475fe3f1469f91d35d9a89796039eed9e74502d8752e19b2bb9466a44d

Observation 6cb8518c-9b3a-4ebd-81fb-8dd3c92faf3f · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

WhisperFlow: speech foundation models in real time Robust speech recognition via large-scale weak supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.866274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.866274Z digest=sha256:d34354800ad8c18bd47bef686e933a535a67c6467c00469e020b083bfd24b415

Observation af9b2efb-3254-4f00-91fc-7315dbf3a9db · outbound

This paper cites Language models are unsupervised multitask learners.

WhisperFlow: speech foundation models in real time Language models are unsupervised multitask learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.868942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.868942Z digest=sha256:c584519c1b4a19385b8b943a8faa513c90e409eb224e1d6d116dd2dbdeffe8e2

Observation 07330927-76d8-4dae-a04f-5db36bfcee3a · outbound

This paper cites Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models.

WhisperFlow: speech foundation models in real time Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.874444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.874444Z digest=sha256:d135ce3df2dd1c722c91ab68df858d376ee5805e6a554e7fead4432c5f6b756e

Observation 392bc3de-f999-4476-b06d-b3ef08c5a579 · outbound

This paper cites Muting whisper: A universal acoustic adversarial attack on speech foundation models.

WhisperFlow: speech foundation models in real time Muting whisper: A universal acoustic adversarial attack on speech foundation models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.877075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.877075Z digest=sha256:df96054c7b27e1b9b5f7c5f2dc370925a64992527619606d73c36a574c289b98

Observation 006ed930-ce33-42b2-9c7b-ad34ce27e3dc · outbound

This paper cites Paul Robinson, and Bradley S.

WhisperFlow: speech foundation models in real time Paul Robinson, and Bradley S

Reference 52

Resolution
verified exact
doi, observed 2026-08-11T15:11:57.982976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.879758Z digest=sha256:e695447df2bafcc92e2dd001291d132db08cfdc55a90f793b78775d047a254eb

Observation c359b220-7008-4c8f-be90-4b7f55c4234c · outbound

This paper cites Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer.

WhisperFlow: speech foundation models in real time Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.704061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.882478Z digest=sha256:467b4abcd895c748e7ff07c93b4bac5c4baf11dafd5dfd8b9394708eef233d18

Observation 93db44f8-3349-4a37-86e6-d80025cece7c · outbound

This paper cites ZeRO-Offload: De- mocratizing Billion-Scale model training.

WhisperFlow: speech foundation models in real time ZeRO-Offload: De- mocratizing Billion-Scale model training

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.695451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.885034Z digest=sha256:74eb1d00f9ac422051f4af27bf9bffb7788b46ccae0b49b9e9ca84eb331daa03

Observation 1a6f2e7e-3144-4ebd-98a7-85a1084640b2 · outbound

This paper cites Enhancing the ted-lium corpus with selected data for language modeling and more ted talks.

WhisperFlow: speech foundation models in real time Enhancing the ted-lium corpus with selected data for language modeling and more ted talks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.687304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.887689Z digest=sha256:6535c31799b4abd6c6d88a1cc97271ed9508791458d20c37af6d044bc3a2c99a

Observation 34169d80-0d54-47e2-a1af-2b5148244e3a · outbound

This paper cites Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding.

WhisperFlow: speech foundation models in real time Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.890946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.890946Z digest=sha256:eff6ed2752547ae9c49a6ba7a12d2e8dbd47500fb5ee0aa54ed8f23bc40a5c73

Observation 7eb371e8-3569-4b44-93e6-ce397c7d49fe · outbound

This paper cites Review of speech-to-text recognition technology for enhancing learning.

WhisperFlow: speech foundation models in real time Review of speech-to-text recognition technology for enhancing learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.679625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.893794Z digest=sha256:9f48b0e328d63b3f96deadd85ef6a90d56caf4c4931b9c8e7dd3178b87d093d5

Observation ff1dba35-b561-490c-8c76-5ade1e72c3a1 · outbound

This paper cites Dissecting User-Perceived Latency of On-Device E2E Speech Recognition.

WhisperFlow: speech foundation models in real time Dissecting User-Perceived Latency of On-Device E2E Speech Recognition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.896422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.896422Z digest=sha256:5932ddc4fc1aa87c81bf9277f90351d0dd88d325fb044b9589dd0acbb891cf51

Observation ada30cfe-0565-4506-b4a1-d835bfc7a6b6 · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

WhisperFlow: speech foundation models in real time PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.899379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.899379Z digest=sha256:fbfa603455e953be67ec4346de73dfbb81ef277e16c3b9aa88d160164e2fc919

Observation b63a376d-5d78-4531-94ae-2497a4566db1 · outbound

This paper cites Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding.

WhisperFlow: speech foundation models in real time Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.904845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.904845Z digest=sha256:ecb2de662c9e236b567340813c62248fcf4da08525e7a1e6f80e7e3a4f8e866d

Observation 73da8f7c-9092-44f6-8757-4afae33bd75c · outbound

This paper cites Intriguing properties of neural networks.

WhisperFlow: speech foundation models in real time Intriguing properties of neural networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.907702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.907702Z digest=sha256:f0f74027afa6615ba2a2786d1eb256eb3d266dc50016d26b1411923ae047ce4a

Observation 4fff794c-ac5e-42d5-b128-04ba5146e1a8 · outbound

This paper cites Streaming trans- former asr with blockwise synchronous beam search.

WhisperFlow: speech foundation models in real time Streaming trans- former asr with blockwise synchronous beam search

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-11T15:11:57.910917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.910917Z digest=sha256:9de96bb5c62cb4fba5b885a07e4f8ea6fcaa45f8ba70864b9bc7dc2f710653cd

Observation 991113f4-2e3a-40c0-b0c3-38b0e68e9f90 · outbound

This paper cites The calo meeting speech recognition and understanding system.

WhisperFlow: speech foundation models in real time The calo meeting speech recognition and understanding system

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.671050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.913485Z digest=sha256:be8a781fce8f36148d5f48d21914e58448478a3b054215120215ef1280c2bc2e

Observation 3dc58165-3aab-40f6-a3b2-0a83fc81a30f · outbound

This paper cites Attention is all you need.

WhisperFlow: speech foundation models in real time Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.916319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.916319Z digest=sha256:d28f3143c15ca9db6d2180b23d6f1e47b3498ccfd6d8bdc9342154535c4df977

Observation d6fff2f0-5bbf-43a4-b963-9db803f49ac5 · outbound

This paper cites Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection.

WhisperFlow: speech foundation models in real time Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.919068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.919068Z digest=sha256:d3808a91dc73a4151f2d64c919ca6ba9a0dd0f44af820134f61ec3ab17c340e5

Observation 03387295-7265-473b-821a-24b96bc67d39 · outbound

This paper cites Turbocharge speech understanding with pilot inference.

WhisperFlow: speech foundation models in real time Turbocharge speech understanding with pilot inference

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-11T15:11:57.921997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.921997Z digest=sha256:e8a7f8c27991cad7bc5088b98f204b4aeddcebbd7bdacdbfb7047699b60d5b84

Observation 03455b10-86cd-4b1a-ad0f-9a48a940dd5f · outbound

This paper cites ESPnet: End-to- end speech processing toolkit.

WhisperFlow: speech foundation models in real time ESPnet: End-to- end speech processing toolkit

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.658540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.924604Z digest=sha256:635c24b3bb93a01f0b79f1af700bf1e4cdff2a6727c6ce8d74be15896ed4c978

Observation 42c84d03-9e8b-4e7b-a97f-b066a9915e7a · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism.

WhisperFlow: speech foundation models in real time Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.649972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.930867Z digest=sha256:9e88ff979bff8a30ab8f6767f4021ad405f27e223c093e343f109f34d4ff167a

Observation f28cfe96-0ce5-4b57-be1b-aa55c22217d2 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

WhisperFlow: speech foundation models in real time Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.936653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.936653Z digest=sha256:ec58b4b3538ffd61a57c3b45cf1c837161073b9a5f2d9e5084f27987da5ec2f3

Observation 5a5b0c0d-2af8-4d98-8484-c827f9c37b3f · outbound

This paper cites Toward human parity in con- versational speech recognition.

WhisperFlow: speech foundation models in real time Toward human parity in con- versational speech recognition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.641768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.940291Z digest=sha256:bdb642268b8c1efdc835ef8c2baf7ca8928b3e765652c4b17d83448298ef5447

Observation 0f39021f-fe3e-452b-af6b-d3d1fd2dd17c · outbound

This paper cites Fast On-device LLM Inference with NPUs.

WhisperFlow: speech foundation models in real time Fast On-device LLM Inference with NPUs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.943619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.943619Z digest=sha256:84d8512dfc1450e159394e5263a4a8c656025de3eee031164ecd6194f47de8fe

Observation ad35efc5-74ed-4041-9ed5-32b43527ef68 · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

WhisperFlow: speech foundation models in real time Inference with Reference: Lossless Acceleration of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.946433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.946433Z digest=sha256:b5d3ecb319e457a1e633bb4bed8489e03b34c387f9d71a289416da0e4cd8d31d

Observation ad388ebc-0738-4f51-a6b9-13e26e4f2cbe · outbound

This paper cites Transformer transducer: A streamable speech recog- nition model with transformer encoders and rnn-t loss.

WhisperFlow: speech foundation models in real time Transformer transducer: A streamable speech recog- nition model with transformer encoders and rnn-t loss

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.949322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.949322Z digest=sha256:cdcb1da344576c146a2c5df154bc6ce4d505a76d80187b18eecf569685837a87

Observation 25c14384-0ad6-4aa9-bafd-b20256e9b58d · outbound

This paper cites Black-box adversarial attacks on commercial speech platforms with minimal information.

WhisperFlow: speech foundation models in real time Black-box adversarial attacks on commercial speech platforms with minimal information

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.633008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.951868Z digest=sha256:ba72783fb0a92de5325b6f6586d3252c20715359fbf30ae968422a3a9068cab5

Observation 85c56d45-b835-4286-8d99-be546aa2d969 · outbound

This paper cites In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , pages 193–210, 2024.

WhisperFlow: speech foundation models in real time In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , pages 193–210, 2024

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:11:58.624249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T15:11:57.955428Z digest=sha256:5f667b8d5e048c2f07878f6e5b213270b7fb56bde8554a736aae09cfb7e58127

Observation 6a3efb1f-8404-416b-9c99-cd82ef5f4442 · outbound

This paper cites an unresolved cited work.

WhisperFlow: speech foundation models in real time Unresolved cited work

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.927426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.927426Z digest=sha256:6edf584692215e33f06b3ed31786c3016d68279b31f376c728a8061ae4ad756d

Observation d164851d-0a5a-4698-92eb-9e96d9578ee8 · outbound

This paper cites doi:10.1145/3694715.3695948.

WhisperFlow: speech foundation models in real time doi:10.1145/3694715.3695948

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.933707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.933707Z digest=sha256:1078523f2e9a9c845d5dbebcfb8b79bfa109a3a6d922ce07ae4f8073a70b4a00

Pith citing papers

Observation 818eced5-2f8e-4812-9425-a9adefa861aa · inbound

WhisperRT -- Turning Whisper into a Causal Streaming Model cites this paper.

WhisperRT -- Turning Whisper into a Causal Streaming Model WhisperFlow: speech foundation models in real time

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:21:53.652242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T22:16:51.917336Z digest=sha256:a2ae8ae5a28921f72e896d1c9d6203a53d8e9d6853da25cca7b0b130423c58b0

Observation 4440a6c6-e16d-4a46-98e2-82eba009ed46 · inbound

Sink or SWIM: Tackling Real-Time ASR at Scale cites this paper.

Sink or SWIM: Tackling Real-Time ASR at Scale WhisperFlow: speech foundation models in real time

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:10:53.532913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T12:09:17.087477Z digest=sha256:3d36adf44ed9c1903cc7e4f80cf2becb5ba1bc601b39d9c2094e63494c502e41