Pith. sign in

Paper Citation Record · LEDGER

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

As of 23 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2607.06831.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06831 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T20:25:33.127942Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact11
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d966a59c-cd8c-476f-a0b9-3b35211818d5 · outbound

This paper cites The application of hidden Markov models in speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs The application of hidden Markov models in speech recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.749417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:c7c34cb47d78811d4f75ac000710fe1b853444f41390c43cd77e8e3d479bd7f6

Observation d45345d8-d7fe-4b76-9a97-2b18a0f7cf52 · outbound

This paper cites A tutorial on hidden Markov models and selected applications in speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs A tutorial on hidden Markov models and selected applications in speech recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.732459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:9ea2c515672d609cc73c89eb6a98a27182e82ab6877ac2cefe0db5ca3ac8fcf3

Observation aeb462a3-2d33-4f03-b2df-157037d21e14 · outbound

This paper cites Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.745973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:f5a5186311bef4da0729c829cd7dd94e91e2b602964c20afe7e483aa6209a043

Observation 6f3e1311-fec4-47c0-93b7-664107fb0ff3 · outbound

This paper cites Less peaky and more accurate CTC forced alignment by label priors,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Less peaky and more accurate CTC forced alignment by label priors,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.747709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:6c81289efaa5922fbe82b8e7069a2effcb9505806c3ceee16fa8cbb01a0e9ecd

Observation 50d9969e-337f-4be5-8045-927e882fac6d · outbound

This paper cites Tradition or inno- vation: A comparison of modern asr methods for forced alignment,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Tradition or inno- vation: A comparison of modern asr methods for forced alignment,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.726454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:bf4fdbccbf95f803dd5d4bed4677fa92015906429f559e03761755d87a113b7f

Observation 810712ad-b30a-4540-a290-e964861f9a83 · outbound

This paper cites End-to-end speech recognition: A survey,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs End-to-end speech recognition: A survey,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.742389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:a6fa1c3f65f559433a11552c009954ac4bd87c2314b9fe99ef34ae47200a77e6

Observation 857bdb36-5b2c-4174-a6c7-da2b75041ccf · outbound

This paper cites Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Connection- ist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.734322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:44b6eaf2893478a93d85c3fd5fd45bbb2e81867ab0ddf1c2360ddd106e9d076c

Observation c54e7385-2112-4562-ae87-af83100b8fae · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Sequence Transduction with Recurrent Neural Networks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.593477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:9063e4c4c54fa2407eda7d3ea3c32aa77c148d2b4531debdfde7feb80a1ab82c

Observation f47e5bcf-50fd-4cae-9ac1-d048fd931d7a · outbound

This paper cites Attention-based models for speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Attention-based models for speech recognition,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.760103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:803157e01003ba41a999253a182ea8a5cfe52065631ff99eb14021c576fafd0e

Observation 3a4df90a-b69d-4449-939a-776b5f7ec512 · outbound

This paper cites Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.758483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:bb9ca24e0a7ae50e9625480a1ade2e54d7ad2d1e92f75157247a133b459fec30

Observation 68e4eecd-db03-4145-b76c-bb713c403e0d · outbound

This paper cites Improved training of end-to- end attention models for speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Improved training of end-to- end attention models for speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.736486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:785358174ddc9b56432d0d53a7130e46927d657bed250181d06724a64f7a57c6

Observation 9c46d745-96f8-469c-91f5-b5f7f68023dc · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.728620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:33241c81d5455079dc7c69f8bf0d0ecf87520d980d695925cc7cffec28294df8

Observation 7ce561c0-0674-4ab2-b590-4b95fa456615 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.581542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:bf41b5f501d6aab54c681937e1cb700eeade16854cd9d5591809916a68c16e7c

Observation 399b1390-7481-4696-a009-023215cc1122 · outbound

This paper cites LLMs and Speech: Integration vs. Combination.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs LLMs and Speech: Integration vs. Combination

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.588062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:95e3a140164fcb23078ab1a3561b904d7a1632cc2519d09333ad7714de568793

Observation ae15985d-1df6-4eb3-bf05-092468164318 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Robust speech recognition via large-scale weak supervision,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.781600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:98ab7d166f7233daaca45ccdd67c955195d1eedf6a06efe35fc734bf96c830f1

Observation 60c69c61-7a2a-4850-a787-41f8ad045c9d · outbound

This paper cites WhisperX: Time-accurate speech transcription of long-form audio,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs WhisperX: Time-accurate speech transcription of long-form audio,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.754796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:133e4b5767bf67ab788a14ff439d3ff437649e60b1f665f8516e8d92d90aee2e

Observation b199a002-115e-49f6-bdbb-7ae779d26b93 · outbound

This paper cites CrisperWhisper: Accurate timestamps on verbatim speech transcriptions,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs CrisperWhisper: Accurate timestamps on verbatim speech transcriptions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.792339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:8fd9b706e2638524de2f89d9704b48c4dffc880fdfd8b664ae2f24d522d4606b

Observation f96fb2a6-db40-4145-a810-8a5fed11c603 · outbound

This paper cites Whisper has an internal word aligner,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Whisper has an internal word aligner,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.788911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:5c6012a24ffebe755354a9467b92ad457b7ca8bad076fd2c463a9ab00790c6e2

Observation fb35d85a-b9c4-4f0e-98b6-f83dc377ac08 · outbound

This paper cites Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.605092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:442761423bbaca87b601e25aa213ed6eac8d3b770190a6e4ecbd6ae185abc4dd

Observation db745102-9d9f-4963-8f66-cf92f78b966d · outbound

This paper cites DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.584712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:8503b3f915f975a0b6dd8e06a4abe681adb99b324ff833b16e779449997816a2

Observation 8fa22560-df82-4b40-b451-dfd3f57a013f · outbound

This paper cites Available: https://arxiv.org/abs/2601.18220.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Available: https://arxiv.org/abs/2601.18220

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-10T20:27:36.599158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:af605d0457ae15e94f2b918cb8639e0f92eeb1be10728d7313d5cd158b2040f7

Observation 46faea88-3537-461b-89de-0bdc9534c20e · outbound

This paper cites Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.610300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:60349d7194ab9776f4d6c5abd7d96593ba4d0aa17ab9383b5b89c3cbc303759a

Observation b5413ae3-61ac-42fb-8423-5534e301fb5b · outbound

This paper cites Right Label Context in End-to-End Training of Time-Synchronous ASR Models.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Right Label Context in End-to-End Training of Time-Synchronous ASR Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.590845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:ba16b228493c18b1676e92e13c957463fa8b0d0dbfbdf624568291f459e72009

Observation 14c126d0-ee23-44f9-996c-8862d3fdc5e3 · outbound

This paper cites Saliency-driven word alignment interpretation for neural machine translation,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Saliency-driven word alignment interpretation for neural machine translation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.756579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:96eb996e48a95a60c5ff5735cf24f4c16af92c03d27c6fafbc174188624d704e

Observation 80b417cf-69a4-4c7a-8385-efd91f5a14d2 · outbound

This paper cites Joint CTC/attention decoding for end-to-end speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Joint CTC/attention decoding for end-to-end speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.786985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:d90327a3ff39147ba6a2bddb9f487b24aa0c3605755509c1fa5c64f7e083fe44

Observation 2b14ce76-a4ce-442e-a644-f717c54a511d · outbound

This paper cites Scaling speech tech- nology to 1,000+ languages,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Scaling speech tech- nology to 1,000+ languages,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.785158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:c1bb3e38710847cbe8174fc7d35b76fa7faf57f609593d395cbef2e47bb72c76

Observation 141c56c5-d111-482b-ad7d-12872cc90159 · outbound

This paper cites TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.724159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:0ba05218655c4d1f0f7b41ece14f279b8b44555a714c2083e40b118809de3601

Observation 7e474b36-d535-4835-af3f-56cce3229774 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.765630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:be082dbe2c7d5b7f5718dcaab7be000129a5a4969e4bfb7180b6f9a1752026e1

Observation 909c4e4a-76a3-4b23-afcd-9256897df7fd · outbound

This paper cites Automatic phoneme recognition on TIMIT dataset with Wav2Vec 2.0,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Automatic phoneme recognition on TIMIT dataset with Wav2Vec 2.0,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.769218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:33495e0aa8921c1d7f1e678b3989b5922fb8280e1380c37d6ec05fdc3f618948

Observation faa55225-00b2-46c4-bc19-88ece0e60d4c · outbound

This paper cites XLS-R: Self-supervised cross-lingual speech representation learning at scale,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs XLS-R: Self-supervised cross-lingual speech representation learning at scale,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.776543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:f8b8a6b9ba0d9761e92177594c40084366ad6e810977fb1bcd2a8532ee3c57dc

Observation f8a77ee5-878e-41e4-9de1-0f4b186747bb · outbound

This paper cites Fast conformer with linearly scalable attention for efficient speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Fast conformer with linearly scalable attention for efficient speech recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.783272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:34506700df238633ec893ef399da8a4b49241b15a5e28ae696823e51dc0acaef

Observation e61fd5aa-ba69-4030-87ff-2ca320b7a803 · outbound

This paper cites OWSM v4: Improving open whisper-style speech models via data scaling and cleaning,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs OWSM v4: Improving open whisper-style speech models via data scaling and cleaning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.790646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:b99acfc2c672a67218b4d31ad4331f7f7c6cd647493f44ef3d004066f61f73dd

Observation 192f1e8d-8680-45ed-9418-f3228e4196ea · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.767507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:0fc019bf2311e0188f6973a0d7cac846bd8bf1ce135c35be29a64d6bc96ee5c6

Observation a9eaa9a3-1064-4281-b3d4-2894d73a38a3 · outbound

This paper cites Stateful conformer with cache-based inference for streaming automatic speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Stateful conformer with cache-based inference for streaming automatic speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.770999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:d23eb660590fd483d56eeb14974d5334ba5ad5620ee5c823bfb50a5e802f6d86

Observation 6c777291-1348-41e2-85d0-09d8e82d52cf · outbound

This paper cites Efficient sequence transduction by jointly predicting tokens and durations,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Efficient sequence transduction by jointly predicting tokens and durations,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.763726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:e6a5b5b848abb96b1532013f4c6f7de02c32b53282143a2283a69b6d705a0c80

Observation 2b966bde-edd6-4365-83b5-d13f50453e29 · outbound

This paper cites Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.753093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:19e14b8d9da17be2de451fa8ab5f6e38e24c544286c87e505ac71f47a7fc31eb

Observation bfa1780e-6102-4e67-b507-77b61a9746c3 · outbound

This paper cites TorchAudio: Building blocks for audio and speech processing,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs TorchAudio: Building blocks for audio and speech processing,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.778265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:4952d676963b74d04777ea0a79551e403a3e755de0c48bd3201fdfd1be78b7e2

Observation 8d6223ec-4015-456d-ad38-511db9c02f77 · outbound

This paper cites OWLS: Scaling laws for multilingual speech recognition and translation models,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs OWLS: Scaling laws for multilingual speech recognition and translation models,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.762010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:544d8e798c6aef776ba4f1b6a3f0fb584f5cc890496d394be9dd8173a475275e

Observation 3db9f10c-f72e-4a63-802f-e476e6864eec · outbound

This paper cites Voxtral.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Voxtral

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.607654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:08938ef71fa2a16fa0a5aab71124789f90db81896ada1222ea9acc9fc11d5497

Observation 0f200d49-b50b-44bd-b3c5-d4edadfefd4d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.602301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:67584536342780792282ce6489fe4a5306a8e2e518994af9038e84c206fa1324

Observation 2d604f11-08f7-4c62-86ce-c490b9e6441f · outbound

This paper cites Less is more: Accurate speech recognition & translation without web-scale data,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Less is more: Accurate speech recognition & translation without web-scale data,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.740428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:b2ec63b9a34c16b7fd99ee74d1755274e306499178569090e0abb6dda85ade04

Observation 5f231df0-a75d-4f6e-885a-dee428eb86c7 · outbound

This paper cites TIMIT acoustic-phonetic continuous speech corpus,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs TIMIT acoustic-phonetic continuous speech corpus,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.779980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:d7f1712ec0610961bfa1a50cdede1181671e1a4773097bc2d4e8c8a5058b7822

Observation 84564faf-fd3a-448e-abb3-4dfb050d909e · outbound

This paper cites The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.744190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:ee328538aeb9bde3b7a9ae1e0921d0cbb91d998b283442a5c7513e62f06c73a0

Observation cc9ab90d-6d58-4ff2-8ea6-9c477c6d02dd · outbound

This paper cites Learning important features through propagating activation differences,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Learning important features through propagating activation differences,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.738421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:8977ea1bddecdefde6415d5e0548eb864c4cda2eb1f5a9d5a9bc26d070de2115

Observation d7e1feb3-8b81-4716-a2eb-1b8bbd2bf388 · outbound

This paper cites Towards better understanding of gradient-based attribution methods for deep neural networks,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Towards better understanding of gradient-based attribution methods for deep neural networks,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.751310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:d90af4c115ba07c82e713956796b31f462c523891ec8ef881a330ec5906ba102

Observation 2290b8b4-2ec2-4619-9cca-56ae8b26d3dc · outbound

This paper cites SmoothGrad: removing noise by adding noise.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs SmoothGrad: removing noise by adding noise

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:27:36.596077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:bac95a28a5df7d01defe2da9de04481e820eb888627b34bd079f5a4c84a4afa7

Observation 5ba70317-85aa-43c0-a86f-3db544ec09e8 · outbound

This paper cites Sanity checks for saliency maps,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Sanity checks for saliency maps,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.772994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:f9689afe765b9a073f0c38140bd257552d3a5e65f86f37a54f77fd69b85b8795

Observation 41457776-1484-4b60-adca-67d6ee3eee18 · outbound

This paper cites Available: https://proceedings.neurips.cc/paper files/ paper/2018/file/294a8ed24b1ad22ec2e7efea049b8737-Paper.pdf.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Available: https://proceedings.neurips.cc/paper files/ paper/2018/file/294a8ed24b1ad22ec2e7efea049b8737-Paper.pdf

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.794010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:47bb36e0434a3b17494e662b570fa901a40f1f637cfac554a6d06c2b81997c9b

Observation c0fb4029-8cc0-4385-9891-18f15ccd52b5 · outbound

This paper cites Axiomatic attribution for deep networks,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Axiomatic attribution for deep networks,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.730491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:119fe750efaf7f93436a359ac38c8a9a0af4100572d8a794086cb90af2f6d03b

Observation bb1834f8-f027-4183-b870-7fb7c94926f4 · outbound

This paper cites Improving performance of deep learning models with axiomatic attribution priors and expected gradients,.

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs Improving performance of deep learning models with axiomatic attribution priors and expected gradients,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T20:27:36.774812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T20:25:33.127942Z digest=sha256:6049baa4364885f8209c70c7a27012118db0dd0dbe6b2acf367aef99f1edd4c7

Pith citing papers

No inbound Pith citation observations are available.