Pith. sign in

Paper Citation Record · LEDGER

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2506.16969.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16969 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:19:59.764201Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:19:59.656455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T19:19:59.861302Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29b7e862-1586-40f1-b092-eaa59ece4617 · outbound

This paper cites State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T19:19:59.866003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.656455Z digest=sha256:548be0a13e06221a00d05859b573e2de465654c3eca3b212e3dbd2576d3806a3

Observation e43acfad-951c-4d1e-9cbd-aa0b1b243dfc · outbound

This paper cites The primary distinction lies in how they are produced.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition The primary distinction lies in how they are produced

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.153712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.661179Z digest=sha256:a64571b846c9c4f718b4c06ac15dbe64b808003b9571160912e2366e147a7989

Observation 72731bd3-fea3-495e-881a-1693122076c0 · outbound

This paper cites Specifically, we introduce the wTIMIT and CHAINS.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Specifically, we introduce the wTIMIT and CHAINS

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.140770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.665829Z digest=sha256:edb67d14bc5428f4a93d49d7f9d4f578352c29404ee0d2ae0e42cc831df1a2d2

Observation 67f7ec60-44c6-4a08-b8cd-6cc40b262965 · outbound

This paper cites As a base- line, we evaluated the performance of the pre-trained Whisper Large-v2 model on the test set to assess the need for a special- ized system for the proposed challenges.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition As a base- line, we evaluated the performance of the pre-trained Whisper Large-v2 model on the test set to assess the need for a special- ized system for the proposed challenges

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:20:00.118240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.673878Z digest=sha256:4d941a54813c75f5541821e121313fdbb0e89046640a66750e2fe95f7e21e047

Observation 85f063d8-b6b9-4e73-99ed-7ab21a4699e6 · outbound

This paper cites These challenges can severely degrade the performance of traditional systems, highlighting the necessity of developing ASR models specifically tailored to handle whispered speech.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition These challenges can severely degrade the performance of traditional systems, highlighting the necessity of developing ASR models specifically tailored to handle whispered speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.105826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.677498Z digest=sha256:752bf45398502583c325a9aa6d7c6e4205090c1ad6a9f2c8f5d21a50b5c5833c

Observation affd1cb9-744f-4efd-a73e-75489f3cc4a2 · outbound

This paper cites Gener- ative models for improved naturalness, intelligibility, and voicing of whispered speech,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Gener- ative models for improved naturalness, intelligibility, and voicing of whispered speech,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.004955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.699284Z digest=sha256:c3d13ee0ef3d13acf30fcb4af6c3a8ab520108c602697d1c5d6d81de142f6742

Observation 738ba118-8258-482e-ad88-17290d9a97ed · outbound

This paper cites Gammatonegram representation for end-to-end dysarthric speech processing tasks: Speech recogni- tion, speaker identification, and intelligibility assessment,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Gammatonegram representation for end-to-end dysarthric speech processing tasks: Speech recogni- tion, speaker identification, and intelligibility assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.092558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.680878Z digest=sha256:ec10871f8dab26fc580a8bf861bf0da4a2672ebfc0423a2dc3f971181a81fec1

Observation 6ad10499-a51d-4611-9320-2e6c29dc32ca · outbound

This paper cites Dysarthric speaker identification with different degrees of dysarthria severity using deep belief networks,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Dysarthric speaker identification with different degrees of dysarthria severity using deep belief networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.078638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.684390Z digest=sha256:24146c36e93478b1411e227a3293208c313a25c2598319b207ff583265376dae

Observation 01904d7e-5fcf-4b9c-831a-a633d837ca66 · outbound

This paper cites Analysis of deep generative model impact on feature extraction and dimension reduction for short ut- terance text-independent speaker verification,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Analysis of deep generative model impact on feature extraction and dimension reduction for short ut- terance text-independent speaker verification,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.059860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.687768Z digest=sha256:ea19fa862c4a0e4def44c3bde176ead0672abbb78f975cd956e63142b42ddd3b

Observation 50adb962-8278-4a59-a9de-fef72ff4b763 · outbound

This paper cites Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:19:59.851655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.692042Z digest=sha256:78d3797202b050c5cd1852ca3b0418ccc737dae2afea5d32532b7ae725c19262

Observation 4943a0d9-d4ad-432e-af54-cb1bbb2372df · outbound

This paper cites Whisper to normal speech conversion using sequence-to-sequence mapping model with auditory attention,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Whisper to normal speech conversion using sequence-to-sequence mapping model with auditory attention,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:00.038672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.696042Z digest=sha256:1ba5f4ecf758b7012fb9be5dbbc26c4d7b36fedd4e914436134de4c61f434ec5

Observation 55a6288d-e757-4447-9c5d-d46c9eec9c55 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.720018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.720018Z digest=sha256:00b409e3daec65840ca039b074fc826b66f3e83b1c8d04e3c951fa5ac3ec457b

Observation 5f77349d-2b6e-4c09-84c3-cd8e6901e049 · outbound

This paper cites End-to- end whispered speech recognition with frequency-weighted ap- proaches and pseudo whisper pre-training,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition End-to- end whispered speech recognition with frequency-weighted ap- proaches and pseudo whisper pre-training,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.990581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.702683Z digest=sha256:b1cc3037bed4a0695bb4c2815297dac6348093a916f455c58041dc4c7ea74e24

Observation a7d48d02-748e-431a-88cf-186b49dd3d7b · outbound

This paper cites Improving whispered speech recognition performance using pseudo-whispered based data augmentation,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Improving whispered speech recognition performance using pseudo-whispered based data augmentation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.979546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.705892Z digest=sha256:f79ec2f5bf20628286e1de24a2f69395a805a37eb36a13d10e8fb5ed8dd8e48a

Observation c8dddfea-b597-45e4-af39-88a511962d43 · outbound

This paper cites Whispered speech recogni- tion using deep denoising autoencoder and inverse filtering,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Whispered speech recogni- tion using deep denoising autoencoder and inverse filtering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.969448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.709042Z digest=sha256:06bbc60a7724ea2441317d05c2f20193d0e87d1342931aadb845200314d7d1cf

Observation 9f526990-9172-46e9-a10f-5ca1c4f7725b · outbound

This paper cites an unresolved cited work.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:20:00.129015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.669766Z digest=sha256:32e4f593ff5114f7544ff40cf05105d4ec5c37de26c7cf6740fa8d759fab9c2a

Observation b7e8fbd9-361e-489d-8881-e8b372672603 · outbound

This paper cites Multi-dialect speech recognition with a single sequence-to-sequence model,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Multi-dialect speech recognition with a single sequence-to-sequence model,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.958953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.712964Z digest=sha256:ec955414785f6bf197819d906fc0e8d80a333f206856b7d9023248163f1e1b8b

Observation 9b0bc379-8d2c-483d-abaf-4f8ad4c63302 · outbound

This paper cites Multi-dialect speech recognition in english using attention on ensemble of experts,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Multi-dialect speech recognition in english using attention on ensemble of experts,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.948819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.716538Z digest=sha256:8fb33ecbee2023e8c413c5b5c43256141bc16e177a24a5fca2feeb0e86cfe4d9

Observation 8167bd77-c85d-4ebe-a24d-47dc0900bcc8 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.723374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.723374Z digest=sha256:d7a07b5e3804078c0a8d4036b07552832dcbbd18a4f9f9e991f82254980a4986

Observation 716cfe03-0fb3-442c-9049-245d6aff0cd3 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.726713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.726713Z digest=sha256:da30d9a6019ea53b148e8a2f1685cb2f588aee895332e9ff103c5e9142baca6c

Observation 8be2b83c-d5b6-4c6a-acdd-c7196df6fdeb · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Robust speech recognition via large-scale weak supervision,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.730071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.730071Z digest=sha256:0f737f7c8d97f93eea4aaf2019d5c6962a169315fb04b2cee4a52e182ea8629c

Observation dc6196fd-7e5b-4638-82c3-9ba663f4d6a4 · outbound

This paper cites wtimit whispered timit dataset,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition wtimit whispered timit dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.915703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.733128Z digest=sha256:214b656fa187f6704f92110e97a448229628e83fff98b0823f91359f45f7f0fd

Observation f52d3c34-3268-4ca5-a590-b987f8bc954c · outbound

This paper cites Acoustic analysis of consonants in whispered speech,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Acoustic analysis of consonants in whispered speech,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.905064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.736135Z digest=sha256:cf78bc49d14da2987966a6ac29f1f1e886e4585d716d9af175d0fec921051e37

Observation b776020f-0336-45ca-a896-51d2636d081e · outbound

This paper cites La- ryngeal adjustment in whispering: magnetic resonance imaging study,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition La- ryngeal adjustment in whispering: magnetic resonance imaging study,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.894509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.739285Z digest=sha256:b00c1c19305cb48df5ab8f5b4caaacb280daf13d8cd9d8785c712333520deb00

Observation bd8b54c9-73bb-40e4-a82c-872769e8dc83 · outbound

This paper cites Acoustic differences between voiced and whispered speech in gender diverse speakers,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Acoustic differences between voiced and whispered speech in gender diverse speakers,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:19:59.883849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.742466Z digest=sha256:17690d416439bec7fcca03a667b05bd3477def394043d8195329ee1b363b22cd

Observation f736476c-e39d-4dea-83bf-601666351fed · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Lib- rispeech: an asr corpus based on public domain audio books,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.745826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.745826Z digest=sha256:f3670eed818624448a59d5cff8d51abedc5f739d2509703d58c6b8d41649f240

Observation 44223e4c-b682-40b9-8e45-7b50687989da · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.749075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.749075Z digest=sha256:d735d266ecaaecfe84ee1eb6923def9d04247ecfdf2168f1c1f6ac43db8f1e89

Observation 882bca90-7434-4f43-a582-00645041aa4c · outbound

This paper cites Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.752685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.752685Z digest=sha256:70f0b642aa13d747a6859ecd862360fbf3e7d7166c3f7b0cfccdeb8485d4149a

Observation 72e2873c-942a-42b9-bb99-33a848072dd0 · outbound

This paper cites SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.756520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.756520Z digest=sha256:abffe144128c2e3cdbdc25c14bc169a91c418a3d2dbc1e01282060dcde9f7054

Observation 05da14a7-142e-4e02-90c1-1a083138f265 · outbound

This paper cites Decoupled Weight Decay Regularization.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:19:59.760069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:19:59.760069Z digest=sha256:26ea342d28dedd01c085583b823d9f46cf24df7fda45795222f8e073a4704d64

Observation 415652aa-8cc6-4254-9525-79464aa22cee · outbound

This paper cites PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:19:59.799570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.764201Z digest=sha256:87d15201ea6a3eefef6c4e8a801e7ec98ca6b51377a4da3ddddc4cd92f3a0f2a

Pith citing papers

Observation 29b7e862-1586-40f1-b092-eaa59ece4617 · inbound

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition cites this paper.

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T19:19:59.866003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:19:59.656455Z digest=sha256:548be0a13e06221a00d05859b573e2de465654c3eca3b212e3dbd2576d3806a3