Pith. sign in

Paper Citation Record · LEDGER

The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2111.09344.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09344 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:38:53.970993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.966776Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 424a669b-2016-4d62-a083-84d59e362c72 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.274992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.274992Z digest=sha256:6105b17224be110b54c2087fb620e09cfcf5c57dd6d71edf1a8fb168761f0556

Observation a0686dc0-edd1-46c0-99f4-d5bd525bcb52 · inbound

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning cites this paper.

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:38:53.970993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:38:53.970993Z digest=sha256:d9d5849ab7f8ed8cfb6dabb8e2a4f1918dfa43da5d121e3307a19897bc108ce2

Observation 52b343b8-f89f-43d4-9dfa-e35a1fa2c00c · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.463964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.463964Z digest=sha256:9024179700cd63c9120dd67c10ee089a36adfa20c727a9fee382a53053323ec3

Observation d94107cd-9278-4f9c-abfd-72870ccb72d7 · inbound

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use cites this paper.

Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:20.652179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:20.652179Z digest=sha256:48ae38e9429a9f696918499557fc179a1fae4ab0f5c42d0dff52c7716ee6be29

Observation 9ec2465d-b686-4696-a457-74c0a22ec0aa · inbound

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning cites this paper.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:19.512688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:19.512688Z digest=sha256:9e60f8e7a9af247c11f84557d25292675df2238f1f558c32bbbc375bdee484f7

Observation 7b443f9f-0199-4440-9394-983dbf2e8ece · inbound

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data cites this paper.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:13.381090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:13.381090Z digest=sha256:ec4aab9c315ef45236df232bb570039f3435ef70a9fd1bd10b9958242f47323d

Observation 007423dc-82e9-40c5-be88-0607471da412 · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:07.524656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:07.524656Z digest=sha256:a3aa5f5c733832d332eb77f3c1fdc745e94cf037569b5a944ad4d0add2126527

Observation 82dc9500-d820-4672-8a3b-a1b4f3ff12c7 · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.432916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.432916Z digest=sha256:f55836ef9e346f2d3fbbd87f49bd4584d56839ef98c5a0da28ef525a999c5e5a

Observation 1f9a1363-5163-4ac3-aa36-633587a2a7a7 · inbound

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit cites this paper.

Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:12.458369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:12.458369Z digest=sha256:dffcd20eeabfe3a8fd03b258b3a8da61c65f1dfd52c9fd5982e0fb4a034a2aff

Observation 1116d9ff-1bb6-45fe-8c3b-2d690506a98d · inbound

TTS-1 Technical Report cites this paper.

TTS-1 Technical Report The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.830714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.830714Z digest=sha256:6d300b86db932981b9ab2cd1ce378e95314d34d4d3c16c5262dd30c2911fb067

Observation 3401423f-1608-4c4e-bf02-fad5fe8224fa · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.415152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.415152Z digest=sha256:adc5a4110ab10e9bff0b56a1473840b8ac1317c9aae81a4e30860e787b048036

Observation 11f03851-e677-4215-b44e-f17e27cdb100 · inbound

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training cites this paper.

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:53:38.926567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:53:38.926567Z digest=sha256:99f4757c942c18d0b3e8e755a4d74293c28dc572545d0047acc6bdf931defd6a

Observation 6f3f1d5d-1fc1-4078-aac8-50ad37a2817d · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.397871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:1f39edf0c0ac5b52eac984356cfa307cf0df2323412fc39661551c94c8906f2f

Observation f08ae745-8c0c-45c9-8e5d-5c4ce0d8d8ab · inbound

Swivuriso: The South African Next Voices Multilingual Speech Dataset cites this paper.

Swivuriso: The South African Next Voices Multilingual Speech Dataset The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T19:05:08.543038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:05:08.543038Z digest=sha256:4d191395d7fe1a282847f47b126a492afc279ecbe2ee34c1a9698e135be67b1f

Observation d7b854e1-cb19-4514-8f8b-5dff47f5ff84 · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:30:57.407031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:13b51da00dbebd109e0f50761818b0039690e481c6d11780f89786ddc0dab1ba

Observation bb045c2c-b638-4482-a623-086d0037e76f · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.964075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:facc3faa77ee9932a673498506cda33d53948846c064cd48c940252eb7a399ff

Observation 068edde5-1384-476f-b6df-1f05db66a483 · inbound

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations cites this paper.

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:27.811824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:34:01.865434Z digest=sha256:c7781cbeedc19f28d4db77ea5b32671ee6b4bea604830d2d582b86023415ecc2

Observation be8a2da7-1b27-480e-87ce-879cc3e59f4b · inbound

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations cites this paper.

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:59:11.692624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T22:57:59.897674Z digest=sha256:39cac84f2868ec8e6602f7e3dcb6b6f6141c943e8c35d292c5ebab42dbd54d30

Observation 49cb625b-4304-4618-9094-15bb842ea5f8 · inbound

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations cites this paper.

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:45:45.419554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T22:53:18.321056Z digest=sha256:5864a55f7311c4da3a245fca62a9483a75ac670dea9601d4c4ff11c31bc330c1

Observation 167ab0d4-6cc9-48fb-9d09-f00e78cea41b · inbound

A Semi-Supervised Framework for Speech Confidence Detection using Whisper cites this paper.

A Semi-Supervised Framework for Speech Confidence Detection using Whisper The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:02:13.237407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T04:00:49.889795Z digest=sha256:83d3b3eba0231ec6d94c6bcc97902c5409d16dace522bd5a5cd74b03f160431f

Observation 0d7c2b33-d648-40b7-8bcc-e3d509e1b74f · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.507165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:7e2dc6c8aa66169bb9a97f27759307aefe97e50a022d4bb5069e8f7c8a88727d

Observation a8d9ac97-39d3-446e-bb53-a2f48a6c3582 · inbound

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails cites this paper.

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:47:19.541629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T21:19:56.932689Z digest=sha256:f111872f7ea6b3ec3e963a59d5e36344333f62f5772c57448aa0e02245172b3e

Observation df4de4f9-255a-4315-90ca-f4423eb35d0b · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:42.903183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:88f9b0aa862c90ba101468622b037529cd434a57389be18d7ff6bf361fbafdaf

Observation 6a8b02e2-f9f1-4c90-8e35-fb7561cdf825 · inbound

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models cites this paper.

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:10:07.969024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T20:36:24.901454Z digest=sha256:c3704c56aad851123c68255eccaaf057c066d49ff9a815a9c7e3b4406041c82f

Observation b47761c7-f254-4d3a-9afa-934c9feee09d · inbound

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR cites this paper.

Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T09:35:44.381743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:35:44.381743Z digest=sha256:c4074698cfeb66aa7d132f6164687d036e2653ced220cbe9d40530f5aa81438f