Pith. sign in

Paper Citation Record · LEDGER

Scaling Speech Technology to 1,000+ Languages

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2305.13516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13516 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:56:13.415115Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

116
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3d61ba3-faed-4c9c-a3ef-41e280caaed2 · inbound

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset cites this paper.

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset Scaling Speech Technology to 1,000+ Languages

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:36:00.889425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T04:34:59.205522Z digest=sha256:99e76d6a389e8e82aa681c25436a76414214c0152839dd240907ef6c456fb3af

Observation c3701935-2526-4bf6-9040-f69e691baa70 · inbound

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models cites this paper.

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models Scaling Speech Technology to 1,000+ Languages

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T12:56:13.415115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:56:13.415115Z digest=sha256:12ed4b15ee8c0219bc45cc52f46f8a3efb58404b7c522f2dc73bb03b150e4b57

Observation fa927900-55be-4a91-a004-7b36b856533a · inbound

Contrastive Learning for Task-Independent SpeechLLM-Pretraining cites this paper.

Contrastive Learning for Task-Independent SpeechLLM-Pretraining Scaling Speech Technology to 1,000+ Languages

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:13:48.535917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:13:48.535917Z digest=sha256:9121c9d676ff094ca33d3cd78812f5cb1ec05adc20aa2921f476cfb30d8de8bb

Observation 553924f0-a986-4f16-b206-ba942c3f43d9 · inbound

FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion cites this paper.

FreeSVC: Towards Zero-shot Multilingual Singing Voice Conversion Scaling Speech Technology to 1,000+ Languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:15:41.008870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:15:41.008870Z digest=sha256:6bb7c49fd9734f6a61a361b192de6935530823300c1420dc3e505b59de6d1d99

Observation f23c4ab2-fcc3-43ed-a3d3-77c25acdf27c · inbound

Continual Learning with Embedding Layer Surgery and Task-wise Beam Search using Whisper cites this paper.

Continual Learning with Embedding Layer Surgery and Task-wise Beam Search using Whisper Scaling Speech Technology to 1,000+ Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:20.399211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:20.399211Z digest=sha256:d40495bead70a5643ddd7a6e5c3c82e235db91da9a743e06eac3116f95558db3

Observation e553957c-9110-494f-a3e7-4b89ac1001d4 · inbound

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models cites this paper.

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models Scaling Speech Technology to 1,000+ Languages

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:14.831683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:14.831683Z digest=sha256:b5479f3a95127c49c348d1e99b81b0590abe7a566a053decc937ced4e532c924

Observation 48978583-1258-41b0-8653-d9ddca2174a2 · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.498003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.498003Z digest=sha256:dcb0cecf00e816cfea1bad66aa63584ae55bc160745a276be3a62dd6e6bbd2b9

Observation 8cfc4100-3c36-4267-9218-487531fe4f54 · inbound

Context-Driven Dynamic Pruning for Large Speech Foundation Models cites this paper.

Context-Driven Dynamic Pruning for Large Speech Foundation Models Scaling Speech Technology to 1,000+ Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:09.278802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:09.278802Z digest=sha256:f709e13cf96534d0f9b1cc8e667d03d056c3bedbf92f1e301a559cb3b4232cd1

Observation 551256bb-0573-477d-a3c7-b06cb9d58e4c · inbound

Improving Language and Modality Transfer in Translation by Character-level Modeling cites this paper.

Improving Language and Modality Transfer in Translation by Character-level Modeling Scaling Speech Technology to 1,000+ Languages

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.487913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.487913Z digest=sha256:78f073388e187c823681e5eaba5ce2f7ab19061089dc55d9aa4689338e5b21b8

Observation 7f4d9657-85cc-41d7-9429-a4ab92bdb808 · inbound

Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages cites this paper.

Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages Scaling Speech Technology to 1,000+ Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:21.990733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:21.990733Z digest=sha256:7dda255fbc7cbd801d661f37ba90c71f5a348b4ef491d82338b7029b427c115f

Observation 307f9388-67b8-41fa-9068-7d207fb2a396 · inbound

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion cites this paper.

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion Scaling Speech Technology to 1,000+ Languages

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:39.174752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:48:39.174752Z digest=sha256:1c536364ed03945674259137d6ba57f86d1822fda8976b7d833f1133ccb53f58

Observation 5fab6ba7-56d8-410e-b917-55ded70f0929 · inbound

On Barriers to Archival Audio Processing cites this paper.

On Barriers to Archival Audio Processing Scaling Speech Technology to 1,000+ Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:08.841029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:14:08.841029Z digest=sha256:aacde09201ce0fb14c07a7c3eba663b6e11e972a54836d236011f7d08d3f7b6a

Observation ab8d4adc-37bd-45eb-920c-f7370d130374 · inbound

A Hybrid Machine Learning Framework for Optimizing Crop Selection via Agronomic and Economic Forecasting cites this paper.

A Hybrid Machine Learning Framework for Optimizing Crop Selection via Agronomic and Economic Forecasting Scaling Speech Technology to 1,000+ Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:56:20.540912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:56:20.540912Z digest=sha256:0e6d2ca2a8bcacd6ef3d2078c46809b85d27ee0d08272a639a1d81faf9af29b1

Observation c7afc37d-fa66-4172-ae08-6e15342c902e · inbound

Coherence in the brain unfolds across separable temporal regimes cites this paper.

Coherence in the brain unfolds across separable temporal regimes Scaling Speech Technology to 1,000+ Languages

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:13:22.862525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T20:11:29.476625Z digest=sha256:5d9e24ade6b0897170e05389933cae0aeca2b30cc70efcfd5ca22b884bbe00d2

Observation ddd22543-ba0a-49e6-a203-fe708de2b2d7 · inbound

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative cites this paper.

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative Scaling Speech Technology to 1,000+ Languages

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:46:12.487399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T02:27:05.776290Z digest=sha256:5ee26edc8759894be3a42f8ab7b6c129948741c0362eef29963ea134db6bc445

Observation 60d89d46-3473-489f-b9c7-f81ed640d2c2 · inbound

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation cites this paper.

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation Scaling Speech Technology to 1,000+ Languages

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.754248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:44:30.762851Z digest=sha256:4568a66b4bd5aa63d3eac73c24bc8ae138adf1750a97d6cda922a45a6a05225c

Observation 5a19eb96-1e3a-4e65-864c-e643fa3a51a4 · inbound

Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations cites this paper.

Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations Scaling Speech Technology to 1,000+ Languages

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.892265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:16:39.890339Z digest=sha256:1d647bc7e23662c89b6205ffedbdad3f6370c0f042669ea57442a40bace027e4

Observation 5e08dc38-63c2-469d-8de7-cc5d58a180b1 · inbound

BlasBench: An Open Benchmark for Irish Speech Recognition cites this paper.

BlasBench: An Open Benchmark for Irish Speech Recognition Scaling Speech Technology to 1,000+ Languages

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.558355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T15:53:54.092426Z digest=sha256:c76ecc703a9cbe8618438da0aa58a543f4b1a0695e7fff691ab7cc601462ec68

Observation d41f838d-44b4-43e7-a11a-56f15e06e0f4 · inbound

Tadabur: A Large-Scale Quran Audio Dataset cites this paper.

Tadabur: A Large-Scale Quran Audio Dataset Scaling Speech Technology to 1,000+ Languages

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:11:03.683889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T02:16:09.216191Z digest=sha256:fcaa9fbee680d6b09e69d860166c51e039ff902965ff8bce57fab75105552649

Observation 1d0a1619-11f1-432b-8a93-44166d0ee087 · inbound

A framework for analyzing concept representations in neural models cites this paper.

A framework for analyzing concept representations in neural models Scaling Speech Technology to 1,000+ Languages

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:18:59.078659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T14:49:22.776209Z digest=sha256:87f8f03519f17828d37dace567499f1d09e9d463d2e0a1cb2b38e8b889af4660

Observation f7355c90-ca9c-4bb5-971f-fb6fc32b7c38 · inbound

A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification cites this paper.

A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification Scaling Speech Technology to 1,000+ Languages

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.361386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T15:12:45.261874Z digest=sha256:1a5c6e1919641a0257209be61c2a6f281557a270d93fede677da523add736f45

Observation 41a4aab9-8e7e-4b78-90f1-602b7d0e7f87 · inbound

Pretrained self-supervised speech models can recognize unseen consonants cites this paper.

Pretrained self-supervised speech models can recognize unseen consonants Scaling Speech Technology to 1,000+ Languages

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.073107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T10:14:47.932613Z digest=sha256:6ee18a32489692ece2dd18cd06c305f12677dd7a1fcd1441e84c37918b830fb1

Observation c7c44471-50e5-4550-bed9-a73c4eed6376 · inbound

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean cites this paper.

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean Scaling Speech Technology to 1,000+ Languages

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:19:50.277522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T05:24:03.268025Z digest=sha256:8ed2d9ab4037f079448f831e0c8e27226418e3fdd6d68cb9d8cdd34393b4a893

Observation 10db3fde-f1e0-4ce7-a1ea-aa0756afa508 · inbound

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling cites this paper.

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling Scaling Speech Technology to 1,000+ Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T00:48:21.716770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:48:21.716770Z digest=sha256:5c073b99bb5f350cbd496d921154552b80bc68deb68b99560c2e00403b3d036d

Observation ba5de31c-da83-4527-8f35-47a42e0f611b · inbound

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language cites this paper.

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language Scaling Speech Technology to 1,000+ Languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T18:11:36.626189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:11:36.626189Z digest=sha256:8ec889f6805d40eff01a9815ed8e81f1e1d8c8deb59dfdeaa4f42e04c82ce7e9

Observation 69f96c9a-58e0-4b13-9b50-9cae96658e6c · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages Scaling Speech Technology to 1,000+ Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.841258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.841258Z digest=sha256:da7f90ed01aa3be4c6b4e9c5de59b0f35a26bfd337662c90762a404c3d252cf2