Pith. sign in

Paper Citation Record · LEDGER

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2505.20564.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20564 v3

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:29.985902Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:56:25.830154Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 11c1482e-737f-49a5-b229-e89af5940eed · outbound

This paper cites While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8].

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages While notable progress has been made in speech processing, African languages – including our focus languages, Igbo, Hausa, and Yoruba – have largely been left behind [ 7, 8]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:37.383518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:25.773223Z digest=sha256:e8015d459bf4107ae8ca880f356a5262ab786ae556c21864c0224c7e67995e44

Observation 3d6feb57-e0c6-4767-b7e1-11026b590d00 · outbound

This paper cites The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:25.830154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:25.830154Z digest=sha256:3d551e92d4d9e7515ce8df320831fed4438fa7ab3b151dfe79a5e1bdf6a46bd6

Observation 9af524c3-5de4-4e79-9692-67bc1d436ff8 · outbound

This paper cites It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages It features a wide range of speech patterns influenced by age, education levels, accents, and speaking styles – from broken to formal speech, ethnic and dialectal influences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:37.166709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:25.900230Z digest=sha256:c15d32ee6f5a137cec56ecf9ca3c8f7a5f8a39bcba87b6a8db108701bc763035

Observation 8ed4b76b-fcba-4576-a741-80f39a82d5fa · outbound

This paper cites Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20].

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Concretely, we finetune three selected ASR models on our dataset and evaluate them on both our test set (NV Test) and the FLEURS test set [20]

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:56:30.806189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:25.988475Z digest=sha256:e83ffbe22c6702fe06ca8973834b7f5eb8866ae4835006f6d704078774e40969

Observation 203c6a02-ac4b-4ebf-bff6-8e7b517a217f · outbound

This paper cites Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Built on the principles of ‘data farming’, our approach fosters a symbiotic relationship with language communities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.866653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.055703Z digest=sha256:0496a62f56dc09cc209400cd9c28443f7e1200ce55ff0d9fd9bc881f8f425ba3

Observation 5db42c02-a711-4017-bd61-39da9762e4ee · outbound

This paper cites an unresolved cited work.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:36.719269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.126261Z digest=sha256:0b717040bcfc03a9c6aa7665230ae8d39dabb4a2655e5b887612e2c5aca00bd1

Observation bb12c951-1bbc-4dd9-9b29-8f898acf4929 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Robust speech recognition via large-scale weak supervision,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.513465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.179296Z digest=sha256:7e883bc667aab3eb6523e5e17090008bec8296d3975a33175df50c80e3d1ade7

Observation f421c7e7-1285-4a25-b393-dc14c5d0898b · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Scaling speech technology to 1,000+ languages,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.276579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.261837Z digest=sha256:fbb952c9a82119ea08a2d08c3e1b6beb1cca28dc931d0d877c05d5783f3b8a47

Observation 83d1bb78-5b5f-407a-81cd-9c9a17bc186b · outbound

This paper cites The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.333107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.333107Z digest=sha256:56e66a3616b2d607118a3f6118a7d9d5939da075116d7c9fae923b8a032068ee

Observation 44c787d9-56e3-4010-9eee-db697332afbe · outbound

This paper cites IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages IndicVoices: Towards building an inclusive multilingual speech dataset for Indian languages,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:36.034575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.416637Z digest=sha256:3bf8e5a2b735b46dccbc769d4956ea5b275d4e264039d6ed7bef61076e4ab238

Observation e1b1d2cb-26c7-4fae-b3ce-93280571f454 · outbound

This paper cites IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages IndicV oices-R: Unlocking a massive multilingual multi- speaker speech corpus for scaling indian TTS,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.846060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.464410Z digest=sha256:0be7e21bbc225b485d3037a1838a41ac66cc93e676f9fc88d63f2dcf7f0e50a9

Observation 89b175f2-8933-40e3-a822-0401697b2d10 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.515597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.515597Z digest=sha256:bb6956a968c90b096cd20c2f35e141ad58c70bbcd063c1d79c43743e9dbf23a3

Observation b8e8a9e4-8227-4f6c-a7af-ae01145eebfa · outbound

This paper cites Replication data for Igbo Natural Language Processing Tasks I,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Replication data for Igbo Natural Language Processing Tasks I,

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T13:56:30.348444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.604471Z digest=sha256:63e7f02737a0c21989e1905f1774ae8c3e113608f70169c5244340a2d0a166cd

Observation bdfb1d8e-de7e-46a0-b35f-567ccf5dbe1a · outbound

This paper cites Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Multi- lingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.621338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.691212Z digest=sha256:caefc5bd2e888696091e9f957c1f9b31290363dff0b57f3d2dc0f1216ffcc7de

Observation 9830b3b2-4efb-4e9c-95c7-851e53957cb6 · outbound

This paper cites Masakhane -- Machine Translation For Africa.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Masakhane -- Machine Translation For Africa

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:26.796632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:26.796632Z digest=sha256:b5dfc73617b39021fb989260f5cfb90c9e933cfc4659976e268a8ea2004a9994

Observation a4114642-5a1d-424e-89c3-e0f5d75afd6b · outbound

This paper cites Partici- patory research for low-resourced machine translation: A case study in African languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Partici- patory research for low-resourced machine translation: A case study in African languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.407955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:26.930277Z digest=sha256:a2b3a6a90b470be1c36066c8068e37cae4a927eeff0299107f07e01bf26d9b01

Observation 4ad9046d-5454-4b94-961f-eb494a51f6a6 · outbound

This paper cites The state and fate of linguistic diversity and inclusion in the NLP world,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The state and fate of linguistic diversity and inclusion in the NLP world,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.176512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.058131Z digest=sha256:27da479a1bebcf2f433e549a8626a64e8d27a83a65e973531efc2ae5d2883fd3

Observation e8546569-57eb-4caa-8ebf-9c30d18b0c80 · outbound

This paper cites A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages A few thousand trans- lations go a long way! Leveraging pre-trained models for African news translation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:35.026735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.228195Z digest=sha256:2d06049a9f542f72a75e685f5d39005a36ff664419f0705674ca272ac23fd2f2

Observation cffa84ed-3bc5-4ca7-bc9d-1a9b04888149 · outbound

This paper cites an unresolved cited work.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:56:34.803033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.343643Z digest=sha256:e615eab677702cabe9769cfcc667f7587ece3ec938d905696b9c85531ff44667

Observation f9a5659e-f323-45da-ad3a-76927b2cf504 · outbound

This paper cites GlobalPhone: A multi- lingual text & speech database in 20 languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages GlobalPhone: A multi- lingual text & speech database in 20 languages,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.531666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.466576Z digest=sha256:57c505321400a1a241f62133080d59439249d592c596177d62e02d8b4ad58dcc

Observation 99dd0dba-51ca-4205-97d5-99d22f6feba7 · outbound

This paper cites YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages YFACC: A Yor`ub´a speech–image dataset for cross-lingual keyword localisation through visual grounding,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.335515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.610794Z digest=sha256:47a10ecae5ae6010e3763e0ae6b5e51deb43212d0eb75ed89b8d6f1d6e654264

Observation 9c46c67d-597e-4272-8e55-8bb5173a78a1 · outbound

This paper cites \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b\'{a} Speech Corpus.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages \`{I}r\`{o}y\`{i}nSpeech: A multi-purpose Yor\`{u}b\'{a} Speech Corpus

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:56:30.593979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.728384Z digest=sha256:a9a7b1154d399b46ac189f58d866ee19117967c27e6c2aea95072ef7261731ac

Observation 305af219-7afc-4346-a0dd-8610269909f9 · outbound

This paper cites V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages V oices Unheard: NLP resources and models for Yor`ub´a regional dialects,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:34.128670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.834694Z digest=sha256:9bbb1e6b8506bb86a2965ec46183dc767603411451d9f104d7321cfe504ee195

Observation 0a70f040-138c-4672-aea9-90a2801964fa · outbound

This paper cites BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.929958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:27.942667Z digest=sha256:d64a9969f0433e077e071c16cf95a15591538e78e8fb1289197ad5503f1b3ce8

Observation 3147b531-225a-4bb9-ace9-2a0d06c27f34 · outbound

This paper cites Common V oice: A massively-multilingual speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Common V oice: A massively-multilingual speech corpus,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.713462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.035532Z digest=sha256:b35a8fc817029da562b2dac8dbf203928f6bb6431bcbfcb8e9aa8f54ca9bbd3b

Observation a7193215-c851-4cc5-ba8c-f30b9d110db3 · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages FLEURS: Few-shot learning evaluation of universal representations of speech,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.504591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.199221Z digest=sha256:e3b6828fbada4700ad58872d87978c5cbf63790bfe0956338c62e3713e6315c0

Observation cc6f6625-dab4-47ea-9537-5f679ec36301 · outbound

This paper cites Quality at a glance: An audit of web-crawled multilingual datasets,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Quality at a glance: An audit of web-crawled multilingual datasets,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.338427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.332611Z digest=sha256:d9377e5bd9bb9ac1287d742190ac7135bb9ae935c7fc50ad12a501973791c4fa

Observation 96873929-315a-446b-b9c0-09bbd1e25e6c · outbound

This paper cites Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Separating grains from the chaff: Using data filtering to improve multilingual translation for low-resourced African languages,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:33.149536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.490123Z digest=sha256:08f53c687b1ddda4ec91a701c7642d57e76cc245994e80f18b5edbbb51498735

Observation 8bc08eac-f04b-4a3f-8b00-637b31757013 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:28.580038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:28.580038Z digest=sha256:17251c97dca331eb8402a55d15829a1b385c56f0064368c563b957a024d9303a

Observation acc0dc36-868a-438c-abd2-87d3171859db · outbound

This paper cites JW300: A wide-coverage parallel corpus for low-resource languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages JW300: A wide-coverage parallel corpus for low-resource languages,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.979742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.668877Z digest=sha256:eb239815c6bec3627c0605762b4a9a2e5fd3c726db46e9df2fc6f0477349dedf

Observation b7d45a26-68e8-4323-a7b1-ff71d126d160 · outbound

This paper cites `Ir`oy`ınspeech: Yor`ub´a speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages `Ir`oy`ınspeech: Yor`ub´a speech corpus,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.809190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.747396Z digest=sha256:3715206d4409d29f76fd17373a15141f95a776cf0787ce2533c8cb3a1254166d

Observation 26073f54-6068-493c-ac4a-fe7e5f0122ef · outbound

This paper cites Hausa speech corpus,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Hausa speech corpus,

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T13:56:30.221593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.816780Z digest=sha256:1455ace1fd8fb7410d70a6605ad9cda2bf15d211ccc3ced944af36ddfef6d6f6

Observation c7af0928-5d8e-4744-88fe-e08c336cb797 · outbound

This paper cites Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Kencorpus: A Kenyan language corpus of Swahili, Dholuo and Luhya for natural language processing tasks,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.526703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:28.892248Z digest=sha256:a3d033a86f9b99510cf30aef2df4590a8bfd04881b3e1f4a756b4b7055d446fb

Observation 2598dc06-a744-406e-8ddd-f238606782fa · outbound

This paper cites LIII. On lines and planes of closest fit to systems of points in space,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages LIII. On lines and planes of closest fit to systems of points in space,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:32.254430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.019144Z digest=sha256:2f4e5c1ee1380213c17c6c783618c40b9ed03f35b4168c116ee36e6bdaafb8ec

Observation cd5e4632-baf8-4fb5-b79a-656297daf12e · outbound

This paper cites Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Robust signal-to-noise ratio estima- tion based on waveform amplitude distribution analysis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.952490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.122151Z digest=sha256:023046064b6c3371db12b9b929d80cd9bbf5e90a4fafcf5ce4a903cbd61e025d

Observation 5cca4861-c392-4c9a-b371-1c33e94ae017 · outbound

This paper cites Audio quality feature,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Audio quality feature,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.618588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.220939Z digest=sha256:2cdc16bb439c1957727dea19bb6e86f71ba1b76555728a63ba77550ddc4c72a0

Observation a364f504-22f4-479d-bc69-db6c22758a11 · outbound

This paper cites (n.d.) Evaluation.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages (n.d.) Evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.388331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.281822Z digest=sha256:0388be7c208f1f3aabc0bdf43f2bb468c452b422e73130fda08872e896836ca9

Observation 8f7bbb81-62db-4296-bed6-c80cd69f1d15 · outbound

This paper cites Unsupervised cross-lingual representation learning for speech recognition,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Unsupervised cross-lingual representation learning for speech recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.170421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.372393Z digest=sha256:9ca71e16830a854f0c94838384f1d6ccacfb8d122e6fbdc5386d34fc4a045cdc

Observation 6dbcbbbd-4370-4a97-9aba-c82c3a6a2db4 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.523320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.523320Z digest=sha256:91aad9cf686d8c482ffd70dfa36e890d264ca7bf949e0c4a066c1f081160c892

Observation b289d64b-ca04-497b-936c-3f7973f29acd · outbound

This paper cites Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Small Data? No Problem! Ex- ploring the viability of pretrained multilingual language mod- els for low-resourced languages,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:31.052468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.620364Z digest=sha256:1849a952017d68db7073fa09e839d1008e5f6d6e8de8007382f54909a99778fa

Observation afc908dd-cfef-403e-8a61-5063e67c5c35 · outbound

This paper cites Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.695164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.695164Z digest=sha256:c0c71373a9222bc969beaf3e7dbdce9ebf4d8494ecea8cfad21a0a60a1b8e082

Observation 1adb9986-0c8f-4d73-b305-abffaafe2324 · outbound

This paper cites What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages What makes a high-quality training dataset for large language mod- els: A practitioners’ perspective,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:56:30.929427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.985902Z digest=sha256:bcbfd8e15481a9e09e87c974bd40c5880d00e6e44794b1455bb67ff4c0f5de8b

Observation 88dd9e44-4744-41c2-b976-6488a62e9e9c · outbound

This paper cites A Proposal to Study "Is High Quality Data All We Need?".

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages A Proposal to Study "Is High Quality Data All We Need?"

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:56:30.468477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:56:29.888274Z digest=sha256:5a403e689ab8cab82d17f0df7429e0c3891d32b8b41cc00f8c34835530f8efd5

Pith citing papers

Observation 3d6feb57-e0c6-4767-b7e1-11026b590d00 · inbound

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages cites this paper.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:25.830154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:25.830154Z digest=sha256:3d551e92d4d9e7515ce8df320831fed4438fa7ab3b151dfe79a5e1bdf6a46bd6

Observation 6771a249-08d4-4e15-a0d4-0f90b1c86d08 · inbound

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI cites this paper.

Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:50:28.342515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T19:27:18.774649Z digest=sha256:6bbd0ef088df39a1427dc82c2dcad1228b88669657476c372c6987665042d355

Observation 983a31b4-e4ab-4f69-af0b-3058fad203f5 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

Reference 207

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.629765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:6b0889e4acfd6f0734f12364c387e557fe6c960a54d41ca591043b3aaa2c4e6f