Pith. sign in

Paper Citation Record · LEDGER

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2505.20506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20506 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:28.634322Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:25.595644Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:00:28.778797Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9f8115b-88fa-4f4b-9ba0-7b5d2b52230e · outbound

This paper cites an unresolved cited work.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:33.316583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:25.491796Z digest=sha256:e37ee964927c805cc852e2b7f79a93aa7b99d6e5357da277e8c0ee0e4cd60e26

Observation b3b01c96-e5a6-431e-a440-48ac43f529a0 · outbound

This paper cites ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:00:28.939908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:25.595644Z digest=sha256:f64f699e3d0fec76ddb94d8ba8fded75877d43f94b5acf7008928e069136bdbf

Observation dfb1b0c9-46c1-48b4-8a93-5b50995bb039 · outbound

This paper cites In this section, we describe each part of ArV oice and provide justifica- tion for design decisions where applicable.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis In this section, we describe each part of ArV oice and provide justifica- tion for design decisions where applicable

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:33.048002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:25.719894Z digest=sha256:4f61bf22c8772f50856a18da469e28c628c81871597b205ed7a3641ae90736be

Observation 53e56dec-cefd-45d0-a40e-a8f05d00d8b3 · outbound

This paper cites an unresolved cited work.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:32.857585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:25.864240Z digest=sha256:20026130b3a2d77edb401eb447a63aa13d256832d1c5b00d586c414833bce4a0

Observation 504024d5-b896-4714-bf34-0f8b1f7a0a16 · outbound

This paper cites The dataset consists of 11 voices in total, 7 of which are human voices, and 4 are syn- thetic with parallel text.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis The dataset consists of 11 voices in total, 7 of which are human voices, and 4 are syn- thetic with parallel text

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:32.522013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.057794Z digest=sha256:4998cdef491d736a7fa3f0fe05b11ba1654a003a430fb5f251ebd5f688862e64

Observation 027b4f30-a22f-4281-9271-061b715c7bcc · outbound

This paper cites This work was partially funded by a Google research award (11/2023).

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis This work was partially funded by a Google research award (11/2023)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:32.321259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.148273Z digest=sha256:c4bd238efb1b826d8b93bc16831b8ee296960afe332e10e8153c49dc0ddc3bab

Observation ab7b15dd-9064-48a4-b593-a83a87fad8c9 · outbound

This paper cites Cmu wilderness multilingual speech dataset,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Cmu wilderness multilingual speech dataset,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.404796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.742493Z digest=sha256:1d0e927bc1c983697713e437ee476c5b46c68a27e6e5762415a3e474388b611a

Observation 86a121d7-422a-40da-86a4-9c12d8eed7b6 · outbound

This paper cites Speech recognition challenge in the wild: Arabic mgb-3,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Speech recognition challenge in the wild: Arabic mgb-3,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:26.224154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:26.224154Z digest=sha256:640d95327579f8e02eb344b78b158523577d4715ba64a559757d3406a30982a7

Observation 5ee82de3-d52f-495f-a59d-c9f82452c7e5 · outbound

This paper cites QASR: QCRI aljazeera speech resource a large scale annotated Arabic speech corpus,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis QASR: QCRI aljazeera speech resource a large scale annotated Arabic speech corpus,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:32.148453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.295851Z digest=sha256:30d5bc37bc8b3761933bd3f2074acfcbf859cff3ec1344088cbf4c171b38bbc4

Observation e69c566f-2255-4adc-82e5-6b2a1379ae4b · outbound

This paper cites Masc: Massive arabic speech corpus,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Masc: Massive arabic speech corpus,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.972932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.415432Z digest=sha256:abe4395d780c88f41ce8dec59e53630e6f4d98c103871c4b3926a381970540f5

Observation 053ef7d6-8f61-4a79-bc8f-df7c3f39f83b · outbound

This paper cites Diacritic recognition perfor- mance in arabic asr,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Diacritic recognition perfor- mance in arabic asr,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.748774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.475402Z digest=sha256:cf15331faf4f81b6ac7784c309c60de572f3abbda7dd260f6db63d84b721dbf8

Observation 9112c2e5-c618-47fe-ad9a-eb3823aa28cf · outbound

This paper cites Automatic restora- tion of diacritics for speech data sets,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Automatic restora- tion of diacritics for speech data sets,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.573106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.595759Z digest=sha256:7c1a0e3096b2a15ae953ce6fa8d69548a39db86eecd411c98a61f4a9597fee59

Observation dd85b093-b194-469c-afc5-51f7b77b07cf · outbound

This paper cites Clartts: An open-source classical arabic text-to-speech corpus,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Clartts: An open-source classical arabic text-to-speech corpus,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.492215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.674924Z digest=sha256:7a16a98403e51fc2ae2b59e2809a5f40c8c9037eb54f67f585abb6021799fc96

Observation 53a911d2-71d0-4978-a8aa-4c5e4be5be7b · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Robust speech recognition via large-scale weak supervision,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:30.089431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.682750Z digest=sha256:c29ff317c66b863cf40df10631e4267d33b8f85f36aa65c6f8767fe3fa10f44a

Observation 9ef20df5-6677-456b-8174-27a1f6c87197 · outbound

This paper cites ArzEn: A speech corpus for code-switched Egyptian Arabic-English,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis ArzEn: A speech corpus for code-switched Egyptian Arabic-English,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.239238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.831943Z digest=sha256:210a01b5e647acb3e3f8b62d6e20ebdf13a10dff236788324d128ef2db7facb8

Observation 1e8c6ac5-a457-4e26-8358-c331085d93d5 · outbound

This paper cites V oxblink: A large scale speaker verification dataset on camera,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis V oxblink: A large scale speaker verification dataset on camera,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:31.089268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:26.916277Z digest=sha256:4734e441fcac106225f61a19a297a67b7f7fc72298456e4d8a92a30f26426e0b

Observation 65af7c26-2278-48fa-9ce9-047a8ad23143 · outbound

This paper cites Modern standard arabic phonetics for speech synthesis,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Modern standard arabic phonetics for speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:30.900960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.019303Z digest=sha256:6c35c15dbfbd2e0b757ba7dfb7d3f18feb81956925ed42a49dee4cb659e0ece5

Observation f620d47a-d187-42f4-93e0-6e13081bcf3c · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:28.199944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:28.199944Z digest=sha256:ee1f7ebcc74e9fdd8ec6b9cd369db4cd3b27c632e7660ac4034d6b59cf0c7705

Observation ffc5d9f1-5ed1-4090-880c-fae2b0018c2f · outbound

This paper cites Tashkeela: Novel corpus of arabic vo- calized texts, data for auto-diacritization systems,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Tashkeela: Novel corpus of arabic vo- calized texts, data for auto-diacritization systems,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:30.602628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.241330Z digest=sha256:25b1b58145f9dade25115a616e8aea21a135014d2b97fdf861dc567e21cb5035

Observation 4f860300-c013-4d5e-93cc-f81f59662c3d · outbound

This paper cites We also fine-tuned KNN-VC [21], a non-parallel VC model that converts source into target speech by replacing each frame (a) VITS (w.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis We also fine-tuned KNN-VC [21], a non-parallel VC model that converts source into target speech by replacing each frame (a) VITS (w

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:32.690016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:25.989216Z digest=sha256:0c5297738fda908b8d0012e27579efe736d1955ea1ebc73bc3ba42efb2cbd9ce

Observation ba7c7e28-106d-4fb4-a4b0-37b977092f62 · outbound

This paper cites Comparison of topic identification methods for arabic language,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Comparison of topic identification methods for arabic language,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:30.443055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.415410Z digest=sha256:1ec5534acfb19909ea71c459b3cdff4747f8254513024fe36d733567cea82fbf

Observation 595761de-0765-4661-924a-299b8c43c8dc · outbound

This paper cites STTATTS: Unified speech-to-text and text-to-speech model,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis STTATTS: Unified speech-to-text and text-to-speech model,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:30.288429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.583427Z digest=sha256:18288b76984a52a64722a6a81397a1d0b7d3f03571332be3e5eee7326791e966

Observation 25212192-0b16-4a2a-86c8-024f93293fcf · outbound

This paper cites ArTST: Arabic text and speech transformer,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis ArTST: Arabic text and speech transformer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:29.737572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.836847Z digest=sha256:730ddb9f502e4c21da5c27029ef80edc3a0e6b3bec2fc55ff1cf53493db9b222

Observation 15679c2d-fd08-41d1-87b3-e00a189b3c50 · outbound

This paper cites Fast conformer with linearly scalable attention for efficient speech recognition,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Fast conformer with linearly scalable attention for efficient speech recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:29.484852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.957211Z digest=sha256:15b1a68e7e401a22a312b254b82808e35ed2d818a39e419fc491c06112b61c77

Observation 68242d48-3888-4ae0-a34e-f2e53dc222d1 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:28.076071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:28.076071Z digest=sha256:33cd7582d4320958ccf071b1374bb15731f3ecdf610a0be48ff91bc1665f4cff

Observation ca834ef0-2608-411c-8690-a878ddb3dd7e · outbound

This paper cites AAS-VC: On the Generalization Ability of Automatic Alignment Search based Non-autoregressive Sequence-to-sequence Voice Conversion.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis AAS-VC: On the Generalization Ability of Automatic Alignment Search based Non-autoregressive Sequence-to-sequence Voice Conversion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:28.372172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:28.372172Z digest=sha256:ee7e14ce45c1bd1c1e7262460b651f945eaab49fdd2f30b0f936cc31366e6481

Observation 6b959bc7-994d-4b22-8339-aeb77a222c66 · outbound

This paper cites Parallel wavegan: A fast waveform generation model based on generative adversarial net- works with multi-resolution spectrogram,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Parallel wavegan: A fast waveform generation model based on generative adversarial net- works with multi-resolution spectrogram,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:29.248170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:28.521972Z digest=sha256:91f82cfb0ad4889d21cea356647ff32434871e274ee768961a89435a8e530f1f

Observation cf97cd27-ac38-4b85-983d-b3806904a469 · outbound

This paper cites V oice conversion with just nearest neighbors,.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis V oice conversion with just nearest neighbors,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:29.075673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:28.634322Z digest=sha256:6ad0e116b76c9148f2ffa43fde82ac012b5bb8841b09690081a15a4b50184d7d

Observation a26d98ca-5205-48f0-99ea-61164e471c45 · outbound

This paper cites Available: https://eprints.soton.ac.uk/409695/.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis Available: https://eprints.soton.ac.uk/409695/

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:30.781759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:27.112518Z digest=sha256:2c23393217843a4aea20d5ce8c7abe9e936b88dc62c367e0d28fd84ca9e0a31d

Pith citing papers

Observation b3b01c96-e5a6-431e-a440-48ac43f529a0 · inbound

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis cites this paper.

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:00:28.939908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:00:25.595644Z digest=sha256:f64f699e3d0fec76ddb94d8ba8fded75877d43f94b5acf7008928e069136bdbf