Pith. sign in

Paper Citation Record · LEDGER

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.08319.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08319 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:27:55.344163Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 118ce986-8683-4f67-b85c-914848249615 · outbound

This paper cites Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:28:00.303298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.174809Z digest=sha256:4b747747277a8c2c9ae60199f895e31fed9cb523c8a025cf932904c905fefcd1

Observation 076116c5-7b12-45cd-a9e9-c86b396fee5d · outbound

This paper cites A vector quantized ap- proach for text to speech synthesis on real-world spontaneous speech,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection A vector quantized ap- proach for text to speech synthesis on real-world spontaneous speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:28:00.139910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.234976Z digest=sha256:2289f276a35d40bde8c43ebe06774bd6dbf35c94864a47a16e7e2ee3b629f427

Observation 0ba30967-01a3-484a-aed5-4987af5df8b6 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:53.306898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:53.306898Z digest=sha256:f8d3e73403e7ef237285a8fc42a879ff75567f7ad9af0ef327eba185e2c70f25

Observation 4cce857c-0983-464f-aa9b-bf3fb64b0bad · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:53.371201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:53.371201Z digest=sha256:5c78220214f647cfaaee40df622ec4620a9e2c96776feaad3acaa40b2c00a62d

Observation a7e6c808-0da9-4274-8872-508aabc41bba · outbound

This paper cites JVS corpus: free Japanese multi-speaker voice corpus.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection JVS corpus: free Japanese multi-speaker voice corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:53.408284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:53.408284Z digest=sha256:cfcabf9f193d35ddba76f1a777c1ae9388fe719d8b063acb9140b786f7809367

Observation 36f5aa05-faf8-4f27-aaf1-02f64daaba7e · outbound

This paper cites SaSLaW: Dialogue speech corpus with audio-visual egocentric information toward environment-adaptive dialogue speech synthesis,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection SaSLaW: Dialogue speech corpus with audio-visual egocentric information toward environment-adaptive dialogue speech synthesis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:28:00.001626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.484224Z digest=sha256:ac5707aa1ba0c93b76efd3a2cefa541b54e18e25ee45512908e3ee84c835a3d6

Observation 9c5e6160-1e19-4dc4-bd19-5887920dc80a · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection LibriTTS: A Corpus Derived from LibriSpeech for Text- to-Speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:59.830933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.556671Z digest=sha256:01b454ba0abf56b7b2de0a584ba09fb81333b93262bc73686af2213929cea739

Observation 6c7597bb-db5c-49e4-abab-d249df0b0910 · outbound

This paper cites HUI-Audio-Corpus-German: A high quality TTS dataset,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection HUI-Audio-Corpus-German: A high quality TTS dataset,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:59.654498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.639708Z digest=sha256:bf8543ac74b74d7eeb001466886449542040658704ef2a062193e48b99c6a3ce

Observation 62f70b8d-8166-4ea9-ab05-5e6223c11989 · outbound

This paper cites Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Text-to-speech synthesis from dark data with evaluation-in-the-loop data selection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:59.459773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.707559Z digest=sha256:b631f06f7f9d73f97c27f666a45ec344ce749e4aa3f30ef347a0d4b94d7f85ed

Observation 3da493c0-b1b5-47cc-afeb-a8ec8a9fa604 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:53.805528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:53.805528Z digest=sha256:c098df2cc868e18df84a435bf1967577c69d40b5043bc1e4f49563e47694f180

Observation 7914dd03-a75c-4d5e-8f23-4166ece8c479 · outbound

This paper cites JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:27:56.070227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.881428Z digest=sha256:af05b41426a0e021aee0471a3b7ba2726fda9b1fcaf1f22a2a80c893c7abed17

Observation 6b85a835-d610-44f8-988e-a93350e38124 · outbound

This paper cites J-CHAT: Japanese large-scale spoken dialogue corpus for spoken dialogue language modeling,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection J-CHAT: Japanese large-scale spoken dialogue corpus for spoken dialogue language modeling,

Reference 12

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:27:55.798236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:53.985871Z digest=sha256:d0a8dd4660de0ba712876b228e9e68f11891a38ad4047e4ee2a5bd3fbb54099f

Observation d69bf310-070b-4da3-860b-9a5ff3263a86 · outbound

This paper cites Diversity-based core-set selection for text-to-speech with linguistic and acoustic fea- tures,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Diversity-based core-set selection for text-to-speech with linguistic and acoustic fea- tures,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:59.307447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.094640Z digest=sha256:d7ff36d734867c91bb8a68e15aa9607d1602f9c088fb70e9d90a3c01bdfb4564

Observation 3c4dc365-f55d-43d8-a1ee-2ef8d31ac6dd · outbound

This paper cites Deepcore: A comprehensive library for coreset selection in deep learning,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Deepcore: A comprehensive library for coreset selection in deep learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:59.104527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.212422Z digest=sha256:8f70d9e0e7fbdcd80d470d83b624100216b1261909edb90762548ce2fa8c10c2

Observation 06b51cf2-0a94-47f6-94e7-b1153d5eb4db · outbound

This paper cites Active learning is a strong baseline for data subset selection,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Active learning is a strong baseline for data subset selection,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:58.887393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.296066Z digest=sha256:9ceb66efae1ae69b33f10b268e574234c0bb91f40dbbd5d116ef20719337bea3

Observation 11ac99fe-fdc0-4a83-b293-35f9693a8e0a · outbound

This paper cites A survey of deep active learning,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection A survey of deep active learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:54.394528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:54.394528Z digest=sha256:c624c1e298b3042e3543a6ab4222427093f48416c45bfc10087a939d484f0785

Observation 9f5f5c20-6de6-40fc-a736-759e88e9676e · outbound

This paper cites X-vectors: Robust DNN embeddings for speaker recognition,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection X-vectors: Robust DNN embeddings for speaker recognition,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:54.435303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:54.435303Z digest=sha256:6e7ace6b544520d8fd002552fe688d7b881bf5a65c5dc5f1aaaf6a67f5fc5fc1

Observation ed20ec0e-f874-44e8-87e5-edbf9143349b · outbound

This paper cites Ctc- segmentation of large corpora for german end-to-end speech recogni- tion,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Ctc- segmentation of large corpora for german end-to-end speech recogni- tion,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:58.648333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.521562Z digest=sha256:5c161966d6483370810d55a92f2e324350e24955ea96d2b61355dc0da6515a23

Observation cb737d24-bba7-41c1-bd09-3916fc06a0f6 · outbound

This paper cites TTSOps: A closed- loop corpus optimization framework for training multi-speaker TTS models from dark data,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection TTSOps: A closed- loop corpus optimization framework for training multi-speaker TTS models from dark data,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:54.611739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:54.611739Z digest=sha256:2bd5bf2f9f3d2994fe73f7713c198dcfd429104d5b09fa1616e154562d5cf8d6

Observation 145f0d5d-2e25-44fc-a068-cd37d6de6161 · outbound

This paper cites Speaker generation,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Speaker generation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:58.324202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.656286Z digest=sha256:5799857d65933fcd65bff7ddcb407b2a7aa9bb24ad94bfebd0b8147bfe4c4d40

Observation 8362134a-5a44-41ed-84d7-06187d788644 · outbound

This paper cites Denoising diffusion probabilistic models,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Denoising diffusion probabilistic models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:58.082996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.696419Z digest=sha256:85ba4728a3d7a92d540ef7221deda3f7c3eb5ceda8d0fb9e21648c4aeff629cb

Observation b77b1d56-2df1-43db-a5f5-5d0b3149cfee · outbound

This paper cites Diffusion models are minimax optimal distribution estimators,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Diffusion models are minimax optimal distribution estimators,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:57.733283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.798528Z digest=sha256:301267aa006db61bfabf90946030f817c9a369265dbea32696b27a963b68cba3

Observation bbc22d8c-fbac-4029-b886-a39061c92b22 · outbound

This paper cites Diffusion models in vision: A survey,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection Diffusion models in vision: A survey,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:54.858512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:54.858512Z digest=sha256:033a300926b4f4b80191f2b0a7099fe785ca6bf79082b351ee9360904b00003d

Observation f265599f-8560-478f-bb09-eed3c25b8428 · outbound

This paper cites ITA corpus,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection ITA corpus,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:57.422205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.902307Z digest=sha256:7f4128c6b8f3826a76014dd36fd3432ba42145acec79ebd05b85659c13e97f93

Observation 9e0de4fd-0987-481c-9a7b-4b4e78aca670 · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection FastSpeech 2: Fast and high-quality end-to-end text to speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:57.190524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:54.971264Z digest=sha256:ed62e9025fc9ec890d56f80a4cb9b0e972b29b227d386af8628a25d507e5b8aa

Observation 7487de70-b9b8-43ba-a8b6-5cab2e462676 · outbound

This paper cites HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:57.046422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:55.058505Z digest=sha256:19ab393d8b93a936d41b9de29e06f10da831923f0e2669e94311edc416dfca19

Observation df9a00ab-2626-49c9-8afe-8d537cc4608e · outbound

This paper cites HiFi-GAN,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection HiFi-GAN,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:56.853905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:55.142489Z digest=sha256:b45e121abfc8cafbe711eb39d8b8cd852e951bee3f152864b562054a77c122c7

Observation 9d0b883e-c82b-455e-89cb-38b3568a60ed · outbound

This paper cites FastSpeech 2-JSUT,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection FastSpeech 2-JSUT,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:56.623802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:55.235654Z digest=sha256:840d67130ae9c7ac8e5a183b885781e2b1fc1f22b85d4182eca1650982997eb6

Observation 3e4dca15-a0b0-4d76-b67a-1ebfdfdbcebf · outbound

This paper cites x-vector,.

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection x-vector,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:27:56.351789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:27:55.344163Z digest=sha256:588e4421a32ca2c057472180cc0ea26b23e5e6d27e0537feeac4bae0a40f8619

Pith citing papers

No inbound Pith citation observations are available.