Pith. sign in

Paper Citation Record · LEDGER

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model

As of 15 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2412.03074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03074 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:50.050472Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28aeab77-1755-46f0-a88a-d109bc6cc833 · outbound

This paper cites A Brief Overview of Unsupervised Neural Speech Representation Learning.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model A Brief Overview of Unsupervised Neural Speech Representation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.884332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.884332Z digest=sha256:c796658b41401e45f0738b2b06b75137aa7ca57118c19fe8d92c3ab3e2586fc9

Observation 24e6162f-c0ef-409a-837a-d5076b57dfee · outbound

This paper cites Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.890550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.890550Z digest=sha256:1fd04515725f4055ed5c9dade92ddfb3877ee6845e7bf0fc8abbb1cd7d12e687

Observation ff9c97c0-aaa2-497e-b7b9-dd602b6bb6bc · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.895990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.895990Z digest=sha256:7ca2871488ec57fe4fd76e6a085609036fceb7536b8b4f0b303d69fc362effca

Observation 2596427f-2627-40a1-8ddc-01f464ec602d · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.610616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.901188Z digest=sha256:61bab35a1bb41aeffcc11c4bda1671cacf8bc5de78f8192d26380aae1aeb2987

Observation 6b041915-aab1-493c-951d-0ce34066dfd5 · outbound

This paper cites Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.906849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.906849Z digest=sha256:c15be88859141b8d3741ee1da94a07d0b97929efa610c2e2bab05e2021c6ffda

Observation 4914c130-3b3c-4289-a90b-8357c6324bc6 · outbound

This paper cites On generative spoken language modeling from raw audio,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model On generative spoken language modeling from raw audio,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.912368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.912368Z digest=sha256:9385437968ce27586898dfaeedcf790e220eafa2fbc84506f236116bd70c7bf6

Observation c9f2b912-7f0a-48b6-8c94-7df3355f3f90 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Representation Learning with Contrastive Predictive Coding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.918052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.918052Z digest=sha256:23cdb59006ea65cf2b6b689e4fa25289baab656bcde4089810485abfd9de313a

Observation 7fd38301-1d33-4570-8c13-dad36764ceb5 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.926589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.926589Z digest=sha256:dff995138332bb08453e03bc829606c69fe4532e5f8e01f6338c2e3e9bf8893a

Observation a807a468-54ee-468e-a89b-502e821a3a4e · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.932289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.932289Z digest=sha256:337caa8e56fc5645d372b66c947b1f216a796ec0e8d8361445c3e539eb9db5f3

Observation 424690e8-0ec4-49c8-bbfc-5cd567793e8c · outbound

This paper cites Attention is all you need,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Attention is all you need,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.532944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.937783Z digest=sha256:95648227f84baeb955f157882df25656a9333c95cc81071ba3b5056a70e335e9

Observation 62613339-6e88-422d-b6f7-bb67e2a1afbc · outbound

This paper cites The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.513108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.943099Z digest=sha256:29632236cd8b6402a0865b6ba6ff2970364c826fb93003ee6fc37604c48b59d4

Observation 426bac3e-2054-43a0-9b94-25b6876fb6bf · outbound

This paper cites The zero resource speech challenge 2021: Spoken language mod- elling,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model The zero resource speech challenge 2021: Spoken language mod- elling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.494312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.948890Z digest=sha256:b98b9705e384d4960f0c1f98c1c3c7a275b00b2ce6b9c5220a51669d32cb1bf6

Observation 9e2a6a93-268c-4b36-aaeb-7d701ee17560 · outbound

This paper cites Self- supervised language learning from raw audio: Lessons from the zero resource speech challenge,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Self- supervised language learning from raw audio: Lessons from the zero resource speech challenge,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.476981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.954166Z digest=sha256:28b043b8545edae48d034125326e1639b9fbba2fae39c7e92bb51a00284e2337

Observation afa51aa1-87a9-461f-9db8-b84826267a1a · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model A comparison of discrete and soft speech units for improved voice conversion,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.457515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.958904Z digest=sha256:f1e5aca8723573ab54493949cef3e90571116fa60d0cc5bd215d13ae03ff571e

Observation b10deed6-9dcb-4f28-95ce-5f99ed039354 · outbound

This paper cites Textless direct speech-to- speech translation with discrete speech representation,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Textless direct speech-to- speech translation with discrete speech representation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.436813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.963908Z digest=sha256:ae6c0dfc0a728249036f6e66a32c4d3303d59dc333ea50fb8f22991cc20fa0db

Observation af9ee82f-fc6c-41e9-8eba-ca1a7d11538d · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.414008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.969832Z digest=sha256:28ab233ae78cccdcca5f47df85888498dcf5d50bf0110c4d73a62d16fe9beb25

Observation be39764a-d380-48bb-bc88-d67353c8bfb5 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model LibriSpeech: An ASR corpus based on public domain audio books,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.394880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.982938Z digest=sha256:9c88c73b44d9a41b162ebc8b222dc720b600716ce523a726a1e254f51091b4e1

Observation 3b8d78e1-02aa-4d7d-9084-518504527449 · outbound

This paper cites Ito and L.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Ito and L

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.373300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.987883Z digest=sha256:00ed7e802fa18e5f42909544b800fb7b0de254ee2b06d6b31556c58f3d970289

Observation 26165c73-cf25-42f4-95ea-76ade09ae3eb · outbound

This paper cites an unresolved cited work.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:50.354281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.992872Z digest=sha256:62312d90867c592eba363c5cf0da880484aabe6e81576ae75b46234cf62e229e

Observation 9af1f3d6-57ea-4a1c-979a-6a410715f544 · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.998068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.998068Z digest=sha256:23eafa70d09a58b071c687f1c6d2e227b153293c5645113c5cd28a3921ca9e5f

Observation aba34485-f38b-4812-a5ff-9bf6ab45dbce · outbound

This paper cites JVS corpus: free Japanese multi-speaker voice corpus.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JVS corpus: free Japanese multi-speaker voice corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.005341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.005341Z digest=sha256:8011d7e6ad778e14d6e4402b4c495bb6d81bfc38bd32c8ed364d17a88a4d5320

Observation 084663e5-cfc0-46c9-b2b8-d3d1fff071c9 · outbound

This paper cites Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.011574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.011574Z digest=sha256:375cbeb1da68b8b70f1455af2222033ad274df9e6328cfc9c179c335e63450ba

Observation 904682ab-e720-4882-9125-c6a52e1496ea · outbound

This paper cites J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.323938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:50.016467Z digest=sha256:d41dc1a78e10a5cfd0ec5aa9dd270cbee6103820573e4508f60e661f35bed6d6

Observation 27ce4de1-9389-43c3-ae3a-80f37022a628 · outbound

This paper cites Radford, K.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Radford, K

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.022731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.022731Z digest=sha256:1c8609f28e717e22602f44683b8efe2aaddb2a376e4f788b8042021cca7cf058

Observation 7b2727b2-2a2b-486f-a526-48cf884fe8a8 · outbound

This paper cites Investi- gation of robustness of hubert features from different layers to domain, accent and language variations,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Investi- gation of robustness of hubert features from different layers to domain, accent and language variations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.294196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:50.027954Z digest=sha256:926a3a930e0f36094199718a1ef1414c2f60295184e086789e310cdd882ce36d

Observation af4827dd-6f50-4f16-99e2-0a9732380725 · outbound

This paper cites ContentVec: An improved self-supervised speech representation by dis- entangling speakers,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model ContentVec: An improved self-supervised speech representation by dis- entangling speakers,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.277654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:50.033061Z digest=sha256:6fa1e3fb92b75ef71809527826f6cfa9257a0c27014eea909ea4f870d3cf19a3

Observation c61e9b87-1dbc-4368-81d9-e557c5c21a8c · outbound

This paper cites JSSS: free Japanese speech corpus for summarization and simplification.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JSSS: free Japanese speech corpus for summarization and simplification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.039933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.039933Z digest=sha256:8aa0d0c9f8a22dde23fbbe6f618a1a671ea1a0005895f289f714c130a28ee7f4

Observation 3828493d-eb62-45b8-a824-91f186f82db4 · outbound

This paper cites Universal phone recognition with a multilingual allophone system,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Universal phone recognition with a multilingual allophone system,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.258903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:50.045137Z digest=sha256:326ed5b5860da149ed41df979788558e99d9e18ed69f6073e1a3b4d5a5439b1c

Observation 7aa3af17-3970-46e3-8975-52bbd7399cf4 · outbound

This paper cites Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.240815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:50.050472Z digest=sha256:f393d9ee1b4a24f3dcecc62d4b49d9064e05ff81220e86a8572a5eb610ba52ca

Observation a362ed75-5413-4e62-9c3f-f8c4e3014f7f · outbound

This paper cites an unresolved cited work.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Unresolved cited work

Reference 3042

Resolution
verified exact
doi, observed 2026-08-11T22:51:50.091253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T22:51:49.977002Z digest=sha256:c4a84258b4f823aa62356cbee1691685eee900ef4e209ca9d582b0b0af8a29af

Pith citing papers

No inbound Pith citation observations are available.