Pith. sign in

Paper Citation Record · LEDGER

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2606.16417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.16417 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:15:41.454335Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f9bebd3-fbfd-41fa-91c8-a841f811f197 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.958536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:b5b1407ab2829104afa4d1bc8f66ddc71aba96567495b001a66c2893d8fa4319

Observation b9a33956-6dad-46f1-9711-f7a52411cb98 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Neural codec language models are zero-shot text to speech synthesizers,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:091e1f6c691c256bfd1efe7bcd95a7e0316b9646be26a6ecbd509d9ecb673080

Observation 85c5fcf4-aafa-43b3-881e-6f170b4b0c79 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Grad-tts: A diffusion probabilistic model for text-to-speech,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:272dca3521046bb3881a624d1e5efb7be141c87daec14b07c2f72b5a24f8c4aa

Observation d9c188d5-f21a-4c39-a4b9-cec8d53f1de1 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:e8cb4cadb4a8042ca47e8b03e4fbc9999165d25512d9ba77fa5ea12a75d27456

Observation f4b6d5e9-3a04-4460-bec4-761851ae30b0 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.949342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:9b227f7dab898da15cc99bd918747a0e52b1a164ae2a75550e42c503b3943d37

Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · outbound

This paper cites IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.958514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:a1fc50715f3d8a82297d08fa7016e99504c1cbb187431b7459e0f5dc53ecfb59

Observation f0ef2729-683e-4fe6-965f-072aa067a56e · outbound

This paper cites Maskgct: Zero-shot text-to-speech with masked generative codec transformer,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Maskgct: Zero-shot text-to-speech with masked generative codec transformer,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:10d24bc30080312bdeceea6031bfa49e08074b5b4460c4c933bed139abb3bfa2

Observation ddf91cb4-1cf7-4c16-9bff-9bb0cab197cb · outbound

This paper cites Macst: Multi-accent speech synthesis via text transliteration for accent conversion,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Macst: Multi-accent speech synthesis via text transliteration for accent conversion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:bf247561486460654fed818927bbaf2e56d4ce64f692d53927bb980447f8f275

Observation 8c739d33-eeb4-4b38-b121-3e12537099e2 · outbound

This paper cites Accent-VITS:accent transfer for end-to-end TTS.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Accent-VITS:accent transfer for end-to-end TTS

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.945348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:87a1a85b20e06a562d49ea9047dc6ee0fab3dfaaf41b078cc5cbe2fcb9cedf2b

Observation 66566f92-d02f-45cb-a2d9-0778f3692911 · outbound

This paper cites L2-GEN: A Neural Phoneme Paraphrasing Approach to L2 Speech Synthesis for Mispronunciation Diagnosis,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction L2-GEN: A Neural Phoneme Paraphrasing Approach to L2 Speech Synthesis for Mispronunciation Diagnosis,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:5911d674b9b83313d70dc6112df68364c99f8f56118cd20f450e78abaf981b3e

Observation 9cb421b7-4b9d-44ad-b43c-246eb5b90b1d · outbound

This paper cites Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.938729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:674b4bfe66ba6e7a4e7f305b2b4109b0f81c0904028e0be5fc761e5dd9d17aed

Observation 3d1e8161-ebd2-4973-9e10-25181bbfde86 · outbound

This paper cites Scalable controllable accented tts,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Scalable controllable accented tts,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:3c780bd5b3b087f2ff12600dc77e2dff3e456a401ca85f009abdf424d0a495d2

Observation ae13f620-be02-4072-97fc-541ad66801b5 · outbound

This paper cites Controllable accented text-to- speech synthesis with fine and coarse-grained intensity rendering,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Controllable accented text-to- speech synthesis with fine and coarse-grained intensity rendering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:5d366f08324d45334bf7427535860caf045377144d3b49ef71bd0aaaebe37c23

Observation e653925e-4724-439d-8124-ad4a5e3b32bf · outbound

This paper cites DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.955633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:903f6c7de37173a6eaa689c818fe281dd6273aa6b8d710ab933ab5d1aecbc12c

Observation 3f634226-f788-4357-9e34-9da8eea8df09 · outbound

This paper cites RAD-MMM: Multilingual Multiaccented Multispeaker Text To Speech,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction RAD-MMM: Multilingual Multiaccented Multispeaker Text To Speech,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:a07689da5406cf6077530c35e41b4ad5e9f7b1f5e750a203521428cca9cdb69e

Observation 3c6b3035-c5cc-4e43-adb6-dbe1bc5a070b · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Robust speech recognition via large-scale weak supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:0b6203eac0d8a627f1542b79a7115ba1850a11fa483d4cb0b2fc8139c837100d

Observation b60e8067-62cc-4fc2-add7-5e48868d6342 · outbound

This paper cites Unsupervised domain adaptation by backpropagation,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Unsupervised domain adaptation by backpropagation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:59466201d5fde133de847f6f8b3fbfdb062bf0dfeac5b41cbbc0e3d62ea26abd

Observation 2d79462a-0ce7-45cd-8803-b5b37c83efc4 · outbound

This paper cites Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:f965e8beecd853482e3099d4f418c720152cfe00665a3ac25d38008a3b7d90ff

Observation 58df358f-9600-450d-a730-92629eae7b43 · outbound

This paper cites Adaspeech: Adaptive text to speech for custom voice,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Adaspeech: Adaptive text to speech for custom voice,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:b4628320dd7fe6a4be18569b1f04302ee2485860164e843337a623861300ba1c

Observation c0f5472c-224d-43c2-8f90-6932194baf50 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:37107093ba316c5513552457819d553439760a73c39a4c65ad03bd9c021fe28e

Observation 347a8adf-719c-45e0-8889-fea43ea76af7 · outbound

This paper cites Con- former: Convolution-augmented Transformer for Speech Recognition,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Con- former: Convolution-augmented Transformer for Speech Recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:9478269ebf7bc8ed071f5a6aabc43bb6d3bd5f3488af53cfaf78691533974f74

Observation e39af0fb-5523-4cbe-8d50-811a253700ab · outbound

This paper cites Accentbox: Towards high- fidelity zero-shot accent generation,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Accentbox: Towards high- fidelity zero-shot accent generation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:6844fe12e8e595f7729e75bb7ad8b330fe28653a881a52804d0f5671e77659d8

Observation 9fbda8ed-d6c5-420c-a552-861d8ecdcf62 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Gaussian Error Linear Units (GELUs)

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.952252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:86dc8a3545a26330594b0f95251ca43ae626d55cde5f0dd4077ea5160e41f29e

Observation 1aaf27e4-99a4-4c2c-a114-f3dce8d6a714 · outbound

This paper cites Denoising diffusion probabilistic models,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Denoising diffusion probabilistic models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:01319cb5536e730ccf3a52ec07653c7a4a53511b7de0490e0ccf806e94a0eab0

Observation 6d2c69e8-7fec-44f8-944e-486a25129a83 · outbound

This paper cites Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi- resolution spectrogram,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi- resolution spectrogram,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:d869e89944738c030ca1edf03a0e2d89bfea446bb2b9202d3b9d658894e1aac6

Observation 1ac402c1-8ddf-4dc5-b291-d4d89fc83152 · outbound

This paper cites AISHELL-3: A multi-speaker mandarin TTS corpus,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction AISHELL-3: A multi-speaker mandarin TTS corpus,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:cfaccccb45ac2d25031b89515b9998affb45c38e2c4d40144a0ba53dff9afbcd

Observation 9140cea1-da72-4a91-87f7-7810f8eade45 · outbound

This paper cites Decoupled weight decay regularization,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Decoupled weight decay regularization,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:8c3a6ca47cb8b40687428e52ff08731f4f3b1bab32bf6e980d254f5347e28d2d

Observation bb663951-e853-45e6-b920-25a2fad96287 · outbound

This paper cites Adam: A method for stochastic optimization,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Adam: A method for stochastic optimization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:282e4cbe6a165a5d2afc17bd4f76d820a692b0fb15fc39e5427a08e430512d8a

Observation 66f881d6-4399-4211-9ab9-77dba00ac3fe · outbound

This paper cites Amphion: An open-source audio, music and speech generation toolkit,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Amphion: An open-source audio, music and speech generation toolkit,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:6ce3a5b045321bd40f2be3eb20ae229a05fe305ba4ba4f000f93b4577b8e39d4

Observation 6887fa67-7c13-4173-8bc5-8bcb0745a79d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:c785f9a632bcbf408f58f6360c40235d41f7558740c63a76c50ebf7fc0ea799e

Observation 896a6cdb-57b1-4ece-a285-f505c582e8c8 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.955943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:2a322940cda3f0d6726198c5b038e66e2c2a336eb9f3e4e2355e63b5db8376c8

Pith citing papers

No inbound Pith citation observations are available.