Pith. sign in

Paper Citation Record · LEDGER

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2606.16417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.16417 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:15:41.454335Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f9bebd3-fbfd-41fa-91c8-a841f811f197 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.958536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:bda5b556dde3652443bfb6b6482fee7c17bdc1d28447f65d41bcf0782d3e9c30

Observation b9a33956-6dad-46f1-9711-f7a52411cb98 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Neural codec language models are zero-shot text to speech synthesizers,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:752f53e91e2d176ad652f120d33fd80be8431ca58b821f8fe5ec026e6c88389a

Observation 85c5fcf4-aafa-43b3-881e-6f170b4b0c79 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Grad-tts: A diffusion probabilistic model for text-to-speech,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:6d66a5c419e13af00eb4a9038b238d38a8bc0562e41f02c3e8754cc473704237

Observation d9c188d5-f21a-4c39-a4b9-cec8d53f1de1 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:1151b4a31fb6d4d47d8930d989fc0027f7eb3331ddf982e24fd83f3cbed3b1ce

Observation f4b6d5e9-3a04-4460-bec4-761851ae30b0 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.949342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:9f186f1bbd681351b02bb9419a7582ff44d3e5c3290f97b0d2cfa05c32728522

Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · outbound

This paper cites IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.958514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:360d6b208fca262df66852d73053b682d93b3e1524d79422be8798c120479922

Observation f0ef2729-683e-4fe6-965f-072aa067a56e · outbound

This paper cites Maskgct: Zero-shot text-to-speech with masked generative codec transformer,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Maskgct: Zero-shot text-to-speech with masked generative codec transformer,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:c159a5b1d78bee6d5bb0ba1eb2780e3370e33c88a8106dbc2894637b8b9ffddb

Observation ddf91cb4-1cf7-4c16-9bff-9bb0cab197cb · outbound

This paper cites Macst: Multi-accent speech synthesis via text transliteration for accent conversion,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Macst: Multi-accent speech synthesis via text transliteration for accent conversion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:64ea55a83bc8f15d4764e53a85bcea7d819f598f15990e87b9532e828976f420

Observation 8c739d33-eeb4-4b38-b121-3e12537099e2 · outbound

This paper cites Accent-VITS:accent transfer for end-to-end TTS.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Accent-VITS:accent transfer for end-to-end TTS

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.945348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:e6efb70bc447080f01ff2838981c0ab5c92715f775f6af2bcf54af1b18fd8d8f

Observation 66566f92-d02f-45cb-a2d9-0778f3692911 · outbound

This paper cites L2-GEN: A Neural Phoneme Paraphrasing Approach to L2 Speech Synthesis for Mispronunciation Diagnosis,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction L2-GEN: A Neural Phoneme Paraphrasing Approach to L2 Speech Synthesis for Mispronunciation Diagnosis,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:6ea06ae0c3ca21edbe8210c9c315c143b3df74867e8d0911d511d13a1f50a1b8

Observation 9cb421b7-4b9d-44ad-b43c-246eb5b90b1d · outbound

This paper cites Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.938729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:e733dc8cbf6a999f80f8d7cb9f93b62933d8b5d47fbf8527bd3a9035d41747ac

Observation 3d1e8161-ebd2-4973-9e10-25181bbfde86 · outbound

This paper cites Scalable controllable accented tts,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Scalable controllable accented tts,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:1e8fbe550209728d7d6609f0e2f5f44e01080d0f603e3cafdf888b35465fcd34

Observation ae13f620-be02-4072-97fc-541ad66801b5 · outbound

This paper cites Controllable accented text-to- speech synthesis with fine and coarse-grained intensity rendering,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Controllable accented text-to- speech synthesis with fine and coarse-grained intensity rendering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:be16273335733c9d815d2140b901680ec76bb5e1bbe49dc9dff847de992ff542

Observation e653925e-4724-439d-8124-ad4a5e3b32bf · outbound

This paper cites DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.955633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:cbace30ed4cc9693f3a760c187901b9aff4912feb3423ae0f8421e62e88dcd2d

Observation 3f634226-f788-4357-9e34-9da8eea8df09 · outbound

This paper cites RAD-MMM: Multilingual Multiaccented Multispeaker Text To Speech,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction RAD-MMM: Multilingual Multiaccented Multispeaker Text To Speech,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:b76864884670feb3a629bbcecd2998de9c96de33a8356c482acf9760d08c02d1

Observation 3c6b3035-c5cc-4e43-adb6-dbe1bc5a070b · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Robust speech recognition via large-scale weak supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:7b53d3d62203becd2263538bdc329ef3f3de1006c2536e65815871e2b749deeb

Observation b60e8067-62cc-4fc2-add7-5e48868d6342 · outbound

This paper cites Unsupervised domain adaptation by backpropagation,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Unsupervised domain adaptation by backpropagation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:f6cd8489c50f2cd73074ad91c96e19d9d199b8d1daff8dbdc1035b594ee7ecfc

Observation 2d79462a-0ce7-45cd-8803-b5b37c83efc4 · outbound

This paper cites Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:250acfea83c8b18818701cb1c377661567e57ec6d2a5a8db726b13be0c6caf4e

Observation 58df358f-9600-450d-a730-92629eae7b43 · outbound

This paper cites Adaspeech: Adaptive text to speech for custom voice,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Adaspeech: Adaptive text to speech for custom voice,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:7f72182ecf906f536bd8d391bfeb1ed66ec5ce411bf57193a6f39e0c2706f739

Observation c0f5472c-224d-43c2-8f90-6932194baf50 · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:507db2bcf241aa9791fdff31c8af4a6070164c7c149772a4926beb09078f7126

Observation 347a8adf-719c-45e0-8889-fea43ea76af7 · outbound

This paper cites Con- former: Convolution-augmented Transformer for Speech Recognition,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Con- former: Convolution-augmented Transformer for Speech Recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:1c4f9eb0f96cdbe6e7b998226c1b8affb74f56d2464930c985323e120156476f

Observation e39af0fb-5523-4cbe-8d50-811a253700ab · outbound

This paper cites Accentbox: Towards high- fidelity zero-shot accent generation,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Accentbox: Towards high- fidelity zero-shot accent generation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:75ad3f54f076e2db9ff2ac6209f06c759f46a2246f261e6cce2e3b03dde2c487

Observation 9fbda8ed-d6c5-420c-a552-861d8ecdcf62 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Gaussian Error Linear Units (GELUs)

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.952252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:f16803bc2945b004dfb2c9261203f04febb2ff4f02bab8c92e907397ddf60cbd

Observation 1aaf27e4-99a4-4c2c-a114-f3dce8d6a714 · outbound

This paper cites Denoising diffusion probabilistic models,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Denoising diffusion probabilistic models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:bf510ea2f9c47ba1eb14663f4316d703c8f0d2028d6739d7e2ade791703d138f

Observation 6d2c69e8-7fec-44f8-944e-486a25129a83 · outbound

This paper cites Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi- resolution spectrogram,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi- resolution spectrogram,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:a1d3c5f6fdd27080dfedd289753e231f882575529aacf73fafc71c763f49e2df

Observation 1ac402c1-8ddf-4dc5-b291-d4d89fc83152 · outbound

This paper cites AISHELL-3: A multi-speaker mandarin TTS corpus,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction AISHELL-3: A multi-speaker mandarin TTS corpus,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:4a908ba4bc4c1b515b50c031c77a8f6184062536365e2918c5235d27e813260c

Observation 9140cea1-da72-4a91-87f7-7810f8eade45 · outbound

This paper cites Decoupled weight decay regularization,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Decoupled weight decay regularization,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:6f6b2a712a6e7529e9249cace97fa369771f681076cd4a06dc880f4219fc76ad

Observation bb663951-e853-45e6-b920-25a2fad96287 · outbound

This paper cites Adam: A method for stochastic optimization,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Adam: A method for stochastic optimization,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:cf0811e7e9ec0ff46dd70c1281577f03edb4dc6dbb594c886f2a08fb51becef6

Observation 66f881d6-4399-4211-9ab9-77dba00ac3fe · outbound

This paper cites Amphion: An open-source audio, music and speech generation toolkit,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Amphion: An open-source audio, music and speech generation toolkit,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:07f23f69c5f8cbf038c5aaa3e72f658f7858cfe1a7320c606ed1431246f31232

Observation 6887fa67-7c13-4173-8bc5-8bcb0745a79d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T03:15:41.454335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:bc5c6f0a73e8efd158af57ce5eaedb13a038a1830553385507755e5930905c9e

Observation 896a6cdb-57b1-4ece-a285-f505c582e8c8 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T18:08:46.955943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:bc27bdb6ad18aabda88423cabd8d4a238708ce745962a584b1ecb1bab65c4f86

Pith citing papers

No inbound Pith citation observations are available.