Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T03:15:41.454335Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2606.16417.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T03:15:41.454335Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2f9bebd3-fbfd-41fa-91c8-a841f811f197 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b9a33956-6dad-46f1-9711-f7a52411cb98 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Neural codec language models are zero-shot text to speech synthesizers,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c5fcf4-aafa-43b3-881e-6f170b4b0c79 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Grad-tts: A diffusion probabilistic model for text-to-speech,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c188d5-f21a-4c39-a4b9-cec8d53f1de1 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b6d5e9-3a04-4460-bec4-761851ae30b0 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c83c0c63-cb90-42d8-827d-dff5dd9adeb3 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f0ef2729-683e-4fe6-965f-072aa067a56e · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Maskgct: Zero-shot text-to-speech with masked generative codec transformer,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf91cb4-1cf7-4c16-9bff-9bb0cab197cb · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Macst: Multi-accent speech synthesis via text transliteration for accent conversion,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c739d33-eeb4-4b38-b121-3e12537099e2 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Accent-VITS:accent transfer for end-to-end TTS
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66566f92-d02f-45cb-a2d9-0778f3692911 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction L2-GEN: A Neural Phoneme Paraphrasing Approach to L2 Speech Synthesis for Mispronunciation Diagnosis,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb421b7-4b9d-44ad-b43c-246eb5b90b1d · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3d1e8161-ebd2-4973-9e10-25181bbfde86 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Scalable controllable accented tts,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae13f620-be02-4072-97fc-541ad66801b5 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Controllable accented text-to- speech synthesis with fine and coarse-grained intensity rendering,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e653925e-4724-439d-8124-ad4a5e3b32bf · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f634226-f788-4357-9e34-9da8eea8df09 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction RAD-MMM: Multilingual Multiaccented Multispeaker Text To Speech,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c6b3035-c5cc-4e43-adb6-dbe1bc5a070b · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Robust speech recognition via large-scale weak supervision,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60e8067-62cc-4fc2-add7-5e48868d6342 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Unsupervised domain adaptation by backpropagation,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d79462a-0ce7-45cd-8803-b5b37c83efc4 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58df358f-9600-450d-a730-92629eae7b43 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Adaspeech: Adaptive text to speech for custom voice,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0f5472c-224d-43c2-8f90-6932194baf50 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Glow-tts: A generative flow for text-to-speech via monotonic alignment search,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 347a8adf-719c-45e0-8889-fea43ea76af7 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Con- former: Convolution-augmented Transformer for Speech Recognition,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39af0fb-5523-4cbe-8d50-811a253700ab · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Accentbox: Towards high- fidelity zero-shot accent generation,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fbda8ed-d6c5-420c-a552-861d8ecdcf62 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Gaussian Error Linear Units (GELUs)
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1aaf27e4-99a4-4c2c-a114-f3dce8d6a714 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Denoising diffusion probabilistic models,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2c69e8-7fec-44f8-944e-486a25129a83 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi- resolution spectrogram,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac402c1-8ddf-4dc5-b291-d4d89fc83152 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction AISHELL-3: A multi-speaker mandarin TTS corpus,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9140cea1-da72-4a91-87f7-7810f8eade45 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Decoupled weight decay regularization,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb663951-e853-45e6-b920-25a2fad96287 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Adam: A method for stochastic optimization,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f881d6-4399-4211-9ab9-77dba00ac3fe · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction Amphion: An open-source audio, music and speech generation toolkit,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6887fa67-7c13-4173-8bc5-8bcb0745a79d · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 896a6cdb-57b1-4ece-a285-f505c582e8c8 · outbound
Joycent: Diffusion-based Accent TTS without Accented Phone Prediction CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.