Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:22:27.950994Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.22746.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:22:27.950994Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6229081c-1426-49c8-90a8-da36704ba759 · outbound
Next Tokens Denoising for Speech Synthesis Better speech synthesis through scaling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b306d4d-dbb8-43b9-b58d-7acdf6d27af1 · outbound
Next Tokens Denoising for Speech Synthesis High Fidelity Neural Audio Compression
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c112359-d552-43bb-9d01-693ec47a680d · outbound
Next Tokens Denoising for Speech Synthesis Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3265b4f7-b4a9-463f-9301-ac12a400f9da · outbound
Next Tokens Denoising for Speech Synthesis Mean Flows for One-step Generative Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce0731a6-61e7-4eb0-bc97-2580a1e39188 · outbound
Next Tokens Denoising for Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b581fd-1e21-41bf-b2d1-564f140f6444 · outbound
Next Tokens Denoising for Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5fa5a1c-84dd-47af-8e12-f28339bbaec8 · outbound
Next Tokens Denoising for Speech Synthesis Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2df25a-9a0d-4912-8631-37d282ff4376 · outbound
Next Tokens Denoising for Speech Synthesis istftnet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16933749-f487-4e28-8067-c097a54726a6 · outbound
Next Tokens Denoising for Speech Synthesis Understanding DDPM Latent Codes Through Optimal Transport
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3763f71c-b841-4b96-8274-d7fe01dee4d1 · outbound
Next Tokens Denoising for Speech Synthesis PromptTTS 2: Describing and Generating Voices with Text Prompt
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8b244c-560b-408d-a460-e5e525c30581 · outbound
Next Tokens Denoising for Speech Synthesis Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1877fbd7-3822-4a5d-95c0-d4b452f46c9f · outbound
Next Tokens Denoising for Speech Synthesis Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6132fe-1c44-4db0-a7a9-dc422b4ccf88 · outbound
Next Tokens Denoising for Speech Synthesis DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3549354-974d-4d53-b100-5142f17f4f53 · outbound
Next Tokens Denoising for Speech Synthesis Autoregressive Speech Synthesis without Vector Quantization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4b868e7-0546-4ed3-be84-a08794a3fc41 · outbound
Next Tokens Denoising for Speech Synthesis Finite Scalar Quantization: VQ-VAE Made Simple
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f01925d6-4100-4867-b4ef-5eb009bfc73c · outbound
Next Tokens Denoising for Speech Synthesis Scaling Transformers for Low-Bitrate High-Quality Speech Coding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d123e31a-3705-44d4-864b-225d16ba3970 · outbound
Next Tokens Denoising for Speech Synthesis Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faba2588-9e08-487e-b264-81415b3455e9 · outbound
Next Tokens Denoising for Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4706dce0-0b29-472a-80c5-38080d13a8c1 · outbound
Next Tokens Denoising for Speech Synthesis Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d503ec9f-d481-4f10-aab6-d5a8b36a0787 · outbound
Next Tokens Denoising for Speech Synthesis Gemini: A Family of Highly Capable Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7001ad77-5a82-47ef-a1b7-01fb19ad16d9 · outbound
Next Tokens Denoising for Speech Synthesis Improving and generalizing flow-based generative models with minibatch optimal transport
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efbba303-682e-4bdb-9678-fe4e6a456f05 · outbound
Next Tokens Denoising for Speech Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2a6f30a-9185-4d39-bc60-d608c1c65b1c · outbound
Next Tokens Denoising for Speech Synthesis Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3d6bb98-fc36-488c-8faf-0489bb1a1767 · outbound
Next Tokens Denoising for Speech Synthesis Lumos-1: On autoregressive video generation from a unified model perspective
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57703305-fd3c-4bd1-8837-dd90a5334c5e · outbound
Next Tokens Denoising for Speech Synthesis Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12782ae2-731c-4d75-bc15-cfd92c294d8f · outbound
Next Tokens Denoising for Speech Synthesis Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53ce42bb-f56c-4f05-b7e7-2ad9bf1494fd · outbound
Next Tokens Denoising for Speech Synthesis Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4517a79-8173-4f65-8551-ffcd886f9f98 · outbound
Next Tokens Denoising for Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e844282-1ca3-4230-b734-949e062f358f · outbound
Next Tokens Denoising for Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebfe630b-423b-4645-a172-64b64bff40ab · outbound
Next Tokens Denoising for Speech Synthesis MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8d21cbe-36a0-4a83-ae5c-37f0360f7ef3 · outbound
Next Tokens Denoising for Speech Synthesis AdaSpeech: Adaptive Text to Speech for Custom Voice
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c650765-46aa-4e7f-868f-ebc596c3d331 · outbound
Next Tokens Denoising for Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572e330a-4104-4d2b-b35e-f211c1787355 · outbound
Next Tokens Denoising for Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9ba83c-6772-4838-9f60-dc93aea1ba61 · outbound
Next Tokens Denoising for Speech Synthesis Language models are few-shot learners
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f10129-0acc-4fc5-a3d2-d6a4e1b0060c · outbound
Next Tokens Denoising for Speech Synthesis One Step Diffusion via Shortcut Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b58e0cd-015b-4b9c-a206-ecd55188d95f · outbound
Next Tokens Denoising for Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.