Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:09.630347Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2507.22612.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:09.630347Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:07.882293Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T11:34:10.011767Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 769360fc-7bb6-4889-b742-f172efc4899d · outbound
Adaptive Duration Model for Text Speech Alignment A typical TTS system includes an encoder, a decoder, and an alignment mechanism linking linguistic and acoustic representations [3–6]
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e822238e-0ba3-457f-bee2-0dc2c323cba7 · outbound
Adaptive Duration Model for Text Speech Alignment Adaptive Duration Model for Text Speech Alignment
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 50b46a03-b979-44fc-86c3-0b452df8c76e · outbound
Adaptive Duration Model for Text Speech Alignment We use Premium and Basic subsets of Wenet- Speech4TTS as our experiment dataset
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 458ac7de-2dbe-4a37-80c5-c02a70b6ea43 · outbound
Adaptive Duration Model for Text Speech Alignment Durformer outperforms baseline methods with respect to efficiency and accuracy
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e2b5885-beda-4b19-8418-c051f936a75c · outbound
Adaptive Duration Model for Text Speech Alignment One tts align- ment to rule them all,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c5d9294-2c5c-4955-820b-fd37b48e7a1b · outbound
Adaptive Duration Model for Text Speech Alignment Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b85757-2026-4848-bdb3-4bc26b481cdf · outbound
Adaptive Duration Model for Text Speech Alignment Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc12977-4893-42a9-9c18-67a74ec62837 · outbound
Adaptive Duration Model for Text Speech Alignment Fastspeech: Fast, robust and controllable text to speech,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 03f6445e-4212-4a15-8f2a-1c798a8804c9 · outbound
Adaptive Duration Model for Text Speech Alignment Fastpitch: Parallel text-to-speech with pitch prediction,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 10183caa-b808-4807-92b6-767070d9fb8e · outbound
Adaptive Duration Model for Text Speech Alignment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849ba4da-4596-46cd-8456-028e7175abd4 · outbound
Adaptive Duration Model for Text Speech Alignment Location-relative attention mechanisms for ro- bust long-form speech synthesis,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 368df377-e60e-4603-b996-af6b64354af5 · outbound
Adaptive Duration Model for Text Speech Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ddde11-7728-46f0-97c8-631b877d30c0 · outbound
Adaptive Duration Model for Text Speech Alignment NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3f2b53-4432-4704-bf13-5ce8b7d01f08 · outbound
Adaptive Duration Model for Text Speech Alignment Non-autoregressive neural text-to-speech,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21c31295-447d-46b1-8bc5-31f1d935c5ab · outbound
Adaptive Duration Model for Text Speech Alignment Durian-e 2: Duration informed attention net- work with adaptive variational autoencoder and adver- sarial learning for expressive text-to-speech synthesis,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6cbc73b0-781d-4ae4-99b2-0c2dc1b537e5 · outbound
Adaptive Duration Model for Text Speech Alignment Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 10f337af-c4fa-42b6-bcc3-fe35ea5ccb08 · outbound
Adaptive Duration Model for Text Speech Alignment Simple-tts: End-to-end text-to-speech synthesis with latent diffusion,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f2d6d4c-ffdd-4fe9-8d08-605b21eeb050 · outbound
Adaptive Duration Model for Text Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f29d60-5a31-40d7-a324-22c2f553eea3 · outbound
Adaptive Duration Model for Text Speech Alignment Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cbfaa89b-f5e3-42c5-8734-c4c308da99e2 · outbound
Adaptive Duration Model for Text Speech Alignment Portaspeech: Portable and high-quality generative text-to-speech,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a65be6c9-43c0-465b-a52b-8a8010999c00 · outbound
Adaptive Duration Model for Text Speech Alignment V oicebox: Text-guided multilingual universal speech generation at scale,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 752c8ba0-a959-4247-8d1e-c3737aee4cf3 · outbound
Adaptive Duration Model for Text Speech Alignment MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb08e55-b39c-4f59-ab51-4d28892893d5 · outbound
Adaptive Duration Model for Text Speech Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56b91bcd-6dbd-4bf0-b070-a152a73ffcc4 · outbound
Adaptive Duration Model for Text Speech Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a86135-f78e-4daf-8f2e-11ed8cd36458 · outbound
Adaptive Duration Model for Text Speech Alignment WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e822238e-0ba3-457f-bee2-0dc2c323cba7 · inbound
Adaptive Duration Model for Text Speech Alignment Adaptive Duration Model for Text Speech Alignment
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.