Pith. sign in

Paper Citation Record · LEDGER

Adaptive Duration Model for Text Speech Alignment

As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2507.22612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22612 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:09.630347Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:07.882293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:34:10.011767Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 769360fc-7bb6-4889-b742-f172efc4899d · outbound

This paper cites A typical TTS system includes an encoder, a decoder, and an alignment mechanism linking linguistic and acoustic representations [3–6].

Adaptive Duration Model for Text Speech Alignment A typical TTS system includes an encoder, a decoder, and an alignment mechanism linking linguistic and acoustic representations [3–6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:12.419908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:07.829227Z digest=sha256:5bcafea9467f5bb18eb023dee5c5e9e15a242d6c2c3088541899426d1d08040f

Observation e822238e-0ba3-457f-bee2-0dc2c323cba7 · outbound

This paper cites Adaptive Duration Model for Text Speech Alignment.

Adaptive Duration Model for Text Speech Alignment Adaptive Duration Model for Text Speech Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:10.121918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:07.882293Z digest=sha256:3b0777870a63c6028d0cc4802f15ba6af1a65801e1a4d70808daac53b353b68b

Observation 50b46a03-b979-44fc-86c3-0b452df8c76e · outbound

This paper cites We use Premium and Basic subsets of Wenet- Speech4TTS as our experiment dataset.

Adaptive Duration Model for Text Speech Alignment We use Premium and Basic subsets of Wenet- Speech4TTS as our experiment dataset

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T11:34:12.271952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:07.942094Z digest=sha256:f13a6d82ab853895d6143a09eaf1311714646c409557566ab4529f74a14e1a82

Observation 458ac7de-2dbe-4a37-80c5-c02a70b6ea43 · outbound

This paper cites Durformer outperforms baseline methods with respect to efficiency and accuracy.

Adaptive Duration Model for Text Speech Alignment Durformer outperforms baseline methods with respect to efficiency and accuracy

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:12.084758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.046867Z digest=sha256:46357703f4771026f65df7c7eba20c6fa2df5f1bd7919a0ae97f3bd806ef228e

Observation 0e2b5885-beda-4b19-8418-c051f936a75c · outbound

This paper cites One tts align- ment to rule them all,.

Adaptive Duration Model for Text Speech Alignment One tts align- ment to rule them all,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.898860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.114869Z digest=sha256:60de94b97a60d945209a31415e809c33e73e6eacbbd85ba2db55c1f9143b39f0

Observation 0c5d9294-2c5c-4955-820b-fd37b48e7a1b · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Adaptive Duration Model for Text Speech Alignment Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.167239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.167239Z digest=sha256:b6f5a409cc7b743e3e5625a9cc5c21f198209305811709e4dd21b2ed902ccf48

Observation d5b85757-2026-4848-bdb3-4bc26b481cdf · outbound

This paper cites Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis.

Adaptive Duration Model for Text Speech Alignment Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.232019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.232019Z digest=sha256:256ea4aee29e87adb7288693b072322bd19cf9e021c95744bc7c9abca3aba1e5

Observation 1dc12977-4893-42a9-9c18-67a74ec62837 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Adaptive Duration Model for Text Speech Alignment Fastspeech: Fast, robust and controllable text to speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.701373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.320168Z digest=sha256:7cec21eb798d06aa4f410afbbcd7eb402907b6706bf70bc3c3f5f961f1442dd6

Observation 03f6445e-4212-4a15-8f2a-1c798a8804c9 · outbound

This paper cites Fastpitch: Parallel text-to-speech with pitch prediction,.

Adaptive Duration Model for Text Speech Alignment Fastpitch: Parallel text-to-speech with pitch prediction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.524555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.378954Z digest=sha256:c5def6ab699c8e49754bfb29ce86e60840ebbb9149ba1ee1cdd9e07924b1daf5

Observation 10183caa-b808-4807-92b6-767070d9fb8e · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Adaptive Duration Model for Text Speech Alignment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.523005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.523005Z digest=sha256:277fe4eb82148b45382abb99e1bcc02ad6ad579006909c59950ccda225876108

Observation 849ba4da-4596-46cd-8456-028e7175abd4 · outbound

This paper cites Location-relative attention mechanisms for ro- bust long-form speech synthesis,.

Adaptive Duration Model for Text Speech Alignment Location-relative attention mechanisms for ro- bust long-form speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.359121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.586589Z digest=sha256:560b753c470fbd277de1b8ea46d5d648a75b69299a981a91dbd2573740852414

Observation 368df377-e60e-4603-b996-af6b64354af5 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Adaptive Duration Model for Text Speech Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.681690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.681690Z digest=sha256:115c7d979f2a4739637158d621e9dcc2fe75350f08950a6591b2a4afdaf721b2

Observation 71ddde11-7728-46f0-97c8-631b877d30c0 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Adaptive Duration Model for Text Speech Alignment NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.767437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.767437Z digest=sha256:4d672ced409a0a0ebc2411a70b024f140c354e56519e437b4aea23301bb697de

Observation 2a3f2b53-4432-4704-bf13-5ce8b7d01f08 · outbound

This paper cites Non-autoregressive neural text-to-speech,.

Adaptive Duration Model for Text Speech Alignment Non-autoregressive neural text-to-speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.191926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.866084Z digest=sha256:5c863148a34e19fbe4923fe6e8e2a593d5f7a3429d98a70deaf0bebe321c6294

Observation 21c31295-447d-46b1-8bc5-31f1d935c5ab · outbound

This paper cites Durian-e 2: Duration informed attention net- work with adaptive variational autoencoder and adver- sarial learning for expressive text-to-speech synthesis,.

Adaptive Duration Model for Text Speech Alignment Durian-e 2: Duration informed attention net- work with adaptive variational autoencoder and adver- sarial learning for expressive text-to-speech synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.038847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:08.959203Z digest=sha256:90135471390650ac3dc5da81e638628db758f50f77eca77e3edb98caff8125c5

Observation 6cbc73b0-781d-4ae4-99b2-0c2dc1b537e5 · outbound

This paper cites Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features.

Adaptive Duration Model for Text Speech Alignment Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:09.863491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:09.044690Z digest=sha256:1b33f6bc7be833fb3feadc2e3cacf3f3283a3a0ec16710e526647576cb8ec730

Observation 10f337af-c4fa-42b6-bcc3-fe35ea5ccb08 · outbound

This paper cites Simple-tts: End-to-end text-to-speech synthesis with latent diffusion,.

Adaptive Duration Model for Text Speech Alignment Simple-tts: End-to-end text-to-speech synthesis with latent diffusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.878139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:09.097184Z digest=sha256:54f5880aff7a2851a49d0dd8847a6c730b3907d3b4a41d625b7e52606b63c989

Observation 7f2d6d4c-ffdd-4fe9-8d08-605b21eeb050 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Adaptive Duration Model for Text Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.160142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.160142Z digest=sha256:8157c1f2118c814846a8569b9944f29336d66efc7caa4cf1830e090e5bf05b5c

Observation 06f29d60-5a31-40d7-a324-22c2f553eea3 · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

Adaptive Duration Model for Text Speech Alignment Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.727750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:09.231739Z digest=sha256:113903c27aebfa4e1fc920b779943efbff7fecbadf1aa9e618978501a80cb3b0

Observation cbfaa89b-f5e3-42c5-8734-c4c308da99e2 · outbound

This paper cites Portaspeech: Portable and high-quality generative text-to-speech,.

Adaptive Duration Model for Text Speech Alignment Portaspeech: Portable and high-quality generative text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.564583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:09.286028Z digest=sha256:3aa75c017ebec469a0d4d159e66b008728b30d46a7ef27fc890c9e0673d5f789

Observation a65be6c9-43c0-465b-a52b-8a8010999c00 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

Adaptive Duration Model for Text Speech Alignment V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.353445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:09.356034Z digest=sha256:fc0ac3a01984b03a72bc657bf71695232d6947835ff0741449c5c5aad9d222ac

Observation 752c8ba0-a959-4247-8d1e-c3737aee4cf3 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Adaptive Duration Model for Text Speech Alignment MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.403232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.403232Z digest=sha256:1341097f1751933657293a2ed649d54c1dc53fa66c95e9cf029791ae4b9a3420

Observation 9cb08e55-b39c-4f59-ab51-4d28892893d5 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

Adaptive Duration Model for Text Speech Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.476364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.476364Z digest=sha256:c04c3e8fd3e476a630046513d2615c2e457db222ec289296df7b5932bc8d57ad

Observation 56b91bcd-6dbd-4bf0-b070-a152a73ffcc4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Adaptive Duration Model for Text Speech Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.568289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.568289Z digest=sha256:0cb49ca36d1b98a904d52a97a5b4f656d07f9fd8966ccd4eb13999801da6687f

Observation f5a86135-f78e-4daf-8f2e-11ed8cd36458 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

Adaptive Duration Model for Text Speech Alignment WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.630347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.630347Z digest=sha256:89a4aad029d2400438c01c33c976f6a60176ad211a190a47855825133a4b2b3f

Pith citing papers

Observation e822238e-0ba3-457f-bee2-0dc2c323cba7 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment Adaptive Duration Model for Text Speech Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:10.121918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:34:07.882293Z digest=sha256:3b0777870a63c6028d0cc4802f15ba6af1a65801e1a4d70808daac53b353b68b