Pith. sign in

Paper Citation Record · LEDGER

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.19595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19595 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:55.714499Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:52.829136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:15:55.827429Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb7fc26a-ef85-4b17-9cd1-4f8198f10e5c · outbound

This paper cites an unresolved cited work.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:01.965263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:52.748091Z digest=sha256:db974c0aa80c96bed4c25aaf808c0842333b047e11973bc1e2f7b74b08c2ed87

Observation 7613c17a-498e-42a3-8d17-dd44120f1d7c · outbound

This paper cites Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:15:55.878979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:52.829136Z digest=sha256:777a9b3942b19fa5c85b4cfac2ee60837062998339b7a7b9b4d73279470da058

Observation 01dc9593-8e57-4e1e-a048-8da73aba0b7c · outbound

This paper cites Dataset We train our model on the LibriTTS [31] dataset, a multi- speaker English corpus containing approximately 585 hours of read speech sampled at 24 kHz.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Dataset We train our model on the LibriTTS [31] dataset, a multi- speaker English corpus containing approximately 585 hours of read speech sampled at 24 kHz

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.813339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:52.919656Z digest=sha256:f7a2bcfba3343ceb6eb41ab000fb2aa5d66920c2cdc668b8deeda41398ae901f

Observation 702f54ac-37a2-4adc-a8b5-3a746a42d4ac · outbound

This paper cites speech align.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment speech align

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.618499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.049253Z digest=sha256:784d7c95b406a9e6ba1aba045cfa25c1021154960d8091a57b2d48d425e3231a

Observation 0faadb96-b8ce-493a-b09d-7274d9b0a12a · outbound

This paper cites This ap- proach significantly accelerates model convergence while en- hancing the quality of generated speech.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment This ap- proach significantly accelerates model convergence while en- hancing the quality of generated speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.472750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.162304Z digest=sha256:8e4d21d4b268995533b81780ecb86691a5d820ac9ffe01a7568960e1dcc9b9f0

Observation 98f04b34-3288-47ce-a147-dd04790540d3 · outbound

This paper cites In future work, we plan to further op- timize the framework by reducing model complexity and accel- erating inference.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment In future work, we plan to further op- timize the framework by reducing model complexity and accel- erating inference

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.307516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.252705Z digest=sha256:5167e6665ca96a590143db52105d9d7708b2f0958935f085758c3f5a4d9414a3

Observation 258d8797-2481-4966-a231-dbe012515b96 · outbound

This paper cites Choi, J.-H.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Choi, J.-H

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.152594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.388149Z digest=sha256:30cec034194f3e9310882d249d39845cd308b823b9f8318c45cf2b4e65d409e5

Observation 6cc1c89c-7940-405c-8b17-9a854be9677f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.471613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.471613Z digest=sha256:141ac55349592ccfc23d0bcb68aaaafe79b5788c185573ffdc47577e9df94060

Observation 1d4d6018-6c07-457a-8bcd-920d8e4b9f3c · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.546922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.546922Z digest=sha256:21dc3b08e1970702950d5760dd3821e6cf9a343e6fa5d1920ad44234be9c67c7

Observation d2e23561-cb4c-4429-b577-f92f37f4fae7 · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V oicecraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.980493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.654282Z digest=sha256:30633298056e27ffca966384f95c36f6300161ccdbd94d9c77f5e40bdff7a589

Observation e6714b00-7472-44e4-8dd7-2f684454f6c5 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.775073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.775073Z digest=sha256:81abd4ceb1757c81a51811b0b8170a8c44f8657d1ca26e9e8249d9864574446b

Observation 94639417-d75f-4193-b59b-dd326c03cc02 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Autoregressive Speech Synthesis without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.849366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.849366Z digest=sha256:e28f7417db74a64c07035cfd921762f2d3f964ceef8ca706a536eaf401b13114

Observation e41f997a-0485-4e23-98ed-f51d6bd54884 · outbound

This paper cites V oice- box: Text-guided multilingual universal speech generation at scale,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V oice- box: Text-guided multilingual universal speech generation at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.771048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.908404Z digest=sha256:9ba30aed136d9bb10ef48850dd0b48f1a69584571c800238cac7684dfa993f6e

Observation 27bab5ff-4fe7-4713-8a79-6587a2a54bbc · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.611342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.958470Z digest=sha256:d7b2973569d23b2421d04e00c002cce6a9ed0079889a379dae50d30c7d5fcbf6

Observation 5c1f7394-f16b-432a-8e9c-fe282429a3cf · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.429665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.966196Z digest=sha256:c5e6018bf9c23f27d0cbf91c3e4998449be6fdf2d626892d35bc24dbebe814db

Observation 9596f8ff-8a8d-4e9a-8684-6b308bf5e031 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.976201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.976201Z digest=sha256:a85f35acf26b31e96df9752ca5f7cea5a5461fe85c2874d2d27e4f1df812553b

Observation c507b2b2-efe6-46d7-8b42-232f6c1a8f0c · outbound

This paper cites Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.150992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:53.988216Z digest=sha256:e72ff50d9f92b7f7ad71ca025f9ad83eddeda1bece3814221fd87b7706b739e4

Observation 5f3b1eba-920e-48f8-b4c0-9c14da64bfbe · outbound

This paper cites Accelerating codec-based speech synthesis with multi- token prediction and speculative decoding,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating codec-based speech synthesis with multi- token prediction and speculative decoding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.913822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.007318Z digest=sha256:85b33751427138cd70d0a1d936cc322a679cbfc048eb97987d255ec452b3aac6

Observation 578d911b-13af-4ec7-899d-516c8e495232 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.026219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.026219Z digest=sha256:c13481f8b34653916417c0a3d4dbfdd8bf6f63e8ce0aa6b1c360cbf81a124050

Observation 37e878e4-81ed-4917-b258-be620ad63813 · outbound

This paper cites Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.607420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.054161Z digest=sha256:90ba15cd716d01f125f00c78deaeacc5f837531bbead59ab0a3e8887559f3ded

Observation e4889d69-d0b6-4e18-94b5-6db5a6357eba · outbound

This paper cites Flow matching for generative modeling,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Flow matching for generative modeling,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.088901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.088901Z digest=sha256:85a1d1e19a6d53abb3fa9e56a9aa6d90e9dda20e01342b6e193f66d9df4b8edb

Observation 8e022587-4e6e-4444-8ab2-63d0c33e41af · outbound

This paper cites Scalable diffusion models with transform- ers,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Scalable diffusion models with transform- ers,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.124824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.124824Z digest=sha256:2230b4f425fc59a63571239bb16d74b548d6e484321c88495512459a07fdaaa9

Observation bf7e6c63-0eae-4340-ae2d-04e1ecd0ec97 · outbound

This paper cites Efficient diffusion training via min-snr weighting strat- egy,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Efficient diffusion training via min-snr weighting strat- egy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.323953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.161703Z digest=sha256:d1ec108ba37a15945d91b7ad1f52fbadc80df4a5276960a8ac48fac637fdb8db

Observation 7e709a39-c473-4410-b4e4-20bf79475049 · outbound

This paper cites Fasterdit: Towards faster diffusion transformers training without architecture modification,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Fasterdit: Towards faster diffusion transformers training without architecture modification,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.205251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.235356Z digest=sha256:4d2b8fedab4bd2cd8d3fee7e18f392acf1787b99ee9a8da9c38f8812230392fa

Observation e9e50506-c6bb-495b-98d0-e26ff9e6bf4d · outbound

This paper cites Representation alignment for generation: Training diffu- sion transformers is easier than you think,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Representation alignment for generation: Training diffu- sion transformers is easier than you think,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.124774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.329789Z digest=sha256:bdae81904023e8818e8fc174cb665da02a52c387458859db3e20873f03641634

Observation c2174071-aecc-4622-8714-9a44830d98e4 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment U-net: Convolutional networks for biomedical image segmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.895075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.427902Z digest=sha256:e27d84bf0ad099edc864e62c711a5d2245f19b41d2a8ad056f8e90f2533cbb8c

Observation 8d6618dd-846d-4943-9b46-f4f7c3b948a5 · outbound

This paper cites Attention is all you need,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.493555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.493555Z digest=sha256:c06bd0d43a98b1441dfa0321570ea126490315e1b9ebbd613cf47b36edd11be4

Observation a3b9ff3e-2fcf-4bf6-b73e-6b2a7743199f · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.617541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.571933Z digest=sha256:8491b556a6944d33edc8ca995cf168865dfec072c0d1fe3e25bcffdc7e111290

Observation 9f3a332f-2942-4462-a4da-9d66d61fe7de · outbound

This paper cites Intermediate loss regularization for ctc- based speech recognition,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Intermediate loss regularization for ctc- based speech recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.310597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.646730Z digest=sha256:471e00f3360c854a8b1a0c5aaeb0ef98e6cacb3e544de38a3b048770c0b12421

Observation 64a576d3-3564-423d-a154-bf4847c85770 · outbound

This paper cites Audio-visual efficient conformer for robust speech recognition,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Audio-visual efficient conformer for robust speech recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.067644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.691350Z digest=sha256:5203add0fa6f0cd53e463955e48d9cc706f5e5e974180635c0b485a2e4b76b3e

Observation 12af8411-e6e6-4a18-afc5-e109460d64a5 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.827126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.760242Z digest=sha256:9f15a658eff15c8036839a9750a995376e548ea2c1bed87f9ac582eb3ea5247d

Observation dfce062a-4b05-4457-ad8e-f42ebb911f27 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.646947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.810329Z digest=sha256:e006a99d5092b8b9ec5944f480462dfa859368f43cb827b37e91f88d56c796aa

Observation afad2a94-9925-4ab6-bb5f-c3a1503d52ce · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.497847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.866978Z digest=sha256:6b16f8ecd12a7e4b16c54a2c29076c3916a863c4f47ec40674fda26fb0948a50

Observation 803692c3-015b-466b-8cbe-2b7465842a7f · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.264673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:54.956368Z digest=sha256:66ad6ddb0fb23ed1c3053c3d5358dbcf0d5c2e69e7c34ba6faa3f391d52e274a

Observation 9d8afcab-636c-4abe-8b1d-ef87767ad15e · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.123506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.056539Z digest=sha256:58a07b47727d907169c36025eb8098a4a045d83e550497fc9e40221848a7967f

Observation e3641322-96fe-4ba8-8d41-73fac6cef310 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.127465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.127465Z digest=sha256:ab6425b498800ef3254231c9f423140da1c0860055ba4e7980210508b1c7510b

Observation 594d5d61-61ff-4221-becd-05815f436794 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.201763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.201763Z digest=sha256:7556f7b1b5925777b5a282567398585efb41a5e57bd185b05f763c7739e32c1d

Observation 35f5bfe1-7bfc-40ea-98c9-3231adacf766 · outbound

This paper cites Libritts: A corpus derived from librispeech for text- to-speech,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Libritts: A corpus derived from librispeech for text- to-speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.922439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.262774Z digest=sha256:66ebc3ce3b7ed69d40a2daba0cc814258e41bc74bf8400cabe47320bb15e3a09

Observation 80e28793-afc4-4762-8796-e84d592efcdd · outbound

This paper cites Lib- rispeech: An asr corpus based on public domain audio books,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Lib- rispeech: An asr corpus based on public domain audio books,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.795579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.319566Z digest=sha256:940cc7305eb180ba89d38ca1bd76bc24cc547b13ae23f8e52df7c6df35df8269

Observation 697a294c-cd83-46f6-bcda-2ebe01d793fb · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.689787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.396629Z digest=sha256:c6887306d83e6c2b881c35561cd99d3829c9c69bc35a87a18b4ed1ded773f97b

Observation 53c028ee-ac48-4ef9-be5f-9f625d05ba04 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Decoupled weight decay regulariza- tion,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.469364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.469364Z digest=sha256:8f92436edba2f0561079a1151f73fd044155f6b5c50e39d2d321a0979b72672c

Observation f1a62001-924d-4777-95bb-7a89c40b6917 · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.516535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.566232Z digest=sha256:4df7a7af72dcce07a519a3beba9bf89ddfefbf5bc39256850d1b037332f48c4d

Observation 03fb12a6-8584-435b-a51b-19c107223dbd · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.364968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.604506Z digest=sha256:d3997b9ece20fae9e08ccedf21bb71c1164f7d6b9f30be2b341e5d033f9ded6b

Observation f1d8e692-c73e-4f5f-b7e5-6a6bb1df7fa6 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Robust speech recognition via large-scale weak su- pervision,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.244826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.657348Z digest=sha256:843b075f6c33e2edd2fd1533ebe63430c82178a964c99aff37d0ba45ab6c083c

Observation 9292437f-a354-4340-a186-9c6b2b4d6cf5 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.033418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:55.714499Z digest=sha256:0529d895c8d4c3f66093eace9ce3206d1a5a8b1b0c8cf9e36bb5fd5f3679c84a

Pith citing papers

Observation 7613c17a-498e-42a3-8d17-dd44120f1d7c · inbound

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment cites this paper.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:15:55.878979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:15:52.829136Z digest=sha256:777a9b3942b19fa5c85b4cfac2ee60837062998339b7a7b9b4d73279470da058